agent-workflow-playbook
Compare original and translation side by side
🇺🇸
Original
English🇨🇳
Translation
ChineseAI Agent Workflow Playbook — 从专家经验到可规模化交付
AI Agent Workflow Playbook — From Expert Experience to Scalable Delivery
适用于:把研究、营销、运营、分析、内容生产等高认知任务,改造成可测量、可纠错、可复用的 Agent 工作流。
Applicable to: Transforming high-cognitive tasks such as research, marketing, operations, analysis, and content production into measurable, correctable, and reusable Agent workflows.
先判断:这个任务该不该 Agent 化
First: Determine if the Task Should Be Agentized
只有同时满足以下多数条件才进入自动化:
- 输入和合格输出可以被描述;
- 专家能说清“什么是对、什么是错”;
- 任务重复发生,或交付成本随客户数近似线性增长;
- 关键数据能合法、稳定取得;
- 错误可以在发布、付款、删除或对外发送前被拦截;
- 结果能通过 rubric、样例集或业务指标复核。
如果任务低频、目标持续变化、没有验收口径,先做人工 SOP,不要先搭多 Agent。
Only proceed with automation if most of the following conditions are met:
- Inputs and qualified outputs can be clearly defined;
- Experts can articulate "what is right and what is wrong";
- The task occurs repeatedly, or delivery costs grow approximately linearly with the number of clients;
- Key data can be obtained legally and stably;
- Errors can be intercepted before publication, payment, deletion, or external sending;
- Results can be reviewed via rubrics, sample sets, or business metrics.
If the task is low-frequency, has constantly changing goals, or lacks acceptance criteria, first create a manual SOP instead of building a multi-Agent system.
1. 建立基线,不要直接写 Prompt
1. Establish a Baseline, Don’t Write Prompts Directly
选择 10–30 个近期真实任务,记录人工基线:
| 指标 | 定义 |
|---|---|
| 任务成功率 | 首次交付通过验收的任务数 / 总任务数 |
| 一次通过率 | 无返工即通过的任务数 / 总任务数 |
| 周期 | 从收到完整输入到可交付输出的 elapsed time |
| 人工工时 | 研究、制作、复核、返工所花人时 |
| 单次成本 | 模型、工具、数据和人工复核成本之和 |
| 重试率 | 发生工具重试或整段重做的任务占比 |
| 严重错误率 | 错误发布、错误付款、数据泄露等高风险事件占比 |
没有这张基线表,就只能证明 Agent “能跑”,不能证明工作流变好了。
Select 10–30 recent real tasks and record the manual baseline:
| Metric | Definition |
|---|---|
| Task Success Rate | Number of tasks passing acceptance on first delivery / Total number of tasks |
| First-Pass Rate | Number of tasks passing without rework / Total number of tasks |
| Cycle Time | Elapsed time from receiving complete input to deliverable output |
| Manual Labor Hours | Total hours spent on research, production, review, and rework |
| Single Task Cost | Sum of model, tool, data, and manual review costs |
| Retry Rate | Percentage of tasks requiring tool retries or full rework |
| Critical Error Rate | Percentage of high-risk incidents such as incorrect publication, wrong payment, or data leakage |
Without this baseline table, you can only prove that the Agent "can run", not that the workflow has improved.
2. 从业务链路拆 Skill
2. Decompose Skills from Business Links
先画业务链路,再按可验收结果拆 skill:
text
需求澄清 → 数据获取 → 证据整理 → 分析 → 产出 → 质检 → 人工批准 → 交付 → 反馈沉淀每个 skill 至少包含:
yaml
name: competitor-evidence-pack
input_contract:
required: [product, market, competitors, time_window]
output_contract:
required: [claims, source_urls, captured_at, confidence, unknowns]
tools:
allow: [search, fetch]
deny: [publish, delete, payment]
acceptance:
- every material claim has a source
- source capture time is recorded
- unknown facts are labeled, not guessed
escalate_when:
- authenticated source is inaccessible
- sources conflict on a decision-critical fact优先做单一职责 skill。只有当步骤间存在清晰依赖时,才增加 orchestrator。
First map the business link, then decompose skills based on acceptable outcomes:
text
Requirement Clarification → Data Acquisition → Evidence Organization → Analysis → Output Generation → Quality Inspection → Human Approval → Delivery → Feedback DocumentationEach skill must include at least:
yaml
name: competitor-evidence-pack
input_contract:
required: [product, market, competitors, time_window]
output_contract:
required: [claims, source_urls, captured_at, confidence, unknowns]
tools:
allow: [search, fetch]
deny: [publish, delete, payment]
acceptance:
- every material claim has a source
- source capture time is recorded
- unknown facts are labeled, not guessed
escalate_when:
- authenticated source is inaccessible
- sources conflict on a decision-critical factPrioritize single-responsibility skills. Only add an orchestrator when there is a clear dependency between steps.
3. Harness:让系统知道边界、记住纠错、持续评测
3. Harness: Let the System Know Boundaries, Remember Corrections, and Continuously Evaluate
Prompt 只描述一次交互;harness 管理长期运行环境。至少包含五层:
- Context:品牌、客户、目标、禁区和数据权限;
- Skills:通用技能与客户专属技能分离,按任务选择调用;
- Memory:只沉淀经过确认的偏好、错误和纠正,不把猜测写成事实;
- Evaluation:固定样例集、rubric、回归测试和业务指标;
- Observability:每步输入摘要、工具调用、证据、成本、耗时、重试和最终批准人。
一次失败的正确处理方式不是无限加提示词,而是:记录失败类型 → 判断是数据、工具、推理还是验收问题 → 修改对应层 → 用旧样例集回归。
A Prompt only describes a single interaction; a harness manages the long-running environment. It should include at least five layers:
- Context: Brand, client, objectives, forbidden areas, and data permissions;
- Skills: Separate general skills from client-specific skills, and select calls based on tasks;
- Memory: Only document confirmed preferences, errors, and corrections, not guesses as facts;
- Evaluation: Fixed sample sets, rubrics, regression tests, and business metrics;
- Observability: Summary of input at each step, tool calls, evidence, cost, time spent, retries, and final approver.
The correct way to handle a failure is not to infinitely add prompts, but to: record the failure type → determine if it is a data, tool, reasoning, or acceptance issue → modify the corresponding layer → regression test with the old sample set.
4. 选择编排方式
4. Choose an Orchestration Method
| 模式 | 适用情况 | 主要风险 |
|---|---|---|
| 顺序 | 后一步严格依赖前一步输出 | 上游错误级联 |
| 并行 | 多个独立来源或方案可同时产生 | 合并冲突、重复成本 |
| 路由 | 不同任务应调用不同专长 | 分类错误 |
| 主管—执行者 | 任务可拆成多个独立子任务 | 主管成为瓶颈 |
| 评审—修订 | 输出有明确 rubric,可迭代改进 | 无界循环、成本失控 |
默认从单 Agent + 多 skill 开始。只有观测数据证明吞吐或专长隔离确实需要并发,才升级为多 Agent。
| Mode | Applicable Scenarios | Main Risks |
|---|---|---|
| Sequential | The next step strictly depends on the output of the previous step | Upstream error cascading |
| Parallel | Multiple independent sources or solutions can be generated simultaneously | Merge conflicts, repeated costs |
| Routing | Different tasks should call different expertise | Classification errors |
| Supervisor-Executor | Tasks can be split into multiple independent subtasks | Supervisor becomes a bottleneck |
| Review-Revision | Output has a clear rubric and can be iteratively improved | Unbounded loops, cost out of control |
Start with a single Agent + multiple skills by default. Only upgrade to multi-Agent when observational data proves that concurrency is truly needed for throughput or expertise isolation.
5. 人工介入与权限
5. Human Intervention and Permissions
以下动作默认需要人工批准:
- 对外发布、群发、私信或代表个人表态;
- 付款、退款、采购和价格承诺;
- 删除、覆盖或批量修改数据;
- 使用未获授权的个人数据;
- 低置信度但会影响客户决策的结论。
连续任务不要重复索取同一授权;记录授权对象、范围和有效期。权限不足时返回缺失项和恢复路径,不要假装完成。
The following actions require human approval by default:
- External publication, mass messaging, private messaging, or representing individuals to express opinions;
- Payments, refunds, procurement, and price commitments;
- Deletion, overwriting, or batch modification of data;
- Use of unauthorized personal data;
- Conclusions with low confidence but that will affect client decisions.
Do not repeatedly request the same authorization for continuous tasks; record the authorized object, scope, and validity period. When permissions are insufficient, return the missing items and recovery path instead of pretending to complete the task.
6. 上线门槛与回滚
6. Launch Threshold and Rollback
按四阶段推进:
- Shadow:Agent 生成结果但不影响人工交付;
- Copilot:人工选择、修改并批准每次输出;
- Guarded automation:低风险步骤自动执行,高风险动作审批;
- Autonomous:仅用于已稳定通过回归测试、可完整审计且可回滚的边界任务。
每次版本变更比较同一批任务的成功率、周期、人工工时、成本和严重错误率。任一安全指标恶化,回滚到上一稳定版本。
Advance in four phases:
- Shadow: The Agent generates results but does not affect manual delivery;
- Copilot: Humans select, modify, and approve each output;
- Guarded automation: Low-risk steps are executed automatically, high-risk actions require approval;
- Autonomous: Only used for boundary tasks that have stably passed regression tests, are fully auditable, and can be rolled back.
Compare the success rate, cycle time, manual labor hours, cost, and critical error rate of the same batch of tasks for each version change. If any safety indicator deteriorates, roll back to the previous stable version.
真实案例:营销洞察交付从项目制走向产品化
Real Case: Transforming Marketing Insight Delivery from Project-Based to Productized
来源:Gingiris 飞书会议纪要《AI agent实践落地困境与规模化尝试分析》,2026-05-10。以下是会议中的经验陈述,不是独立审计或随机对照实验。
Source: Gingiris Feishu Meeting Minutes "Analysis of AI Agent Practice Implementation Dilemmas and Scaling Attempts", May 10, 2026. The following are experience statements from the meeting, not independent audits or randomized controlled trials.
旧链路
Old Workflow
- 运营人员手动浏览小红书约 10 天到 2 周;
- 早期自动化把信息收集压缩到数小时,但报告仍依赖人工思考、迭代和制图;
- 一个高质量 PPT 交付需要约 15 人、3–4 周,难以复制到 100 或 1,000 家客户。
- Operations staff manually browsed Xiaohongshu for about 10 days to 2 weeks;
- Early automation reduced information collection to a few hours, but reports still relied on human thinking, iteration, and charting;
- Delivering a high-quality PPT required about 15 people and 3–4 weeks, making it difficult to replicate for 100 or 1,000 clients.
新链路
New Workflow
- 将达人筛选、赛道分析、人群识别、内容与投放建议拆成通用 skill 和客户专属 skill;
- 把正例、反例、验收标准、记忆和反馈链路放进 harness;
- 系统先识别任务,再组合调用技能;专家负责策略判断与最终验收;
- 结果、洞察过程和纠错留在系统中,供下一次复用。
- Decomposed talent screening, track analysis, crowd identification, content and delivery suggestions into general skills and client-specific skills;
- Integrated positive examples, negative examples, acceptance criteria, memory, and feedback links into the harness;
- The system first identifies the task, then combines and calls skills; experts are responsible for strategic judgment and final acceptance;
- Results, insight processes, and corrections are retained in the system for reuse in future tasks.
已报告结果
Reported Results
- 约 15 人 × 3–4 周 的 PPT 项目,变为 1 名策略师 + AI 系统,5 天一次通过;
- 一份竞品分析与投放建议可在 不到 1 天 内交付;
- 会议报告的服务成本相较早期方式下降 两个数量级;
- 专业策略师仍需约 2–3 个月 学习并迁移到新工作方式,说明专家并没有被“零成本替代”。
- A PPT project that originally required 15 people × 3–4 weeks became 1 strategist + AI system, passing on the first try in 5 days;
- A competitor analysis and delivery suggestion can be delivered in less than 1 day;
- The service cost of meeting reports has decreased by two orders of magnitude compared to early methods;
- Professional strategists still need about 2–3 months to learn and transition to the new working method, indicating that experts are not "zero-cost replaced".
这个案例真正验证了什么
What This Case Actually Validates
它支持“把专家判断编码为 skill + harness,可以减少交付周期和边际人工”的判断;它不证明所有行业都能获得相同幅度,也没有披露统一口径下的错误率、模型成本和客户长期留存。因此复用时必须重新建立本团队基线,并补测质量、重试和严重错误率。
It supports the judgment that "coding expert judgment into skill + harness can reduce delivery cycles and marginal labor"; it does not prove that all industries can achieve the same magnitude of improvement, nor does it disclose error rates, model costs, and long-term client retention under a unified standard. Therefore, when replicating, you must re-establish your team's baseline and supplement tests for quality, retries, and critical error rates.
7 天落地清单
7-Day Implementation Checklist
- Day 1:选一个高频任务,收集 10–30 个真实样例和人工基线;
- Day 2:定义输入、输出、rubric、禁区与人工审批点;
- Day 3:拆成 3–7 个单一职责 skill;
- Day 4:接入证据记录、日志、成本与失败分类;
- Day 5:Shadow replay,修复最高频失败;
- Day 6:Copilot 小流量运行,比较人工基线;
- Day 7:决定继续、回滚或只自动化其中一段。
- Day 1: Select a high-frequency task, collect 10–30 real samples and manual baselines;
- Day 2: Define inputs, outputs, rubrics, forbidden areas, and human approval points;
- Day 3: Decompose into 3–7 single-responsibility skills;
- Day 4: Integrate evidence recording, logs, cost, and failure classification;
- Day 5: Shadow replay, fix the most frequent failures;
- Day 6: Run small traffic with Copilot, compare with manual baselines;
- Day 7: Decide to continue, roll back, or automate only a segment.
最终输出模板
Final Output Template
每次工作流评审必须交付:
- 业务链路和自动化边界;
- skill 清单与输入/输出 contract;
- 编排图和工具权限;
- 评测集、基线和本次结果;
- 人工介入、失败恢复与回滚方案;
- 下一轮只改一个变量的实验计划。
Built by Gingiris. Historical figures are labeled as reported case evidence; do not present them as guaranteed outcomes.
Each workflow review must deliver:
- Business links and automation boundaries;
- Skill list and input/output contracts;
- Orchestration diagram and tool permissions;
- Evaluation set, baseline, and current results;
- Human intervention, failure recovery, and rollback plans;
- Experimental plan to modify only one variable in the next round.
Built by Gingiris. Historical figures are labeled as reported case evidence; do not present them as guaranteed outcomes.