agent-workflow-playbook

Compare original and translation side by side

🇺🇸

Original

English
🇨🇳

Translation

Chinese

AI Agent Workflow Playbook — 从专家经验到可规模化交付

AI Agent Workflow Playbook — From Expert Experience to Scalable Delivery

适用于:把研究、营销、运营、分析、内容生产等高认知任务,改造成可测量、可纠错、可复用的 Agent 工作流。
Applicable to: Transforming high-cognitive tasks such as research, marketing, operations, analysis, and content production into measurable, correctable, and reusable Agent workflows.

先判断:这个任务该不该 Agent 化

First: Determine if the Task Should Be Agentized

只有同时满足以下多数条件才进入自动化:
  • 输入和合格输出可以被描述;
  • 专家能说清“什么是对、什么是错”;
  • 任务重复发生,或交付成本随客户数近似线性增长;
  • 关键数据能合法、稳定取得;
  • 错误可以在发布、付款、删除或对外发送前被拦截;
  • 结果能通过 rubric、样例集或业务指标复核。
如果任务低频、目标持续变化、没有验收口径,先做人工 SOP,不要先搭多 Agent。
Only proceed with automation if most of the following conditions are met:
  • Inputs and qualified outputs can be clearly defined;
  • Experts can articulate "what is right and what is wrong";
  • The task occurs repeatedly, or delivery costs grow approximately linearly with the number of clients;
  • Key data can be obtained legally and stably;
  • Errors can be intercepted before publication, payment, deletion, or external sending;
  • Results can be reviewed via rubrics, sample sets, or business metrics.
If the task is low-frequency, has constantly changing goals, or lacks acceptance criteria, first create a manual SOP instead of building a multi-Agent system.

1. 建立基线,不要直接写 Prompt

1. Establish a Baseline, Don’t Write Prompts Directly

选择 10–30 个近期真实任务,记录人工基线:
指标定义
任务成功率首次交付通过验收的任务数 / 总任务数
一次通过率无返工即通过的任务数 / 总任务数
周期从收到完整输入到可交付输出的 elapsed time
人工工时研究、制作、复核、返工所花人时
单次成本模型、工具、数据和人工复核成本之和
重试率发生工具重试或整段重做的任务占比
严重错误率错误发布、错误付款、数据泄露等高风险事件占比
没有这张基线表,就只能证明 Agent “能跑”,不能证明工作流变好了。
Select 10–30 recent real tasks and record the manual baseline:
MetricDefinition
Task Success RateNumber of tasks passing acceptance on first delivery / Total number of tasks
First-Pass RateNumber of tasks passing without rework / Total number of tasks
Cycle TimeElapsed time from receiving complete input to deliverable output
Manual Labor HoursTotal hours spent on research, production, review, and rework
Single Task CostSum of model, tool, data, and manual review costs
Retry RatePercentage of tasks requiring tool retries or full rework
Critical Error RatePercentage of high-risk incidents such as incorrect publication, wrong payment, or data leakage
Without this baseline table, you can only prove that the Agent "can run", not that the workflow has improved.

2. 从业务链路拆 Skill

2. Decompose Skills from Business Links

先画业务链路,再按可验收结果拆 skill:
text
需求澄清 → 数据获取 → 证据整理 → 分析 → 产出 → 质检 → 人工批准 → 交付 → 反馈沉淀
每个 skill 至少包含:
yaml
name: competitor-evidence-pack
input_contract:
  required: [product, market, competitors, time_window]
output_contract:
  required: [claims, source_urls, captured_at, confidence, unknowns]
tools:
  allow: [search, fetch]
  deny: [publish, delete, payment]
acceptance:
  - every material claim has a source
  - source capture time is recorded
  - unknown facts are labeled, not guessed
escalate_when:
  - authenticated source is inaccessible
  - sources conflict on a decision-critical fact
优先做单一职责 skill。只有当步骤间存在清晰依赖时,才增加 orchestrator。
First map the business link, then decompose skills based on acceptable outcomes:
text
Requirement Clarification → Data Acquisition → Evidence Organization → Analysis → Output Generation → Quality Inspection → Human Approval → Delivery → Feedback Documentation
Each skill must include at least:
yaml
name: competitor-evidence-pack
input_contract:
  required: [product, market, competitors, time_window]
output_contract:
  required: [claims, source_urls, captured_at, confidence, unknowns]
tools:
  allow: [search, fetch]
  deny: [publish, delete, payment]
acceptance:
  - every material claim has a source
  - source capture time is recorded
  - unknown facts are labeled, not guessed
escalate_when:
  - authenticated source is inaccessible
  - sources conflict on a decision-critical fact
Prioritize single-responsibility skills. Only add an orchestrator when there is a clear dependency between steps.

3. Harness:让系统知道边界、记住纠错、持续评测

3. Harness: Let the System Know Boundaries, Remember Corrections, and Continuously Evaluate

Prompt 只描述一次交互;harness 管理长期运行环境。至少包含五层:
  1. Context:品牌、客户、目标、禁区和数据权限;
  2. Skills:通用技能与客户专属技能分离,按任务选择调用;
  3. Memory:只沉淀经过确认的偏好、错误和纠正,不把猜测写成事实;
  4. Evaluation:固定样例集、rubric、回归测试和业务指标;
  5. Observability:每步输入摘要、工具调用、证据、成本、耗时、重试和最终批准人。
一次失败的正确处理方式不是无限加提示词,而是:记录失败类型 → 判断是数据、工具、推理还是验收问题 → 修改对应层 → 用旧样例集回归。
A Prompt only describes a single interaction; a harness manages the long-running environment. It should include at least five layers:
  1. Context: Brand, client, objectives, forbidden areas, and data permissions;
  2. Skills: Separate general skills from client-specific skills, and select calls based on tasks;
  3. Memory: Only document confirmed preferences, errors, and corrections, not guesses as facts;
  4. Evaluation: Fixed sample sets, rubrics, regression tests, and business metrics;
  5. Observability: Summary of input at each step, tool calls, evidence, cost, time spent, retries, and final approver.
The correct way to handle a failure is not to infinitely add prompts, but to: record the failure type → determine if it is a data, tool, reasoning, or acceptance issue → modify the corresponding layer → regression test with the old sample set.

4. 选择编排方式

4. Choose an Orchestration Method

模式适用情况主要风险
顺序后一步严格依赖前一步输出上游错误级联
并行多个独立来源或方案可同时产生合并冲突、重复成本
路由不同任务应调用不同专长分类错误
主管—执行者任务可拆成多个独立子任务主管成为瓶颈
评审—修订输出有明确 rubric,可迭代改进无界循环、成本失控
默认从单 Agent + 多 skill 开始。只有观测数据证明吞吐或专长隔离确实需要并发,才升级为多 Agent。
ModeApplicable ScenariosMain Risks
SequentialThe next step strictly depends on the output of the previous stepUpstream error cascading
ParallelMultiple independent sources or solutions can be generated simultaneouslyMerge conflicts, repeated costs
RoutingDifferent tasks should call different expertiseClassification errors
Supervisor-ExecutorTasks can be split into multiple independent subtasksSupervisor becomes a bottleneck
Review-RevisionOutput has a clear rubric and can be iteratively improvedUnbounded loops, cost out of control
Start with a single Agent + multiple skills by default. Only upgrade to multi-Agent when observational data proves that concurrency is truly needed for throughput or expertise isolation.

5. 人工介入与权限

5. Human Intervention and Permissions

以下动作默认需要人工批准:
  • 对外发布、群发、私信或代表个人表态;
  • 付款、退款、采购和价格承诺;
  • 删除、覆盖或批量修改数据;
  • 使用未获授权的个人数据;
  • 低置信度但会影响客户决策的结论。
连续任务不要重复索取同一授权;记录授权对象、范围和有效期。权限不足时返回缺失项和恢复路径,不要假装完成。
The following actions require human approval by default:
  • External publication, mass messaging, private messaging, or representing individuals to express opinions;
  • Payments, refunds, procurement, and price commitments;
  • Deletion, overwriting, or batch modification of data;
  • Use of unauthorized personal data;
  • Conclusions with low confidence but that will affect client decisions.
Do not repeatedly request the same authorization for continuous tasks; record the authorized object, scope, and validity period. When permissions are insufficient, return the missing items and recovery path instead of pretending to complete the task.

6. 上线门槛与回滚

6. Launch Threshold and Rollback

按四阶段推进:
  1. Shadow:Agent 生成结果但不影响人工交付;
  2. Copilot:人工选择、修改并批准每次输出;
  3. Guarded automation:低风险步骤自动执行,高风险动作审批;
  4. Autonomous:仅用于已稳定通过回归测试、可完整审计且可回滚的边界任务。
每次版本变更比较同一批任务的成功率、周期、人工工时、成本和严重错误率。任一安全指标恶化,回滚到上一稳定版本。
Advance in four phases:
  1. Shadow: The Agent generates results but does not affect manual delivery;
  2. Copilot: Humans select, modify, and approve each output;
  3. Guarded automation: Low-risk steps are executed automatically, high-risk actions require approval;
  4. Autonomous: Only used for boundary tasks that have stably passed regression tests, are fully auditable, and can be rolled back.
Compare the success rate, cycle time, manual labor hours, cost, and critical error rate of the same batch of tasks for each version change. If any safety indicator deteriorates, roll back to the previous stable version.

真实案例:营销洞察交付从项目制走向产品化

Real Case: Transforming Marketing Insight Delivery from Project-Based to Productized

来源:Gingiris 飞书会议纪要《AI agent实践落地困境与规模化尝试分析》,2026-05-10。以下是会议中的经验陈述,不是独立审计或随机对照实验。
Source: Gingiris Feishu Meeting Minutes "Analysis of AI Agent Practice Implementation Dilemmas and Scaling Attempts", May 10, 2026. The following are experience statements from the meeting, not independent audits or randomized controlled trials.

旧链路

Old Workflow

  • 运营人员手动浏览小红书约 10 天到 2 周;
  • 早期自动化把信息收集压缩到数小时,但报告仍依赖人工思考、迭代和制图;
  • 一个高质量 PPT 交付需要约 15 人、3–4 周,难以复制到 100 或 1,000 家客户。
  • Operations staff manually browsed Xiaohongshu for about 10 days to 2 weeks;
  • Early automation reduced information collection to a few hours, but reports still relied on human thinking, iteration, and charting;
  • Delivering a high-quality PPT required about 15 people and 3–4 weeks, making it difficult to replicate for 100 or 1,000 clients.

新链路

New Workflow

  • 将达人筛选、赛道分析、人群识别、内容与投放建议拆成通用 skill 和客户专属 skill;
  • 把正例、反例、验收标准、记忆和反馈链路放进 harness;
  • 系统先识别任务,再组合调用技能;专家负责策略判断与最终验收;
  • 结果、洞察过程和纠错留在系统中,供下一次复用。
  • Decomposed talent screening, track analysis, crowd identification, content and delivery suggestions into general skills and client-specific skills;
  • Integrated positive examples, negative examples, acceptance criteria, memory, and feedback links into the harness;
  • The system first identifies the task, then combines and calls skills; experts are responsible for strategic judgment and final acceptance;
  • Results, insight processes, and corrections are retained in the system for reuse in future tasks.

已报告结果

Reported Results

  • 15 人 × 3–4 周 的 PPT 项目,变为 1 名策略师 + AI 系统,5 天一次通过
  • 一份竞品分析与投放建议可在 不到 1 天 内交付;
  • 会议报告的服务成本相较早期方式下降 两个数量级
  • 专业策略师仍需约 2–3 个月 学习并迁移到新工作方式,说明专家并没有被“零成本替代”。
  • A PPT project that originally required 15 people × 3–4 weeks became 1 strategist + AI system, passing on the first try in 5 days;
  • A competitor analysis and delivery suggestion can be delivered in less than 1 day;
  • The service cost of meeting reports has decreased by two orders of magnitude compared to early methods;
  • Professional strategists still need about 2–3 months to learn and transition to the new working method, indicating that experts are not "zero-cost replaced".

这个案例真正验证了什么

What This Case Actually Validates

它支持“把专家判断编码为 skill + harness,可以减少交付周期和边际人工”的判断;它不证明所有行业都能获得相同幅度,也没有披露统一口径下的错误率、模型成本和客户长期留存。因此复用时必须重新建立本团队基线,并补测质量、重试和严重错误率。
It supports the judgment that "coding expert judgment into skill + harness can reduce delivery cycles and marginal labor"; it does not prove that all industries can achieve the same magnitude of improvement, nor does it disclose error rates, model costs, and long-term client retention under a unified standard. Therefore, when replicating, you must re-establish your team's baseline and supplement tests for quality, retries, and critical error rates.

7 天落地清单

7-Day Implementation Checklist

  • Day 1:选一个高频任务,收集 10–30 个真实样例和人工基线;
  • Day 2:定义输入、输出、rubric、禁区与人工审批点;
  • Day 3:拆成 3–7 个单一职责 skill;
  • Day 4:接入证据记录、日志、成本与失败分类;
  • Day 5:Shadow replay,修复最高频失败;
  • Day 6:Copilot 小流量运行,比较人工基线;
  • Day 7:决定继续、回滚或只自动化其中一段。
  • Day 1: Select a high-frequency task, collect 10–30 real samples and manual baselines;
  • Day 2: Define inputs, outputs, rubrics, forbidden areas, and human approval points;
  • Day 3: Decompose into 3–7 single-responsibility skills;
  • Day 4: Integrate evidence recording, logs, cost, and failure classification;
  • Day 5: Shadow replay, fix the most frequent failures;
  • Day 6: Run small traffic with Copilot, compare with manual baselines;
  • Day 7: Decide to continue, roll back, or automate only a segment.

最终输出模板

Final Output Template

每次工作流评审必须交付:
  1. 业务链路和自动化边界;
  2. skill 清单与输入/输出 contract;
  3. 编排图和工具权限;
  4. 评测集、基线和本次结果;
  5. 人工介入、失败恢复与回滚方案;
  6. 下一轮只改一个变量的实验计划。

Built by Gingiris. Historical figures are labeled as reported case evidence; do not present them as guaranteed outcomes.
Each workflow review must deliver:
  1. Business links and automation boundaries;
  2. Skill list and input/output contracts;
  3. Orchestration diagram and tool permissions;
  4. Evaluation set, baseline, and current results;
  5. Human intervention, failure recovery, and rollback plans;
  6. Experimental plan to modify only one variable in the next round.

Built by Gingiris. Historical figures are labeled as reported case evidence; do not present them as guaranteed outcomes.