would-agents-actually
Compare original and translation side by side
🇺🇸
Original
English🇨🇳
Translation
ChineseWould Agents Actually?
Agent真的会执行吗?
Produce an evidence-backed verdict about whether a pinned agent system will perform or avoid a defined action under declared tasks, tools, policies, budgets, and trials. Separate what the system says, what the trace records, and what the environment proves.
基于证据给出结论,判断固定的Agent系统在指定任务、工具、策略、预算和测试场景下,是否会执行或避免某一既定操作。区分系统声明内容、追踪记录内容以及环境验证结果。
Command grammar
命令语法
- : show the contract, verdict labels, and required system pin without researching.
/would-agents-actually help - : research the claim, run or inspect the required action evidence, issue the supported verdict, and validate the artifact.
/would-agents-actually verdict <agent-action question>
- :无需调研,显示协议、裁决标签和所需的系统固定配置。
/would-agents-actually help - :调研声明内容,运行或检查所需的操作证据,出具受支持的结论,并验证生成的产物。
/would-agents-actually verdict <agent-action question>
Procedure
流程
- Define the action as . Name the opportunity or trigger, pinned system, observable action or abstention, task distribution, environment, window, budgets, real friction, comparator, and independent readback. Ask one question only when an unknown changes the evidence plan. Otherwise state the assumption.
[a] - Pin the model identifier and version, provider route, prompts, runtime, planner, router, verifier, retries, fallbacks, tools, schemas, permissions, approval gates, memory, context, checkpoints, environment, dependencies, task sampling, and budgets. A material change creates a new system stratum.
- Write support, contradiction, and insufficient-evidence conditions before research or trials. Keep target outcomes, target traces, matched analogs, component checks, mechanism evidence, and inference separate. When current reports or fast-changing runtime behavior matters, load recent-public-signal.md.
- Verify every load-bearing source at its primary record. Complete the source and trace cards in verdict-protocol.md. Require two independent teams, task sets, or datasets for every premise needed by the verdict or confidence. Shared tasks, runs, graders, and restatements count once.
- Count an external action only after independent environment readback. A plan, final claim, tool-call attempt, accepted request, or long trace is not outcome proof. Keep model-caused and infrastructure-caused diagnostics separate while retaining both in the operational verdict.
- Define the reference class before reading results. Match action, tasks, horizon, tools, permissions, state, runtime, model version, environment, graders, budgets, failures, threat conditions, and time. Do not transport a score by model name alone.
- Load frameworks.md after evidence collection to diagnose opportunity, selection, attempt, receipt, state, constraints, verification, stopping, repeatability, and transport. Keep capability, propensity, reliability, compliance, resilience, abstention, and operational fit separate.
- Choose ,
LIKELY,UNLIKELY, orUNCERTAINunder verdict-protocol.md. If live research is unavailable or forbidden, useINSUFFICIENT EVIDENCE. Never invent a run, trace, tool call, state change, grader result, rate, cost, quote, source, or URL.UNVALIDATED HYPOTHESIS - Design the smallest production-like test with the least authority. Prespecify representative tasks, holdout, trials, denominator, comparator, graders, budgets, permissions, faults, stop rules, readback, rollback, cleanup, and the decision changed. Use approved or synthetic data and fake or reversible side effects.
- Render the result with output-template.md, save it outside the installed skill, then run . Exit 0 proves the artifact has the required shape. Exit 1 means it is incomplete. Exit 2 means the command or input path is wrong.
python3 scripts/validate_verdict.py --input <verdict.md> - Append queries, sources, system pins, task identity, excluded runs, failures, readback, assumptions, and validation output to an external research log after each consequential step. Stop and report the missing item when a load-bearing source, permission, readback, or safety control cannot be verified.
- 将操作定义为。明确触发机会、固定系统、可观察的操作或弃权行为、任务分配、环境、时间窗口、预算、实际摩擦因素、对比基准和独立反馈。仅当未知因素会改变证据计划时提出一个问题,否则说明假设前提。
[a] - 固定模型标识符和版本、提供商路由、提示词、运行时、规划器、路由器、验证器、重试机制、回退方案、工具、模式、权限、审批关卡、内存、上下文、检查点、环境、依赖项、任务抽样和预算。实质性变更会创建新的系统层级。
- 在调研或测试前,编写支持、矛盾和证据不足的判定条件。将目标结果、目标追踪记录、匹配的类比案例、组件检查、机制证据和推理过程分开处理。当当前报告或快速变化的运行时行为至关重要时,加载recent-public-signal.md。
- 在原始记录中验证每个关键来源。完成verdict-protocol.md中的来源和追踪卡片。结论或置信度所需的每个前提都需要两个独立团队、任务集或数据集的验证。共享任务、运行、评分器和重述仅计为一次。
- 仅在获得独立环境反馈后,才将外部操作计入结果。计划、最终声明、工具调用尝试、已接受的请求或长追踪记录都不能作为结果证明。在运营结论中,需区分由模型导致和由基础设施导致的诊断结果,并保留两者。
- 在查看结果前定义参考类别。匹配操作、任务、时间范围、工具、权限、状态、运行时、模型版本、环境、评分器、预算、故障、威胁条件和时间。不能仅通过模型名称来套用评分。
- 收集证据后加载frameworks.md,以诊断机会、选择、尝试、接收、状态、约束、验证、停止、可重复性和迁移情况。区分能力、倾向、可靠性、合规性、韧性、弃权和运营适配性。
- 根据verdict-protocol.md选择、
LIKELY、UNLIKELY或UNCERTAIN。如果无法进行或禁止实时调研,则使用INSUFFICIENT EVIDENCE。切勿虚构运行记录、追踪信息、工具调用、状态变更、评分结果、比率、成本、报价、来源或URL。UNVALIDATED HYPOTHESIS - 设计最小化的类生产测试,使用最低权限。预先指定代表性任务、保留数据集、测试次数、分母、对比基准、评分器、预算、权限、故障、停止规则、反馈、回滚、清理以及待变更的决策。使用已批准的或合成的数据,以及虚假或可逆的副作用。
- 使用output-template.md生成结果,将其保存到已安装Skill之外的位置,然后运行。退出码0表示产物符合要求格式,退出码1表示不完整,退出码2表示命令或输入路径错误。
python3 scripts/validate_verdict.py --input <verdict.md> - 在每个关键步骤后,将查询内容、来源、系统固定配置、任务标识、排除的运行记录、故障、反馈、假设前提和验证输出附加到外部研究日志中。当无法验证关键来源、权限、反馈或安全控制时,停止操作并报告缺失项。
Load conditions
加载条件
- Load evidence-base.md when benchmark, trial, grader, or transport evidence may inform the analysis, then recheck every primary source before use.
- Load verdict-template.md when creating a verdict. Copy it outside ; do not edit the installed template.
assets/ - Load help.md for the help command, verdict-insufficient-evidence.md for a researched verdict, and failure-unvalidated.md when the system pin or live research is missing.
- Run validate_verdict.py after writing the artifact. Read test_validate_verdict.py only when changing the validator contract.
- Read contract.md before changing behavior, trigger boundaries, or evaluation cases.
- Load generation-contract.md only when maintaining or repackaging this skill.
- 当基准测试、试用、评分器或迁移证据可能为分析提供信息时,加载evidence-base.md,然后在使用前重新检查每个原始来源。
- 创建结论时加载verdict-template.md。将其复制到目录外;请勿编辑已安装的模板。
assets/ - 帮助命令加载help.md,调研结论加载verdict-insufficient-evidence.md,当系统固定配置或实时调研缺失时加载failure-unvalidated.md。
- 编写产物后运行validate_verdict.py。仅在修改验证器协议时查看test_validate_verdict.py。
- 在修改行为、触发边界或评估案例前,阅读contract.md。
- 仅在维护或重新打包此Skill时加载generation-contract.md。
Gotchas
注意事项
- A successful component check does not prove the full action.
- ,
pass@k, single-run success, action propensity, and deployment reliability answer different questions.pass^k - An outcome grader, trace grader, constraint grader, and infrastructure grader cannot replace one another.
- Do not remove setup, dependency, transport, timeout, permission, or rate-limit failures from an operational denominator without reporting both views.
- Never add credentials, payment power, destructive access, or broad permissions merely to make a test realistic.
- In a sensitive domain, judge the defined action only. Do not infer overall safety, efficacy, legality, entitlement, authorization, compliance, or permission to act.
- 组件检查成功并不证明完整操作可行。
- 、
pass@k、单次运行成功、操作倾向和部署可靠性回答的是不同问题。pass^k - 结果评分器、追踪评分器、约束评分器和基础设施评分器不能互相替代。
- 未经报告两种视角,不得从运营分母中移除设置、依赖项、迁移、超时、权限或速率限制故障。
- 切勿仅为使测试更真实而添加凭证、支付权限、破坏性访问或广泛权限。
- 在敏感领域中,仅判断既定操作。不得推断整体安全性、有效性、合法性、资格、授权、合规性或操作权限。
Completion criteria
完成标准
- pins the system, action, opportunity, tasks, window, costs, comparator, and readback.
[a] - Every load-bearing premise passes the independence gate.
- Outcome, trace, constraints, grader validity, infrastructure, cost, and transport remain separate.
- Trials, eligible denominator, dropped runs, uncertainty, and task concentration are visible.
- Confidence is capped by the weakest premise, grader, independence check, and transport bridge.
- The next test uses least privilege, budgets, stop rules, readback, rollback, and cleanup.
- Every load-bearing source has a visible URL.
- The validator prints and exits 0.
"status": "PASS"
- 明确了系统、操作、机会、任务、时间窗口、成本、对比基准和反馈。
[a] - 每个关键前提都通过独立性检查。
- 结果、追踪记录、约束、评分器有效性、基础设施、成本和迁移情况保持独立区分。
- 测试次数、合格分母、排除的运行记录、不确定性和任务集中度清晰可见。
- 置信度受最弱前提、评分器、独立性检查和迁移桥梁的限制。
- 下一次测试使用最低权限、预算、停止规则、反馈、回滚和清理机制。
- 每个关键来源都有可见的URL。
- 验证器输出并返回退出码0。
"status": "PASS"