harness-master
Compare original and translation side by side
🇺🇸
Original
English🇨🇳
Translation
ChineseHarness Master
Harness Master
Audit AI harness configuration quality, usage posture, and ecosystem options, then apply approved config fixes in the same session.
NOT for: creating agents (agent-conventions), building MCP servers (mcp-creator), generic application telemetry (observability-advisor), or generic code review without harness/config focus (review).
Input: harness names, level, or natural language. Modes: Audit, Discover (gap + bounded research), Usage, Apply/Install, Journal. Run when mode is ambiguous.
scripts/discovery/classify_intent.py --args "$ARGUMENTS" --json审计AI Harness的配置质量、使用状态和生态系统选项,并在同一会话中应用已获批的配置修复方案。
不适用于: 创建Agent(遵循agent-conventions规范)、搭建MCP服务器(使用mcp-creator)、通用应用遥测(使用observability-advisor)或不聚焦Harness/配置的通用代码评审(使用review)。
输入: Harness名称、级别或自然语言。模式:审计(Audit)、发现(Discover)(缺口分析+限定范围研究)、使用情况(Usage)、应用/安装(Apply/Install)、日志记录(Journal)。当模式不明确时,运行。
scripts/discovery/classify_intent.py --args "$ARGUMENTS" --jsonDispatch
调度规则
| Mode | First move |
|---|---|---|
| Intake | Ask for harnesses or |
| Intake | Ask only for level |
| Audit | Dry-run all supported harnesses in deterministic order |
| Intake | Resolve one or more harnesses, then ask only for missing level |
| Intake | Resolve level, then ask only for missing harnesses |
| Audit | Run dry-run review for the selected harnesses and level |
| Apply Approved | Apply the last matching dry-run batch only if the user already approved it in this session |
| Apply Approved | Apply only if the same scope was already reviewed and approved in this session; otherwise rerun audit first |
| Install Guidance | Show exact |
| Discover full | W0 scans → scouts → ideate → report per |
| Discover W0 | Gap report only; no scouts |
| Journal | Load or list |
| Usage Review | Collect safe aggregate usage/cost/token/tool signals; read-only |
| NL: missing skills / expand / find best X for harness | Discover (inferred) | Depth: full, focused, candidate, or compare — see discovery-pipeline.md |
| Natural language: "review/audit/check/tune <harness> config" | Audit | Normalize harnesses + level, then run dry-run review |
| Natural language: "review/analyze usage/cost/tokens/tools for <harness>" | Usage Review | Normalize harnesses + level, collect safe metadata only, and report optimization lanes |
| Natural language: "find/latest/best plugins/extensions/MCPs/skills for <harness>" | Discover focused | W0 + category scouts + |
| Natural language approval like "approved", "do it", "apply those changes" | Apply Approved | Continue only if the immediately preceding |
| Requests to create agents or MCP servers | Refuse + redirect | Redirect to the correct specialized skill |
| Unsupported harness names | Refuse + clarify | List supported harnesses and ask the user to choose from that set |
| 模式 | 第一步操作 |
|---|---|---|
| 初始引导 | 询问要评审的Harness或输入 |
| 初始引导 | 仅询问范围级别 |
| 审计 | 按确定顺序对所有支持的Harness执行试运行审计 |
| 初始引导 | 解析一个或多个Harness,然后仅询问缺失的范围级别 |
| 初始引导 | 解析范围级别,然后仅询问缺失的Harness |
| 审计 | 对选定的Harness和级别执行试运行评审 |
| 应用已获批方案 | 仅当用户已在本次会话中批准过匹配的试运行批次时,应用最后一次匹配的试运行结果 |
| 应用已获批方案 | 仅当相同范围已在本次会话中完成评审并获批准时执行应用;否则先重新运行审计 |
| 安装指导 | 仅展示确切的 |
| 完整发现模式 | 执行W0扫描 → 侦察 → 构思 → 按 |
| 基础发现(W0) | 仅生成缺口报告;不执行侦察操作 |
| 日志记录 | 加载或列出 |
| 使用情况评审 | 收集安全的聚合使用/成本/令牌/工具数据信号;仅只读操作 |
| 自然语言:missing skills / expand / find best X for harness | 推断发现模式 | 深度:完整、聚焦、候选或对比 — 详见discovery-pipeline.md |
| 自然语言:"review/audit/check/tune <harness> config" | 审计 | 标准化Harness和级别,然后运行试运行评审 |
| 自然语言:"review/analyze usage/cost/tokens/tools for <harness>" | 使用情况评审 | 标准化Harness和级别,仅收集安全元数据,并报告优化方向 |
| 自然语言:"find/latest/best plugins/extensions/MCPs/skills for <harness>" | 聚焦发现模式 | 执行W0 + 分类侦察 + |
| 自然语言批准如"approved", "do it", "apply those changes" | 应用已获批方案 | 仅当紧接的 |
| 创建Agent或MCP服务器的请求 | 拒绝并重定向 | 重定向至对应的专用技能 |
| 不支持的Harness名称 | 拒绝并澄清 | 列出支持的Harness并要求用户从中选择 |
Empty-Args Handler
空参数处理
If the user invokes with no arguments:
/harness-master- Ask which harnesses to review or whether to use .
all - Ask for ,
project, orglobal.both - Explain that the first pass is always dry-run only.
如果用户调用时未传入参数:
/harness-master- 询问要评审的Harness或是否使用。
all - 询问范围是、
project还是global。both - 说明首次执行始终为试运行模式。
Input Normalization
输入标准化
- Supported harnesses: ,
claude-code,claude-desktop,chatgpt,codex,cursor,cursor-cli,grok-build,opencode,perplexity-desktop,cherry-studiolm-studio - Supported levels: ,
project,globalboth - Supported research categories: ,
config,plugin,extension,mcp,skillall - Supported usage windows: positive day counts such as ,
7,14, or30; default todays=14when absent14 - Harness aliases:
- ,
claude->claude-codeclaude-code - ->
claude-desktopclaude-desktop - ,
chatgpt,chatgpt-desktop->openai-chatgptchatgpt - ->
codexcodex - ,
cursor,cursor-editor,cursor-desktop,cursor-web,cursor-cloud,cursor-cloud-agent->cursor-background-agentcursor - ,
cursor-cli,cursor-agent,cursor-agent-cli,agent-cli->agentcursor-cli - ,
grok,grok-build->grok-cligrok-build - ,
opencode->open-codeopencode - ,
perplexity,perplexity-desktop->perplexity-macperplexity-desktop - ,
cherry,cherrystudio,cherry-ai->cherry-studiocherry-studio - ,
lm-studio,lmstudio->lmslm-studio
- Level aliases:
- ,
project,repo->localproject - ,
global->userglobal - ,
both->all-levelsboth
- Split multiple harnesses on commas and whitespace.
- If the user supplies both and named harnesses, ask which form they want.
all - Deterministic order:
all,claude-code,claude-desktop,chatgpt,codex,cursor,cursor-cli,grok-build,opencode,perplexity-desktop,cherry-studiolm-studio - Unknown tokens are never guessed. Ask one focused clarification question.
- 支持的Harness:,
claude-code,claude-desktop,chatgpt,codex,cursor,cursor-cli,grok-build,opencode,perplexity-desktop,cherry-studiolm-studio - 支持的级别:,
project,globalboth - 支持的研究分类:,
config,plugin,extension,mcp,skillall - 支持的使用周期:正整数天数,如,
7,14或30;未指定时默认days=14天14 - Harness别名:
- ,
claude->claude-codeclaude-code - ->
claude-desktopclaude-desktop - ,
chatgpt,chatgpt-desktop->openai-chatgptchatgpt - ->
codexcodex - ,
cursor,cursor-editor,cursor-desktop,cursor-web,cursor-cloud,cursor-cloud-agent->cursor-background-agentcursor - ,
cursor-cli,cursor-agent,cursor-agent-cli,agent-cli->agentcursor-cli - ,
grok,grok-build->grok-cligrok-build - ,
opencode->open-codeopencode - ,
perplexity,perplexity-desktop->perplexity-macperplexity-desktop - ,
cherry,cherrystudio,cherry-ai->cherry-studiocherry-studio - ,
lm-studio,lmstudio->lmslm-studio
- 级别别名:
- ,
project,repo->localproject - ,
global->userglobal - ,
both->all-levelsboth
- 多个Harness用逗号和空格分隔。
- 如果用户同时提供和指定Harness,询问用户选择哪种形式。
all - 的确定顺序:
all,claude-code,claude-desktop,chatgpt,codex,cursor,cursor-cli,grok-build,opencode,perplexity-desktop,cherry-studiolm-studio - 不猜测未知令牌。询问一个明确的澄清问题。
Grok Build (grok-build
)
grok-buildGrok Build (grok-build
)
grok-build- Instruction entry: ->
AGENTS.md->instructions/grok-global.mdinstructions/global.md - Policy template: merged into
config/grok-config.tomlvia repo Grok platform adapter sync~/.grok/config.toml - Ownership: replace-owned (,
models,mcp_servers,plugins,compat) vs blend-owned (telemetry,ui,features,session,tools,toolset,subagents)memory - Skills: no native Skills CLI adapter; install/sync uses alias + mirror to
claude-code; inventory also scans repo~/.grok/skills(project scope).grok/skills - Plannotator: CLI + skills + hooks (no Grok npm plugin). Install: (repo CLI). Policy:
grok plannotator install->config/grok-plannotator-hooks.json; home shim maps~/.grok/hooks/plannotator.json->block. Curated indeny. Contrast OpenCodedocs/src/authoring/skills/*.mdx.@plannotator/opencode@latest - Env (not in TOML): ,
GROK_WEB_FETCH,GROK_MEMORY,GROK_SUBAGENTS— document inGROK_LSP_TOOLS; verify withconfig/grok-env.sh/grok-delegate preflight - Cross-harness delegate: — native headless
/grok-delegate/grok -p/ worktrees / leader for Codex/OpenCode task graphs (not config sync)-r - Isolated sync: repo stack sync with (skips other harness home merges; includes Plannotator hook refresh)
--platforms grok --targets home
- 指令入口: ->
AGENTS.md->instructions/grok-global.mdinstructions/global.md - 策略模板: 通过仓库Grok平台适配器同步合并到
config/grok-config.toml~/.grok/config.toml - 所有权: replace-owned(,
models,mcp_servers,plugins,compat)vs blend-owned(telemetry,ui,features,session,tools,toolset,subagents)memory - Skills: 无原生Skills CLI适配器;安装/同步使用别名 + 镜像至
claude-code;清单同时扫描仓库~/.grok/skills(项目范围).grok/skills - Plannotator: CLI + skills + hooks(无Grok npm插件)。安装:(仓库CLI)。策略:
grok plannotator install->config/grok-plannotator-hooks.json;本地垫片映射~/.grok/hooks/plannotator.json->block。在deny中管理。对比OpenCode的docs/src/authoring/skills/*.mdx。@plannotator/opencode@latest - 环境变量(不在TOML中): ,
GROK_WEB_FETCH,GROK_MEMORY,GROK_SUBAGENTS— 记录在GROK_LSP_TOOLS中;通过config/grok-env.sh验证/grok-delegate preflight - 跨Harness代理: — 原生无界面
/grok-delegate/grok -p/ 工作区 / Codex/OpenCode任务图的领导者(非配置同步)-r - 隔离同步: 使用进行仓库栈同步(跳过其他Harness本地合并;包含Plannotator钩子刷新)
--platforms grok --targets home
Cursor Agent CLI (cursor-cli
)
cursor-cliCursor Agent CLI (cursor-cli
)
cursor-cli- Binary: (also
agent). Skills CLI adapter remainscursor-agent(cursor).-a cursor - Instruction entry: +
AGENTS.md+CLAUDE.md+.cursor/rules(thininstructions/cursor-cli-global.mdpointer).cursor/rules/cursor-cli.mdc - Shared SSOT with editor: ,
.cursor/hooks.json,.cursor/mcp.json,.cursor/skills,.cursor/agents— rendered once by the Cursor adapter (~/.cursor/{hooks,mcp,skills,agents})harness="cursor" - CLI-only: project permissions; user-owned
.cursor/cli.json(sync SHALL NOT write it). Doctor: repo CLI~/.cursor/cli-config.jsoncursor-cli doctor - Plugins: and
--plugin-dir(operator-opt-in; never sync-add)agent plugin marketplace - ACP / worktrees / worker: cite ;
cursor-acp;agent --worktreeis out of repo syncagent worker - Hooks: consume shared hooks.json + plugin hooks. Do not add to hook-registry render lists
cursor-cli
- 二进制文件: (也叫
agent)。Skills CLI适配器仍为cursor-agent(cursor)。-a cursor - 指令入口: +
AGENTS.md+CLAUDE.md+.cursor/rules(存在instructions/cursor-cli-global.md指针).cursor/rules/cursor-cli.mdc - 与编辑器共享单一可信源: ,
.cursor/hooks.json,.cursor/mcp.json,.cursor/skills,.cursor/agents— 由Cursor适配器(~/.cursor/{hooks,mcp,skills,agents})统一渲染harness="cursor" - 仅CLI: 项目权限;用户所有的
.cursor/cli.json(同步不得修改)。诊断:仓库CLI~/.cursor/cli-config.jsoncursor-cli doctor - 插件: 和
--plugin-dir(操作员选择加入;从不通过同步添加)agent plugin marketplace - ACP / 工作区 / 工作进程: 参考;
cursor-acp;agent --worktree不在仓库同步范围内agent worker - 钩子: 消费共享hooks.json + 插件钩子。不要将添加到钩子注册表渲染列表
cursor-cli
Classification Gate
分类校验
- Parse into harnesses, level,
$ARGUMENTS, install intent, apply intent, and unresolved tokens.all - If unresolved tokens remain, ask a single clarification question before continuing.
- If the request is , gap/missing/expand intent, or find-best-for-harness NL, run Discover (infer depth via
discover ...or discovery-pipeline.md).classify_intent.py - If the request is , run Usage Review only.
usage ... - If the request is , run Install Guidance only.
install ... - If ad-hoc "find skill for X", redirect to or
skill-router search. If systematic skill/MCP expansion, stay in Discover. If create skill, redirect to skill-creator. If open-ended non-harness research, redirect tonpx skills find./research - If the request is or an approval phrase, run Apply Approved only if the matching dry-run review already exists in this session. Otherwise rerun audit first.
apply ... - If the request includes unsupported harnesses, refuse cleanly and list the supported set.
- If harnesses are missing and was not supplied (audit/usage), ask only for harnesses or
all.all - If the level is missing (audit/usage), ask only for ,
project, orglobal.both - Otherwise run Audit in dry-run mode.
- 将解析为Harness、级别、
$ARGUMENTS、安装意图、应用意图和未解析令牌。all - 如果存在未解析令牌,先询问一个明确的澄清问题再继续。
- 如果请求是、缺口/缺失/扩展意图,或“为Harness寻找最佳X”类自然语言,运行发现模式(通过
discover ...或discovery-pipeline.md推断深度)。classify_intent.py - 如果请求是,仅运行使用情况评审。
usage ... - 如果请求是,仅运行安装指导。
install ... - 如果是临时“为X寻找skill”,重定向至或
skill-router search。如果是系统性skill/MCP扩展,留在发现模式。如果是创建skill,重定向至skill-creator。如果是开放式非Harness研究,重定向至npx skills find。/research - 如果请求是或批准短语,仅当本次会话中已存在匹配的试运行评审时,运行应用已获批方案。否则先重新运行审计。
apply ... - 如果请求包含不支持的Harness,明确拒绝并列出支持的集合。
- 如果缺少Harness且未提供(审计/使用情况),仅询问Harness或
all。all - 如果缺少级别(审计/使用情况),仅询问、
project或global。both - 否则以试运行模式运行审计。
Classification Logic
分类逻辑
- Route explicit mode tokens first: ,
discover,usage,install,apply,resume.list - Route gap/missing/expand/find-best NL to Discover depths (not Audit).
- Route complete harness/level scopes to Audit when config tune/fix intent and no discover signal.
- Route usage/cost/token NL to Usage Review.
- Route approvals to Apply Approved only when the immediately preceding dry-run audit matches the requested scope.
- 优先路由显式模式令牌:,
discover,usage,install,apply,resume。list - 将缺口/缺失/扩展/寻找最佳类自然语言路由至不同深度的发现模式(非审计)。
- 当存在配置调优/修复意图且无发现信号时,将完整的Harness/级别范围路由至审计模式。
- 将使用/成本/令牌类自然语言路由至使用情况评审。
- 仅当紧接的试运行审计匹配请求范围时,将批准请求路由至应用已获批方案。
Workflow
工作流程
Mode A: Audit (default, dry-run first)
模式A:审计(默认,首次为试运行)
-
Gate 0 — Discover surfaces
- Run:
bash
uv run python scripts/discover_surfaces.py --repo-root . --level <level> --harness <canonical-harness> [--harness ...] - Parse the JSON output and classify each surface as ,
present,missing,generated,merged, orrepo-observed.blind-spot - If the script cannot run, fall back to manual discovery with ,
Glob, andRead.Grep
- Run:
-
Gate 1 — Inspect project context before recommending fixes
- Read the repo's intent and operating model first: ,
README.md, harness-facing project files, key manifests, CI/workflow signals, and repo-native orchestration logic when present.AGENTS.md - In this repository, treat the repo-level harness sync script together with and
config/tooling-policy.jsonas canonical context for managed/generated/merged harness surfaces.config/sync-manifest.json
- Read the repo's intent and operating model first:
-
Gate 2 — Refresh latest official guidance
- Read .
references/latest-doc-sources.md - Use first when available, then first-party docs, then canonical vendor repo docs, then web fallback only if needed.
llms.txt - If official docs are unavailable, enter degraded mode and lower confidence.
- Read
-
Gate 3 — Audit each selected harness independently
- Load ,
references/harness-surfaces.md, andreferences/harness-checklists.mdselectively for the selected harnesses only.references/evidence-boundaries.md - For , compare project/global precedence before recommending edits.
both - Mark generated surfaces explicitly. Prefer fixing the canonical source over directly editing generated output.
- Treat missing native surfaces as evidence to assess, not automatic failure.
- Load
-
Gate 4 — Report dry-run findings
- Use the structure in .
references/output-format.md - Every finding must include severity, affected surface, evidence tag, why it matters for this project, and the correct home for the fix.
- Show concrete patch previews when the correct change is clear.
- Stop after the dry-run report and wait for explicit user approval.
- Use the structure in
-
校验0 — 发现配置面
- 运行:
bash
uv run python scripts/discover_surfaces.py --repo-root . --level <level> --harness <canonical-harness> [--harness ...] - 解析JSON输出并将每个配置面分类为、
present、missing、generated、merged或repo-observed。blind-spot - 如果脚本无法运行,回退至使用、
Glob和Read手动发现。Grep
- 运行:
-
校验1 — 推荐修复前检查项目上下文
- 先读取仓库的意图和运营模型:,
README.md、面向Harness的项目文件、关键清单、CI/工作流信号,以及仓库原生编排逻辑(如果存在)。AGENTS.md - 在本仓库中,将仓库级Harness同步脚本与和
config/tooling-policy.json一起视为托管/生成/合并Harness配置面的标准上下文。config/sync-manifest.json
- 先读取仓库的意图和运营模型:
-
校验2 — 刷新最新官方指南
- 读取。
references/latest-doc-sources.md - 优先使用(如果可用),然后是官方文档,再是标准供应商仓库文档,仅在必要时使用网络回退。
llms.txt - 如果官方文档不可用,进入降级模式并降低置信度。
- 读取
-
校验3 — 独立审计每个选定的Harness
- 针对选定的Harness,选择性加载、
references/harness-surfaces.md和references/harness-checklists.md。references/evidence-boundaries.md - 对于级别,在推荐修改前比较项目/全局优先级。
both - 显式标记生成的配置面。优先修改标准源而非直接编辑生成的输出。
- 将缺失的原生配置面视为评估证据,而非自动判定为问题。
- 针对选定的Harness,选择性加载
-
校验4 — 报告试运行结果
- 使用中的结构。
references/output-format.md - 每个发现必须包含严重性、受影响的配置面、证据标签、对本项目的影响,以及修复的正确位置。
- 当明确知道正确修改内容时,展示具体的补丁预览。
- 试运行报告后停止,等待用户明确批准。
- 使用
Mode B: Apply Approved
模式B:应用已获批方案
- Apply only after an explicit approval gate.
- Approval is valid only when the immediately preceding dry-run review in the same session already covered the same harness set and level.
harness-master - If the reviewed scope is stale, missing, ambiguous, or contradicted by new evidence, rerun Audit before editing.
- Restate the approved batch before making changes.
- Apply the smallest approved batch only. Do not broaden scope silently.
- If the approved remediation is an install step, use the exact guidance from
npx skills add ....references/install-guidance.md - After edits, rerun surface discovery and targeted dry-run checks for the touched harnesses, then summarize remaining gaps.
- 仅在明确的批准校验后执行应用。
- 仅当本次会话中紧接的试运行评审已覆盖相同的Harness集合和级别时,批准才有效。
harness-master - 如果评审范围已过期、缺失、模糊或与新证据矛盾,先重新运行审计再编辑。
- 执行修改前重申已获批的批次。
- 仅应用最小的获批批次。不得静默扩大范围。
- 如果获批的修复是安装步骤,使用中确切的
references/install-guidance.md指导。npx skills add ... - 编辑后,重新运行配置面发现并对受影响的Harness执行针对性试运行检查,然后总结剩余缺口。
Mode C: Discover
模式C:发现
Read-only gap expansion and harness-bounded research. Load and .
references/discovery-pipeline.mdreferences/discovery/coordinator-contract.md- Classify depth (full, focused, candidate, compare, w0only, ideate, journal).
- Run W0 scripts before any scouts when depth requires gaps.
- Use and
source_probe.pyinside W2/W2b — not as separate user-facing modes.candidate_score.py - Block apply/install until user confirms; config fixes require Audit + .
apply approved
只读的缺口扩展和Harness限定范围研究。加载和。
references/discovery-pipeline.mdreferences/discovery/coordinator-contract.md- 分类深度(完整、聚焦、候选、对比、仅W0、构思、日志)。
- 当深度要求分析缺口时,先运行W0脚本再执行任何侦察操作。
- 在W2/W2b中使用和
source_probe.py— 不作为独立的用户可见模式。candidate_score.py - 阻止应用/安装操作直到用户确认;配置修复需要先执行审计 + 。
apply approved
Mode D: Usage Review
模式D:使用情况评审
Use this mode for or natural-language requests to review token cost, quota, tool friction, MCP use, skill fit, plugin fit, context health, session attribution, or harness telemetry posture.
usage- Treat the entire mode as read-only. Do not edit files, install packages, create jobs, export raw sessions, print raw prompts, print raw traces, or mutate harness configs.
- Normalize harnesses, level, and usage window. Default missing window to days.
14 - Run surface discovery plus the safe usage probe:
bash
uv run python scripts/usage_probe.py --repo-root . --harness <canonical-harness> --level <level> --days <days> --json - Load ,
references/usage-review.md, andreferences/evidence-boundaries.md.references/output-format.md - Use runtime tools such as ,
token_stats,token_history,token_export,insights_collect,agent_attribution,quota_status,gemini_quota, andworkspace-summaryonly when they match the requested harness and privacy class.git-smart-status - Normalize findings into usage signals and score source coverage, privacy safety, cost/token posture, tool efficiency, MCP hygiene, skill fit, plugin fit, context health, approval safety, and recommendation actionability.
- Report recommendation lanes: ,
keep,tune-config,tune-skill,tune-mcp,tune-plugin,tune-workflow,instrument, ordefer.do-not-change - If the user asks to apply a Usage Review recommendation, block the apply step and require a matching dry-run Audit plus explicit approval first.
此模式适用于或自然语言请求,用于评审令牌成本、配额、工具摩擦、MCP使用、skill适配、插件适配、上下文健康度、会话归因或Harness遥测状态。
usage- 整个模式为只读。不得编辑文件、安装包、创建任务、导出原始会话、打印原始提示、打印原始跟踪或修改Harness配置。
- 标准化Harness、级别和使用周期。未指定周期时默认天。
14 - 运行配置面发现以及安全使用探针:
bash
uv run python scripts/usage_probe.py --repo-root . --harness <canonical-harness> --level <level> --days <days> --json - 加载、
references/usage-review.md和references/evidence-boundaries.md。references/output-format.md - 仅当运行时工具匹配请求的Harness和隐私类别时,才使用、
token_stats、token_history、token_export、insights_collect、agent_attribution、quota_status、gemini_quota和workspace-summary。git-smart-status - 将发现标准化为使用信号,并对源覆盖范围、隐私安全性、成本/令牌状态、工具效率、MCP卫生度、skill适配、插件适配、上下文健康度、批准安全性和建议可操作性进行评分。
- 报告建议方向:、
keep、tune-config、tune-skill、tune-mcp、tune-plugin、tune-workflow、instrument或defer。do-not-change - 如果用户要求应用使用情况评审的建议,阻止应用步骤并要求先执行匹配的试运行审计 + 明确批准。
Scaling Strategy
扩展策略
| Scope | Strategy |
|---|---|
| small | 1 harness and 1 level: inline review with one per-harness report |
| medium | 2-3 harnesses or one |
| large | 4+ harnesses or |
| Discover (full/focused) | Use Pattern F semantics: explore sources, score candidates, synthesize evidence, and stop before apply |
| Team-capable Discover | Use Pattern E or nested waves for source-family scouts, with accounting for every dispatched scout |
| Usage review | Prefer aggregate probes and runtime summaries first; avoid raw transcripts, raw traces, secret files, and message-table reads |
| 范围 | 策略 |
|---|---|
| 小型 | 1个Harness和1个级别:内联评审,每个Harness一份报告 |
| 中型 | 2-3个Harness或一个 |
| 大型 | 4个以上Harness或 |
| 发现(完整/聚焦) | 使用Pattern F语义:探索源、评分候选、综合证据,在应用前停止 |
| 团队级发现 | 使用Pattern E或嵌套波进行源类别侦察,并记录每个已调度的侦察任务 |
| 使用情况评审 | 优先使用聚合探针和运行时摘要;避免原始转录本、原始跟踪、机密文件和消息表读取 |
Latest-Doc Lookup Policy
最新文档查找策略
Use the most current authoritative guidance available:
- or equivalent official index
llms.txt - First-party docs pages for config, instructions, rules, MCP, and permissions
- Canonical vendor repo docs if first-party product docs are incomplete
- Web fallback only when the above are unavailable
Never claim without evidence from a current source.
latest使用可用的最新权威指南:
- 或等效官方索引
llms.txt - 配置、指令、规则、MCP和权限的官方文档页面
- 如果官方产品文档不完整,使用标准供应商仓库文档
- 仅当上述来源不可用时才使用网络回退
若无当前来源的证据,不得声称是“最新”。
Per-Harness Review Contract
每个Harness的评审约定
- Review each harness independently before writing any synthesis.
- Tag every finding as ,
verified-file,verified-doc, orrepo-observed.blind-spot - Surface the authoritative, secondary, generated, and merged config surfaces for the selected level.
- For , report project/global conflicts before recommending changes.
both - If a harness-specific native surface is not observable from the current session, mark it as a blind spot instead of inventing behavior.
- 在撰写任何综合分析前,独立评审每个Harness。
- 将每个发现标记为、
verified-file、verified-doc或repo-observed。blind-spot - 展示选定级别的权威、次要、生成和合并配置面。
- 对于级别,在推荐修改前报告项目/全局冲突。
both - 如果当前会话无法观察到Harness特定的原生配置面,标记为盲点而非臆测行为。
Install Guidance Contract
安装指导约定
- Load only when the user asks how to install
references/install-guidance.md, or when missing skill availability is the actual root cause.harness-master - Default surfaced form:
bash
npx skills add <source> --skill harness-master -y -g --agent claude-code --agent codex --agent crush --agent cursor --agent opencode - Do not recommend install commands for native config problems.
- Do not suggest project-local install by default. Mention only if the user explicitly asks for project-local installation.
python scripts/install_skills.py --local --execute
- 仅当用户询问如何安装,或缺失skill可用性是实际根本原因时,加载
harness-master。references/install-guidance.md - 默认展示格式:
bash
npx skills add <source> --skill harness-master -y -g --agent claude-code --agent codex --agent crush --agent cursor --agent opencode - 不为原生配置问题推荐安装命令。
- 默认不建议项目本地安装。仅当用户明确要求项目本地安装时,提及。
python scripts/install_skills.py --local --execute
Validation Contract
验证约定
After changing this skill, run:
bash
python scripts/check.py
uv run python scripts/usage_probe.py --repo-root . --harness opencode --level both --days 14 --json
uv run python scripts/discover_surfaces.py --repo-root . --level both --harness claude-code --harness claude-desktop --harness chatgpt --harness codex --harness cursor --harness cursor-cli --harness grok-build --harness opencode --harness perplexity-desktop --harness cherry-studio --harness lm-studioCompletion criteria:
- metadata and eval manifests validate
- packaging remains portable
- surface discovery smoke-check accepts every canonical harness
- dry-run/apply behavior still requires approval before edits
- Discover remains read-only and blocks apply until a matching audit is approved
- usage review remains read-only, redacts sensitive sources, and blocks apply until a matching audit is approved
修改本技能后,运行:
bash
python scripts/check.py
uv run python scripts/usage_probe.py --repo-root . --harness opencode --level both --days 14 --json
uv run python scripts/discover_surfaces.py --repo-root . --level both --harness claude-code --harness claude-desktop --harness chatgpt --harness codex --harness cursor --harness cursor-cli --harness grok-build --harness opencode --harness perplexity-desktop --harness cherry-studio --harness lm-studio完成标准:
- 元数据和评估清单验证通过
- 打包保持可移植性
- 配置面发现冒烟测试接受所有标准Harness
- 试运行/应用行为仍需批准后才能编辑
- 发现模式保持只读,且需匹配审计获批准后才能执行应用
- 使用情况评审保持只读,编辑敏感源,且需匹配审计获批准后才能执行应用
Cross-Harness Synthesis Contract
跨Harness综合分析约定
After all per-harness reviews, synthesize:
- Shared issues across harnesses
- Conflicting conventions or precedence rules
- Duplicated instruction sources or MCP definitions
- Generated-vs-canonical drift patterns
- Highest-leverage cleanup order if the user wants a follow-up apply batch
完成所有单个Harness评审后,综合分析:
- 跨Harness的共享问题
- 冲突的约定或优先级规则
- 重复的指令源或MCP定义
- 生成源与标准源的漂移模式
- 如果用户需要后续应用批次,给出最高优先级的清理顺序
Output Contract
输出约定
Every per-harness report must include:
- Harness
- Level
- Files Reviewed
- Docs Checked
- Project Context Summary
- Blind Spots
- Scorecard
- Findings
- Recommended Changes
- Proposed Patch Preview
- Confidence
Then add a cross-harness synthesis section when 2+ harnesses were reviewed.
每个Harness的报告必须包含:
- Harness
- 级别
- 已评审文件
- 已检查文档
- 项目上下文摘要
- 盲点
- 评分卡
- 发现结果
- 推荐修改
- 补丁预览建议
- 置信度
当评审2个以上Harness时,添加跨Harness综合分析部分。
Reference File Index
参考文件索引
| File | Content | Read When |
|---|---|---|
| Gate-by-gate audit/apply workflow, precedence rules, degraded mode, and approval gate details | Audit, Apply Approved |
| Official | Latest-doc refresh |
| Project/global surfaces, precedence, install agent names, generated/merged notes | Surface interpretation |
| Per-harness audit checklist and edge cases | Per-harness review |
| Evidence tags, blind-spot handling, and contradiction policy | Reporting findings |
| Exact | Install Guidance |
| Per-harness and cross-harness report templates | Final output |
| W0–W4 discover waves, depth inference, evidence vs contracts, redirects | Discover |
| Coordinator contract, scout templates, output formats, research integration | Discover scouts and reports |
| Legacy source-family notes; prefer discovery-pipeline.md for dispatch | Discover W2 source planning |
| Source families, programmatic access, confidence roles, and degraded-source behavior | Discover W2 source planning |
| Read-only usage sources, privacy classes, scorecard, signal schema, and recommendation lanes | Usage Review |
| Machine-readable source registry for source planning and probe scripts | Discover W2 source planning |
Read only the references needed for the active step. Do not preload all references.
| 文件 | 内容 | 读取时机 |
|---|---|---|
| 按校验步骤划分的审计/应用工作流程、优先级规则、降级模式和批准校验细节 | 审计、应用已获批方案 |
| 每个Harness的官方 | 最新文档刷新 |
| 项目/全局配置面、优先级、安装Agent名称、生成/合并说明 | 配置面解释 |
| 每个Harness的审计清单和边缘情况 | 单个Harness评审 |
| 证据标签、盲点处理和矛盾处理策略 | 发现结果报告 |
| 确切的 | 安装指导 |
| 单个Harness和跨Harness的报告模板 | 最终输出 |
| W0–W4发现阶段、深度推断、证据与约定、重定向 | 发现模式 |
| 协调器约定、侦察模板、输出格式、研究集成 | 发现侦察和报告 |
| 遗留源类别说明;优先使用discovery-pipeline.md进行调度 | 发现W2源规划 |
| 源类别、程序化访问、置信度角色和降级源行为 | 发现W2源规划 |
| 只读使用源、隐私类别、评分卡、信号 schema 和建议方向 | 使用情况评审 |
| 用于源规划和探针脚本的机器可读源注册表 | 发现W2源规划 |
仅读取当前步骤所需的参考文件。不要预加载所有参考文件。
Progressive Disclosure
渐进式披露
Load reference files as indicated by the active mode. Do not load all references at once; use the dispatch table, source plan, and final output needs to choose the smallest relevant set.
根据当前活动模式加载参考文件。不要一次性加载所有参考文件;使用调度表、源计划和最终输出需求选择最小的相关集合。
Canonical Vocabulary
标准词汇
These terms are canonical for reports and should be used exactly.
harness-masterUse these terms exactly throughout:
| Term | Meaning | NOT |
|---|---|---|
| One supported agent/runtime target | |
| | |
| Repo-local file or generated artifact used by a harness | |
| User-level harness config outside the repo | |
| Findings + patch preview only; no edits | |
| Explicit user consent required before edits | |
| A surface or behavior that is not observable in the current session | |
| Behavior inferred from the current codebase's real harness wiring | |
| Proposed diff or snippet shown before edits | |
| The file or config that should be changed instead of a generated output | |
| Read-only list of sources, URLs/commands, env vars, evidence fields, and failure modes | |
| Normalized evidence bundle for one ecosystem candidate | |
| Adoption readiness class backed by evidence and validation state | |
| Candidate class for credential, proxy, destructive, offensive, or broad-permission risk | |
| Sanitized aggregate evidence about token, cost, quota, tool, MCP, skill, plugin, or workflow behavior | |
| Sensitivity class that determines whether a source can be collected, summarized, or must stay manual-only | |
这些术语是报告的标准词汇,应严格使用。
harness-master严格使用以下术语:
| 术语 | 含义 | 禁用术语 |
|---|---|---|
| 一个受支持的Agent/运行时目标 | |
| 评审范围: | |
| Harness使用的仓库本地文件或生成工件 | |
| 仓库外的用户级Harness配置 | |
| 仅展示发现结果+补丁预览;不修改 | |
| 修改前需要用户明确同意 | |
| 当前会话中无法观察到的配置面或行为 | |
| 从当前代码库的实际Harness连接推断出的行为 | |
| 修改前展示的建议差异或片段 | |
| 应修改的文件或配置,而非生成输出 | |
| 只读的源列表、URL/命令、环境变量、证据字段和失败模式 | |
| 一个生态系统候选者的标准化证据包 | |
| 基于证据和验证状态的采用就绪类 | |
| 存在凭证、代理、破坏性、冒犯性或广泛权限风险的候选类别 | |
| 关于令牌、成本、配额、工具、MCP、skill、插件或工作流行为的 sanitized 聚合证据 | |
| 决定源是否可收集、汇总或必须保持手动的敏感度类别 | |
Critical Rules
关键规则
- Every new run starts in dry-run audit mode.
harness-master - Never edit before explicit user approval.
- is valid only after a matching dry-run review in the current session.
Apply Approved - Ask only for missing inputs. Never re-ask resolved harnesses or levels.
- Always inspect project context before recommending harness changes.
- Always refresh latest official guidance before claiming a configuration is current or stale.
- Mark blind spots explicitly. Do not guess hidden or UI-only settings.
- Prefer edits to canonical sources over generated or merged outputs when this is clear from evidence.
- Every finding must carry an evidence tag: ,
verified-file,verified-doc, orrepo-observed.blind-spot - Show concrete patch previews whenever the correct edit is clear.
- Use only when installation is genuinely the right next step.
npx skills add ... - Unsupported harnesses must be refused cleanly with the supported set listed.
- Ecosystem research is advisory and read-only. Never apply from a research report alone.
- Credentialed APIs are optional enrichments. Missing credentials lower confidence but do not block baseline source planning.
- GitHub GraphQL enriches known candidate repositories; use REST/search or code/file lookup for broad discovery.
- Usage Review is advisory and read-only. Never apply config changes, install tools, export raw sessions, or print raw prompts/traces from usage findings alone.
- Usage Review may use aggregate-safe and sanitized metadata sources, but secret files, credential values, raw transcripts, raw traces, and raw message-table bodies are out of scope.
- 每次新的运行都从试运行审计模式开始。
harness-master - 获得用户明确批准前不得修改。
- 仅当本次会话中存在匹配的试运行评审时,才有效。
Apply Approved - 仅询问缺失的输入。不得重复询问已解析的Harness或级别。
- 推荐Harness修改前必须检查项目上下文。
- 声称配置是最新或过时前必须刷新最新官方指南。
- 显式标记盲点。不得猜测隐藏或仅UI可见的设置。
- 当证据明确时,优先修改标准源而非生成或合并输出。
- 每个发现必须带有证据标签:、
verified-file、verified-doc或repo-observed。blind-spot - 当明确知道正确修改内容时,展示具体的补丁预览。
- 仅当安装确实是正确的下一步时,才使用。
npx skills add ... - 必须明确拒绝不支持的Harness并列出支持的集合。
- 生态系统研究仅为建议且只读。不得仅基于研究报告执行应用。
- 带凭证的API是可选的增强功能。缺少凭证会降低置信度但不阻止基线源规划。
- GitHub GraphQL可丰富已知候选仓库;广泛发现使用REST/search或代码/文件查找。
- 使用情况评审仅为建议且只读。不得仅基于使用发现结果执行配置修改、安装工具、导出原始会话或打印原始提示/跟踪。
- 使用情况评审可使用聚合安全和 sanitized 的元数据源,但机密文件、凭证值、原始转录本、原始跟踪和原始消息表内容不在范围内。