Loading...
Loading...
Autonomously optimize an existing AI skill by running it repeatedly against binary evals, mutating one instruction at a time, and keeping only changes that improve pass rate. Based on Karpathy-style autoresearch, but applied to SKILL.md iteration instead of ML training. Use when optimizing a skill, benchmarking prompt quality, building evals for a skill, or running self-improvement loops on reusable agent instructions. Triggers on: skill-autoresearch, optimize this skill, improve this skill, benchmark this skill, eval my skill, run autoresearch on this skill, self-improve skill.
npx skill4agent add akillness/oh-my-skills skill-autoresearch52mSKILL.mdreferences/EVAL 1: Short name
Question: Yes/no question about the output
Pass: Specific condition that counts as yes
Fail: Specific condition that counts as noskill-autoresearch-[skill-name]/
dashboard.html
results.json
results.tsv
changelog.md
SKILL.md.baselineresults.tsvresults.jsondashboard.htmlSKILL.md.baselineSKILL.md.baselineN0results.tsvexperiment score max_score pass_rate status descriptionSKILL.mdresults.tsvresults.jsonchangelog.mdresults.json## Experiment N — keep|discard
Score: X/Y
Change: one-sentence mutation summary
Reasoning: why this mutation was tried
Result: what improved or regressed
Remaining failures: what still breaksskill-autoresearch-[skill-name]/
dashboard.html
results.json
results.tsv
changelog.md
SKILL.md.baseline