Loading...
Loading...
Found 2,091 Skills
AI-powered deep stock analysis engine for A-share/HK/US markets with 51 investor personas, 22 data dimensions, 180 quantitative rules, and 17 institutional methods
Score and compare images using vision LLMs as judges. YAML-defined criteria presets for 11 use cases (text-to-image, photorealism, document OCR, charts, UI, portrait, product, scientific, invoice, alt-text, artistic style). Supports OpenAI, Anthropic, Gemini, Mistral, and OpenRouter as judge providers. Keys auto-decrypted via SOPS + age.
Master context engineering principles for building production-grade AI agent systems with effective context management, multi-agent architectures, and memory systems.
Iteratively inspect an agent repository and optional traces, interview the user, and create, run, and audit Harbor evals one at a time. Use for agent evals, benchmark tasks, regression cases, trace-informed evals, verifier design, or controlled agent environments.
Own assessment-only manuscript deliverables without rewriting: full scientific review, scoring, reviewer reports, issue diagnosis, AC/meta-review, readiness judgment, writing/format review, and cross-version comparison. Use for full review, scientific review, do not rewrite, assessment-only, paper review, score drift, moving-target review, 完整审稿, 不要改写, 模拟审稿, 论文评分, 版本对比, 复审一致性, 写作评审, LaTeX检查. Requests for revised or polished prose, including rewrite based on reviews with no new review, belong to ccf-paper-writer; rebuttals belong to ccf-rebuttal-writer; visual/table styling belongs to ccf-visual-composer.
Analyze lending products including mortgages, HELOCs, and personal loans with amortization and comparison tools. Use when the user asks about mortgage comparison, fixed vs ARM rates, loan qualification, amortization schedules, extra payments, or buying points. Also trigger when users mention 'monthly payment calculation', '15-year vs 30-year mortgage', 'PMI', 'APR vs interest rate', 'HELOC', 'home equity', 'should I buy down the rate', 'biweekly payments', or ask how much house they can afford.
Web search and research specialist for finding and synthesizing information
Reviews a change by running the mission, architecture, implementation, craft, security, and performance passes, then weighing them into a verdict.
Convert W&B Table artifacts into non-destructive EvalTable previews with scan-first planning, typed input/output/score columns, bounded batches, verification, and safe removal. Use when a coding agent needs to create, inspect, compare, verify, or remove W&B EvalTable previews.
Use when auditing an existing product, app, or feature across all four Product Judgement scales: screen structure (Focal), multi-screen journeys (Compass), relationship value and retention (Flywheel), and memorable moments (Soul). Run for a holistic app audit, cross-scale critique, or prioritized UX review using a codebase, live product, prototype, Figma/Paper frames, screenshots, or a description. Prefer a codebase because it exposes behavior, state, and lifecycle context. Do not use for a single-screen, single-flow, or single-stage review; invoke the corresponding Skill instead, or for implementation, design-system analysis, visual styling, animation implementation, research, or analytics.
Publish the nurb model leaderboard from merged benchmark submissions. Sanity-checks every run landed since the last regeneration, writes or refreshes the editorial verdicts, regenerates evals/REPORT.md and site/benchmarks.html, and opens the publish PR. Use when the user says "update the leaderboard", "publish the benchmarks", "regenerate the benchmark page", or after merging submission PRs.
Gate fine-tuned checkpoints with drift budgets, paired comparison, and forgetting checks before promotion. Use after a training run produces a checkpoint, when deciding whether a tuned model ships, or when a promoted model needs re-gating against updated goldens.