illo

Compare original and translation side by side

🇺🇸

Original

English
🇨🇳

Translation

Chinese

Illo

Illo

Make original, distinctive editorial illustrations for written content. One image explains one idea: a key judgment, a flow, a before/after, a trap, a loop. A recurring mascot is the one performing the idea in every scene — the subject, never decoration. When one idea advances through stages, it can be a mini-comic: 2–4 panels inside a single image. And when the idea is itself a traceable structure — a pipeline, labeled stages, a fan-out, a timeline, a loop — it can be an explainer: the same mascot and look drawing the structure as a hand-built sketch-diagram with arrows and callouts (
references/composition.md
, "Two registers" and "Pick the diagram type"; editorial scene is always the default). A named pipeline or recipe is labeled stages inside that register — named phases in order, one connected system, pack-solved for this body, never a new look. Or a character cutout: the mascot alone on a transparent PNG for downstream overlay — pose and contact continuity only, no idea, no text, no environment (
references/cutout.md
).
This is a configurable house style, not a generic image generator. The methodology is the constant; the character pack and palette are the parameters — and a character pack carries its style with it: one look per pack, chosen from the bundled look library (riso — grainy halftone, ink-layer offset, paper grain, one bold softly-rounded outline — plus blueprint, woodcut, pixel, clay, manila, chalk, phosphor, enamel, gouache, felt, diorama, sketchbook, bricks, fizz, bloom, and snes) or a custom style file. The default mascot is Blot, a deadpan ink-drop in riso. Palettes come from presets, the user's own palette file, or one derived color. Whatever the parameters, it is intentionally not a photo — with one deliberate exception, the
bricks
look, a toy-brick photography style — not a logo, not a corporate infographic, not a formal boxes-and-diamonds flowchart look, not a UI mockup. Asking for a flowchart still means labeled stages in the pack's look — the formality ban is a look constraint, not a refusal of the word.
为书面内容生成原创、独特的编辑插画。每张图阐释一个核心创意:关键判断、流程、前后对比、陷阱、循环。重复出现的吉祥物是每个场景中演绎创意的主体——绝非装饰元素。当创意需分阶段推进时,可生成迷你漫画:单张图片内含2-4个分镜。若创意本身是可追踪的结构——如流水线、标注阶段、扇形展开、时间线、循环——则可生成说明图:同一吉祥物以相同风格绘制结构,呈现为带箭头和标注的手绘草图(详见
references/composition.md
中的“两种呈现形式”和“选择图表类型”;编辑场景为默认选项)。命名的流水线或流程对应说明图中的标注阶段——按顺序排列的命名阶段,构成一个互联系统,适配当前角色风格,不会使用新风格。此外还支持角色抠图:仅保留吉祥物的透明PNG文件,用于后续叠加——仅保证姿势和连贯性,不含创意、文字或环境(详见
references/cutout.md
)。
这是一款可配置的专属风格工具,而非通用图像生成器。方法论固定不变角色包和调色板为可配置参数——每个角色包自带其风格:每个包对应一种风格,选自内置风格库(riso风格——颗粒状半色调、墨层偏移、纸张纹理、粗软圆角轮廓——外加蓝图、木刻、像素、黏土、马尼拉纸、粉笔、荧光、珐琅、水粉、毛毡、立体模型、速写本、积木、气泡、 bloom、snes风格)或自定义风格文件。默认吉祥物为Blot,一个面无表情的riso风格墨滴。调色板可来自预设、用户自定义调色板文件,或从单一颜色衍生而来。无论参数如何,生成的内容刻意避免写实风格——唯一例外是
bricks
风格,即积木摄影风格;也不会生成logo、企业信息图、正式的方框菱形流程图或UI原型。若请求流程图,仍会以当前角色包的风格生成标注阶段——禁止正式风格是外观约束,而非拒绝流程图需求。

Use cases — route the request

用例——请求路由

The user wantsThe path
Illustrate an article / post / newsletter / URLSteps 0–7: route the source first (thesis → coverage: hero / hero+set / set / mini-comic —
references/composition.md
, "Source routing"), then shot list (hero row + anchors), one image per anchor, interleave by placement.
One image for a single conceptStep 1 concept branch (up to ~3 quick questions if the idea is thin), then a single image.
Surprise / random — "surprise me", "random", "surprise me with art quote using bray", "surprise me --autopick"Read
references/surprise.md
in full: Step 0 first, then character + provenance (ignore
defaultCharacter
;
* quote
forces a cited quote; else ~1/3 roll), build three safe candidates, interactive picker or auto-pick-best (
--autopick
preferred for schedulers), then register from the locked saying, then Steps 3–7 as one image. Deliver saying + image. Poster titles default off; mini-comics still get per-panel labels.
A sequence — story beat, before→after, fail→fixOne mini-comic when the progression sits in one place (shape routing in
references/composition.md
— the idea picks the shape, the destination never does). A specified process diagram / flowchart / labeled workflow is labeled stages, not this row.
A traceable structure — "show the flow", "as labeled stages", "label the steps", "walk the stages", "diagram the pipeline", "like that factory diagram", "map the steps", "as an explainer", or specified flowchart / labeled-workflow / process-diagram intentionThe explainer register (
references/composition.md
, "Pick the diagram type" and "The explainer register"): a hand-built labeled-stages / flow / fan-out / timeline / loop / stack / system slice in the active look, the mascot a working part of it. Specified flowchart / labeled-workflow / process-diagram intention locks labeled stages in the pack's look — the formal-flowchart ban is a look constraint (no Visio, no title/legend/grid), not a refusal of the word. Labeled stages is a structure type inside explainer, not a new register or look — pack-solve it for the active character before the prompt. BEST when a unit's thesis IS a named pipeline, recipe, or staged process; never the automatic choice for every explainer.
Social-ready art for X posts / article body images16:9 (or 1:1 when square is explicitly useful), bold
ink-punch
, watermark with the
x
handle if configured or asked.
X Article banner / hero imageUse the unique banner format: 1536 × 640 px when the user asks for an X Article hero/banner. Prompt and render through the normal
illo.py generate
image pipeline, with normal, undistorted character/object proportions and crop-safe breathing room. Do not satisfy this by manually compositing or rebuilding crops from another image unless the user explicitly asks for post-processing.
Blog / brand / site-matched artA named or custom palette, or derive the palette from one dominant color (
references/palettes.md
).
Their own mascot — "make me a character", "use our mascot", "replace Blot"The character builder: read
references/character-builder.md
in full and follow it end to end.
Community characters — "what characters are available", "install blip", "install all characters", "update mole", "publish my character"
references/pack-sharing.md
— engine
packs list/show/install/update
, including
packs install --all
; publish via a GitHub PR.
A different look — "in blueprint", "woodcut style", "pixel version of blip"Styles travel with character packs: build a style variant pack via
references/character-builder.md
, "Style variants".
Options to pick from, or "which model is best"Step 5b:
--count
variations or a model loop →
gallery
with a recommendation.
Fix an existing image (stray title, recolor, mascot too decorative)Edit prompts in
references/prompt-recipe.md
, passing the image back as
--ref
.
Character cutout / transparent PNG / overlay sticker — "just the mascot", "no background", "paste on something else"The cutout register (
references/cutout.md
): read in full, prompt from
references/prompt-recipe.md
"Cutout variant", generate with
--cutout
and
--aspect 1:1
. OpenRouter cutouts default to GPT Image 2 (not Grok). Not for explaining an idea — reroute to editorial if the ask needs a scene.
Animated idle / bot avatar / looping GIF of the mascotThe cutout register plus
references/cutout.md
, "Idle loop / bot avatar": one transparent 1:1 cutout with
--cutout
and the character sheet as
--ref
, then programmatic motion on that PNG.
用户需求处理路径
为文章/帖子/通讯/URL生成插画步骤0-7:先对来源进行路由(核心论点→覆盖范围:主图/主图+系列图/系列图/迷你漫画——详见
references/composition.md
中的“来源路由”),再生成镜头列表(主图+锚点图),每个锚点对应一张图,按位置交错排列。
为单个概念生成一张图步骤1的概念分支(若创意模糊,可提最多3个简短问题),然后生成单张图片。
惊喜/随机模式——“给我惊喜”、“随机生成”、“用Bray生成艺术名言的惊喜插画”、“惊喜模式 --autopick”完整阅读
references/surprise.md
:先执行步骤0,再确定角色和创意来源(忽略
defaultCharacter
* quote
强制使用带引用的名言;否则约1/3概率随机生成),构建3个安全候选方案,通过交互式选择或自动最优选择(调度器优先使用
--autopick
)确定,再根据锁定的文案选择呈现形式,最后执行步骤3-7生成单张图片。交付文案+图片。默认关闭海报标题;迷你漫画仍保留分镜标注。
序列内容——故事节拍、前后对比、失败→修复当流程在同一场景内推进时,生成迷你漫画(详见
references/composition.md
中的形状路由——创意决定形状,而非目标载体)。指定的流程图/工作流标注图属于标注阶段,不属于此类序列。
可追踪的结构——“展示流程”、“标注阶段”、“标注步骤”、“分步演示”、“绘制流水线图”、“类似工厂示意图”、“映射步骤”、“作为说明图”,或明确要求流程图/标注工作流/流程示意图说明图模式(详见
references/composition.md
中的“选择图表类型”和“说明图模式”):以当前风格生成手工制作的标注阶段/流程/扇形展开/时间线/循环/堆叠/系统切片,吉祥物作为结构的一部分。明确要求流程图/标注工作流/流程示意图时,锁定说明图模式中的标注阶段,并使用当前角色包的风格——禁止正式流程图是外观约束(无Visio风格、无标题/图例/网格),而非拒绝需求。标注阶段是说明图模式内的一种结构类型,而非新的呈现形式或风格——在生成提示前需适配当前角色。最适合核心论点为命名流水线、流程或分阶段过程的场景;绝非所有说明图的默认选择。
适合X平台帖子/文章正文的社交友好型插画16:9比例(若明确需要方形则用1:1),采用醒目
ink-punch
调色板,若已配置或被要求则添加X平台账号水印。
X平台文章横幅/主图使用独特横幅格式:当用户请求X平台文章主图/横幅时,尺寸为1536 × 640 px。通过常规
illo.py generate
图像流水线生成提示并渲染,保持角色/物体比例正常且无变形,预留安全裁剪空间。除非用户明确要求后期处理,否则不要通过手动合成或从其他图片裁剪来满足需求。
匹配博客/品牌/网站风格的插画使用命名或自定义调色板,或从单一主色调衍生调色板(详见
references/palettes.md
)。
自定义吉祥物——“帮我创建一个角色”、“使用我们的吉祥物”、“替换Blot”使用角色构建器:完整阅读
references/character-builder.md
并按流程操作。
社区角色——“有哪些可用角色”、“安装Blip”、“安装所有角色”、“更新Mole”、“发布我的角色”详见
references/pack-sharing.md
——引擎支持
packs list/show/install/update
命令,包括
packs install --all
;通过GitHub PR发布角色包。
不同风格——“蓝图风格”、“木刻风格”、“Blip的像素版”风格随角色包绑定:通过
references/character-builder.md
中的“风格变体”构建风格变体包
选项选择,或“哪个模型最好”步骤5b:使用
--count
参数生成变体或循环测试模型→生成
gallery
并给出推荐。
修改现有图片(多余标题、重新上色、吉祥物过于装饰化)编辑
references/prompt-recipe.md
中的提示,将图片作为
--ref
参数传入。
角色抠图/透明PNG/叠加贴纸——“只要吉祥物”、“无背景”、“可粘贴到其他内容上”抠图模式(详见
references/cutout.md
):完整阅读该文档,使用
references/prompt-recipe.md
中的“抠图变体”提示,添加
--cutout
--aspect 1:1
参数生成。OpenRouter抠图默认使用GPT Image 2(而非Grok)。该模式不用于阐释创意——若需求包含场景,需重新路由到编辑场景模式。
吉祥物的动态 idle/机器人头像/循环GIF抠图模式结合
references/cutout.md
中的“Idle循环/机器人头像”:生成一张1:1比例的透明抠图,添加
--cutout
参数并将角色表作为
--ref
,然后对该PNG进行程序化动画处理。

Prerequisites

前置条件

The engine (
scripts/illo.py
, stdlib Python, no installs) renders through one of three engine backends;
python3
and network access are the only hard requirements. Grok Bot (Cursor's Grok Bot / the Grok desktop assistant) is a fourth, agent-side transport: use its built-in Grok image tool directly, not
illo.py generate
, when no user config explicitly selects an engine backend.
Running the engine — set
$SKILL_DIR
inline in each block.
Every engine command below is
python3 "$SKILL_DIR/scripts/illo.py" …
. Set
SKILL_DIR
to the absolute path of the directory this
SKILL.md
was loaded from (it contains
scripts/illo.py
and
assets/
) in the same command block that uses it — shell state does not persist between separate command runs, so a value set in an earlier block is gone by the next. If the harness does not expose that path, find the installed
scripts/illo.py
and use its parent; if neither resolves, stop rather than guessing the working directory. The engine self-locates its own bundled assets, so
$SKILL_DIR
only has to be right enough to launch
illo.py
and to point
--ref
at the bundled character sheet.
Write the block flatten-safe — some hosts (Codex observed) collapse a fenced block to one line, turning a newline into a space. Terminate the assignment with
;
(
SKILL_DIR="…";
— without it, a flattened
SKILL_DIR="…" python3 "$SKILL_DIR/…"
becomes an env-prefix whose
$SKILL_DIR
expands to empty before the assignment applies, so the path collapses to
/scripts/illo.py
). Put no comment on an assignment or command line (a flattened
#
comments out the rest of the line and the command silently vanishes), and keep each invocation on one line (a flattened
\
continuation injects stray arguments). A wrong or unset value makes
doctor
(Workflow step 0) fail loudly (
can't open file …/scripts/illo.py
) — the signal to fix the path, not a skill fault.
  • Codex backend (free for Codex subscribers). When the host has a usable Codex CLI — installed,
    codex login
    -ed, with the
    image_generation
    feature — illo can generate through the user's Codex subscription at no per-image charge (it draws on their Codex quota). No API key, no token: illo only shells out to the user's own CLI. Detected, not assumed; gpt-image-2 is automatic; unsupported on Windows/WSL.
  • Grok CLI backend (free for Grok/xAI subscribers). When the host has a usable Grok CLI — installed and
    grok login
    -ed — illo can generate through the user's Grok subscription via
    grok -p
    (headless), drawing on their Grok quota. Same env-free, token-free subprocess design as Codex. Grok returns JPEG with no alpha, so it cannot make transparent cutouts — those auto-fall back to a cutout-capable backend. The image tool exposes no model selector.
  • Grok Bot native transport (agent-side, free for Grok Bot users). When you are Grok Bot — specifically Cursor's Grok Bot / the Grok desktop assistant with the built-in Grok image tool — build the illo prompt and call that tool with the active character's model sheet as a reference image. Do not require the Grok CLI, Codex CLI, or an OpenRouter key; do not treat a missing engine backend as a reason to run
    init
    . This is not a generic "host image API" rule and not an
    illo.py --backend
    value.
  • OpenRouter backend (paid, direct or explicit fallback). Needs an OpenRouter API key in the user's config file — the single credential channel — written once by the user-run
    init
    (mode 600). The engine never reads secrets from the environment and never accepts them as command-line arguments. A host without a subscription CLI can select this engine path directly. A failed Codex/Grok CLI render does not spend money automatically: paid fallback requires
    --allow-paid-fallback
    . It is model-selectable (
    --model
    ).
Capsule of the backend/transport model (resolution and precedence, the CLI requirements, the Grok Bot native path, the built-in image tool being automatic, quota vs. charge, cutout limits, Windows/WSL, fallback): read
references/backends.md
in full before choosing or explaining a backend
— the mechanics live there, once.
引擎(
scripts/illo.py
,基于Python标准库,无需额外安装)通过三种引擎后端渲染;
python3
和网络访问是唯一硬性要求。Grok Bot(Cursor的Grok Bot/Grok桌面助手)是第四种代理端传输方式:当用户未明确选择引擎后端时,直接使用其内置的Grok图像工具,而非
illo.py generate
运行引擎——在每个命令块中内联设置
$SKILL_DIR
。以下所有引擎命令格式均为
python3 "$SKILL_DIR/scripts/illo.py" …
。将
SKILL_DIR
设置为加载此
SKILL.md
的目录的绝对路径(该目录包含
scripts/illo.py
assets/
且需在使用该变量的同一命令块中设置——shell状态不会在不同命令块间保留,因此前一个命令块设置的值会在下一个命令块中失效。若工具未暴露该路径,找到已安装的
scripts/illo.py
并使用其父目录;若均无法解析,应停止操作而非猜测工作目录。引擎会自动定位内置资源,因此
$SKILL_DIR
只需足够准确以启动
illo.py
并指向
--ref
参数对应的内置角色表。
编写命令块时需支持扁平化处理——部分宿主(如Codex)会将代码块压缩为一行,将换行符转为空格。赋值语句需以
;
结尾(
SKILL_DIR="…";
——若省略,扁平化后的
SKILL_DIR="…" python3 "$SKILL_DIR/…"
会变为环境前缀,其中
$SKILL_DIR
在赋值生效前会被展开为空,导致路径变为
/scripts/illo.py
)。不要在赋值或命令行添加注释——扁平化后的
#
会注释掉该行剩余内容,导致命令无声失效;每个调用需保持在一行(扁平化后的
\
换行符会注入多余参数)。错误或未设置的值会导致
doctor
(工作流步骤0)报错(
can't open file …/scripts/illo.py
)——这是修复路径的信号,而非技能故障。
  • Codex后端(Codex订阅用户免费)。当宿主有可用的Codex CLI——已安装、完成
    codex login
    、具备
    image_generation
    功能——illo可通过用户的Codex订阅生成图片,无需按次付费(消耗用户的Codex配额)。无需API密钥或令牌:illo仅调用用户本地的CLI。引擎会自动检测,而非默认启用;默认使用gpt-image-2;不支持Windows/WSL系统。
  • Grok CLI后端(Grok/xAI订阅用户免费)。当宿主有可用的Grok CLI——已安装并完成
    grok login
    ——illo可通过用户的Grok订阅,使用
    grok -p
    (无头模式)生成图片,消耗用户的Grok配额。与Codex采用相同的无环境变量、无令牌的子进程设计。Grok返回的JPEG无透明通道,因此无法生成透明抠图——此类请求会自动回退到支持抠图的后端。图像工具不支持模型选择。
  • Grok Bot原生传输(代理端,Grok Bot用户免费)。当你是Grok Bot——特指Cursor的Grok Bot/Grok桌面助手且具备内置Grok图像工具——构建illo提示并调用该工具,传入当前角色的模型表作为参考图片。无需Grok CLI、Codex CLI或OpenRouter密钥;不要将缺失引擎后端作为运行
    init
    的理由。这并非通用的“宿主图像API”规则,也不是
    illo.py --backend
    的有效值。
  • OpenRouter后端(付费,直接使用或显式回退)。需要用户配置文件中的OpenRouter API密钥——这是唯一的凭证渠道——由用户运行
    init
    命令写入(权限为600)。引擎从不从环境变量读取密钥,也不接受命令行传入的密钥。无订阅CLI的宿主可直接选择该引擎路径。Codex/Grok CLI渲染失败不会自动产生费用:付费回退需添加
    --allow-paid-fallback
    参数。该后端支持模型选择
    --model
    参数)。
后端/传输模型概要(优先级、CLI要求、Grok Bot原生路径、内置图像工具自动性、配额vs付费、抠图限制、Windows/WSL支持、回退机制):在选择或解释后端前,完整阅读
references/backends.md
——详细机制均在此文档中。

Setup is the user's job (never enter the key yourself)

配置由用户完成(切勿自行输入密钥)

Entering an API key is something the user does. Do not type, paste, print, or store the user's key — direct them to bootstrap it:
  • Bootstrap (user runs it):
    python3 "$SKILL_DIR/scripts/illo.py" init
    — prompts for the key at a hidden prompt (never echoed) and writes the YAML config
    ${XDG_CONFIG_HOME:-~/.config}/illo/config.yaml
    (mode 600). It can also store non-secret defaults:
    --model
    ,
    --palette
    ,
    --aspect
    ,
    --character
    ,
    --watermark
    . Use
    --no-key
    to update preferences without touching the stored key. (The config is read via PyYAML when installed; without it a minimal built-in parser still reads the flat keys —
    apiKey
    ,
    model
    , … — so generation needs no installs. Only nested settings like
    watermark
    need PyYAML:
    python -m pip install 'PyYAML==6.0.2'
    .)
  • Non-secret prefs may be seeded for the user with the same command and
    --no-key
    , but the key itself is theirs to enter.
输入API密钥是用户的操作。切勿输入、粘贴、打印或存储用户的密钥——指导用户自行完成初始化:
  • 初始化(用户运行)
    python3 "$SKILL_DIR/scripts/illo.py" init
    ——在隐藏提示符下请求密钥(不会回显),并将YAML配置写入
    ${XDG_CONFIG_HOME:-~/.config}/illo/config.yaml
    (权限为600)。该命令还可存储非机密默认值:
    --model
    --palette
    --aspect
    --character
    --watermark
    。使用
    --no-key
    参数可更新偏好设置而不修改已存储的密钥。(配置文件在安装PyYAML后通过PyYAML读取;若无PyYAML,内置的极简解析器仍可读取扁平密钥——
    apiKey
    model
    等——因此生成图片无需额外安装。仅嵌套设置如
    watermark
    需要PyYAML:
    python -m pip install 'PyYAML==6.0.2'
    。)
  • 非机密偏好可预先配置:使用相同命令加
    --no-key
    参数为用户预设,但密钥需由用户自行输入。

Hermes Agent only: binary asset repair preflight

仅Hermes Agent:二进制资源修复预检查

Some Hermes versions corrupt binary files (the bundled character sheets) when installing multi-file skills from GitHub — text files survive, binaries don't, and a corrupted sheet silently breaks the character lock. Under Hermes Agent, run this once before first use (and whenever
doctor
reports
assets: CORRUPTED
):
bash
bash ${HERMES_SKILL_DIR}/scripts/repair-hermes-assets.sh
It verifies every bundled binary against known-good SHA256 hashes (
assets/checksums.txt
) and re-downloads only mismatched files from pinned, immutable URLs — a no-op when everything checks out. Under Claude Code, Codex, OpenClaw, or any runtime that installs faithfully: skip this;
doctor
checks asset integrity everywhere and will say if repair is ever needed.
部分Hermes版本在从GitHub安装多文件技能时会损坏二进制文件(内置角色表)——文本文件不受影响,但二进制文件会损坏,且损坏的角色表会导致角色锁定失效。在Hermes Agent下,首次使用前(以及每当
doctor
报告
assets: CORRUPTED
时)运行一次以下命令:
bash
bash ${HERMES_SKILL_DIR}/scripts/repair-hermes-assets.sh
该命令会将每个内置二进制文件与已知正确的SHA256哈希值(
assets/checksums.txt
)进行校验,并仅重新下载不匹配的文件(从固定的不可变URL)——若所有文件均正常,则无任何操作。在Claude Code、Codex、OpenClaw或任何可正常安装的运行时环境中:跳过此步骤;
doctor
会在所有环境中检查资源完整性,并在需要修复时提示。

Read these references as needed

按需阅读以下参考文档

Do not load everything at once. Pull the file that matches the step:
  • references/visual-style.md
    — riso, the house default look: the risograph technique, line language, paper/ink, hard do/don'ts.
  • references/styles/<name>.md
    — the rest of the look library (
    blueprint
    ,
    woodcut
    ,
    pixel
    ,
    clay
    ,
    manila
    ,
    chalk
    ,
    phosphor
    ,
    enamel
    ,
    gouache
    ,
    felt
    ,
    diorama
    ,
    sketchbook
    ,
    bricks
    ,
    fizz
    ,
    bloom
    ,
    snes
    ), consumed by character packs. Read the active character's style file in full before generating.
  • references/character.md
    — the character rules (the load-bearing test, anti-complexity guardrails, value-follows-palette, the interaction model — declared per pack or derived conservatively from the locked design and reference sheet), the default character Blot, and the custom-pack format. Read before any character work.
  • references/character-builder.md
    — the guided flow for designing and installing a user's own mascot. Read in full before building or replacing a character.
  • references/pack-sharing.md
    — installing characters from the community repo and publishing a pack via PR. Read before any install/publish request.
  • references/palettes.md
    — named presets, default resolution, custom palettes, and the derive-a-palette-from-one-color algorithm. Read in full before choosing or deriving any palette.
  • references/composition.md
    — the two registers (editorial scene / explainer diagram), the diagram-type picker, the explainer's structure types and budget (including labeled stages, arrow notes, and its pack-solve), stagings, turning an idea into a move, the anatomy-action feasibility gate (validate the contact map against the character's interaction model before rendering), the no-recycled-composition rule, and the shot-list format.
  • references/cutout.md
    — the cutout register: transparent compositing assets, contact continuity, pose vocabulary, and generate flags. Read in full before any cutout request.
  • references/surprise.md
    — surprise / random mode: preflight-first, scope parse, random character, provenance variety + three saying candidates (optional parallel verify for sourced modes), interactive picker or
    --autopick
    / auto-pick-best, full re-roll on refresh, register after the locked saying, saying bar + sense bar, multi-source quote verification, safety-before-offer, headless contract. Read in full before any surprise/random request.
  • references/backends.md
    — the three-backend image engine plus the Grok Bot native transport: how the engine backend resolves (precedence Codex > Grok > OpenRouter, and the self-identify rule), when Grok Bot bypasses
    illo.py generate
    , the Codex/Grok CLI requirements, artifact-first success, the built-in image tool being automatic (no model selection), quota-vs-charge, Grok's no-cutout limit, Windows/WSL, and opt-in paid fallback. Read before choosing or explaining a backend.
  • references/models.md
    — the model lineup (OpenRouter backend only): friendly-name → OpenRouter id map, traits, aspect caveats, 404/fallback handling. Read before passing any
    --model
    .
  • references/prompt-recipe.md
    — the generation prompt template and the edit/recolor prompts.
  • references/quality-bar.md
    — the post-generation checklist and iteration rules. Read before delivering.
assets/character-reference.webp
is the default character's canonical model sheet — the consistency anchor (used by the engine, below); a custom pack brings its own. Style-calibration examples are not bundled — each style file links its own by URL (fetch when needed): study line density, negative space, and accent restraint. Never copy their compositions — invent a fresh metaphor for the current piece.
无需一次性加载所有内容。根据步骤选择对应文档:
  • references/visual-style.md
    ——riso风格,默认内置风格:risograph技术、线条语言、纸张/墨水、严格的注意事项。
  • references/styles/<name>.md
    ——其余风格库(
    blueprint
    woodcut
    pixel
    clay
    manila
    chalk
    phosphor
    enamel
    gouache
    felt
    diorama
    sketchbook
    bricks
    fizz
    bloom
    snes
    ),供角色包调用。生成前需完整阅读当前角色的风格文件。
  • references/character.md
    ——角色规则(核心测试、反复杂度约束、调色板匹配值、交互模型——每个包声明或从锁定设计和参考表保守推导)、默认角色Blot、自定义包格式。进行任何角色相关操作前需阅读。
  • references/character-builder.md
    ——设计和安装用户自定义吉祥物的引导流程。构建或替换角色前需完整阅读。
  • references/pack-sharing.md
    ——从社区仓库安装角色包,通过PR发布角色包。进行任何安装/发布请求前需阅读。
  • references/palettes.md
    ——命名预设、默认分辨率、自定义调色板、从单一颜色衍生调色板的算法。选择或衍生调色板前需完整阅读。
  • references/composition.md
    ——两种呈现形式(编辑场景/说明图)、图表类型选择器、说明图的结构类型和限制(包括标注阶段、箭头注释、角色包适配)、场景布置、将创意转化为画面、动作可行性检查(渲染前验证角色交互模型的接触映射)、禁止重复构图规则、镜头列表格式。
  • references/cutout.md
    ——抠图模式:透明合成素材、接触连贯性、姿势词汇、生成参数。进行任何抠图请求前需完整阅读。
  • references/surprise.md
    ——惊喜/随机模式:预检查优先、范围解析、随机角色、创意来源多样性+3种文案候选(可选并行验证来源模式)、交互式选择或
    --autopick
    /自动最优选择、刷新时重新生成、根据锁定文案选择呈现形式、文案栏+意义栏、多来源名言验证、安全优先、无头模式约定。进行任何惊喜/随机请求前需完整阅读。
  • references/backends.md
    ——三种图像引擎后端加Grok Bot原生传输:引擎后端解析方式(优先级Codex > Grok > OpenRouter,以及自识别规则)、Grok Bot何时绕过
    illo.py generate
    、Codex/Grok CLI要求、 artifact优先成功、内置图像工具自动性(无模型选择)、配额vs付费、Grok的抠图限制、Windows/WSL支持、付费回退选项。选择或解释后端前需阅读。
  • references/models.md
    ——模型列表(仅OpenRouter后端):友好名称→OpenRouter ID映射、特性、比例限制、404/回退处理。传入任何
    --model
    参数前需阅读。
  • references/prompt-recipe.md
    ——生成提示模板和编辑/重上色提示。
  • references/quality-bar.md
    ——生成后检查清单和迭代规则。交付前需阅读。
assets/character-reference.webp
是默认角色的标准模型表——一致性锚点(供引擎使用);自定义包自带模型表。风格校准示例未内置——每个风格文件通过URL链接到对应示例(按需获取):研究线条密度、负空间和装饰克制。切勿复制示例构图——为当前内容发明全新隐喻。

Workflow

工作流

0. Preflight

0. 预检查

Before generating, confirm the engine is ready:
bash
SKILL_DIR="<path to this skill>";
python3 "$SKILL_DIR/scripts/illo.py" doctor
Run the
illo.py
call standalone — never chained with
&&
— so the displayed exit code is the readiness signal itself (0 = ready): a chained neighbor's failure paints a healthy check as an error.
It reports python, the config path, the resolved model/palette defaults, whether a custom character pack or custom palettes file exists, Codex/Grok CLI detection and the resolved backend/transport, and whether an OpenRouter key is found (without revealing it); exit 0 = the resolved backend is ready. An OpenRouter-only install (no subscription CLI) stays exit 0 — readiness follows the resolved backend, not a hardwired key check (
references/backends.md
).
Grok Bot native path (agent-side). If you are Grok Bot (Cursor's Grok Bot / the Grok desktop assistant with the built-in Grok image tool) and the user has not explicitly chosen
backend: openrouter
,
backend: codex
, or
backend: grok
, initialize the agent-side transport before relying on
doctor
:
bash
SKILL_DIR="<path to this skill>";
python3 "$SKILL_DIR/scripts/illo.py" init --backend grok-bot --no-key
Then run
doctor
normally: exit 0 is meaningful readiness for this path. Missing Codex CLI, Grok CLI, or OpenRouter key are not failures when
backend: grok-bot
; corrupted assets, unreadable custom packs, malformed palettes, or the wrong
SKILL_DIR
still fail and must be fixed before rendering. Generate in Step 5 by calling Grok Bot's built-in Grok image tool with the active model sheet reference. This route is only for Grok Bot's own Grok image tool; other agents with unrelated image tools must use the engine backends below.
Short path for Grok Bot: run
init --backend grok-bot --no-key
once when backend is unset/auto, run
doctor
for assets/config/packs, use
packs
commands normally (including
packs install --all
after install so community characters are local), read the same references, build the same prompt, then call Grok Bot's built-in Grok image tool with the active character reference. Skip
illo.py init
for OpenRouter unless the user explicitly wants OpenRouter or another engine backend default, and skip
illo.py generate
unless the user explicitly selected an engine backend.
After the first successful Grok Bot install and bulk character install, ask once whether the user wants periodic checks for skill updates (
npx skills update
) and character updates (
packs update
). Default is off: if they say no, do not answer, or the host has no recurring-job mechanism, set up nothing. Only create a recurring check/reminder on an explicit yes, and never silently update — surface the proposed skill/character update and get consent before applying it.
Config migration — surface the backend choice interactively. When you are going to use
illo.py generate
, if
doctor
reports
backend: NEEDS CHOICE
(or
generate
hard-stops saying the config "is out of date"), this user's config predates the backend choice — they have an older install and have never been offered a subscription CLI. Do not pick for them silently. Surface an interactive choice using the platform's blocking-question capability (
AskUserQuestion
in Claude Code, the equivalent elsewhere; where the host has none — e.g. a plain chat session — ask the same one choice as a concise message and wait for the reply, never picking silently): "illo now has image backends/transports — which would you like?" with four options — Codex (free, your Codex subscription), Grok CLI (free, your Grok subscription; no transparent cutouts), Grok Bot (agent-side native tool; use only when you are Grok Bot), and OpenRouter (pick the model: Grok Imagine, Nano Banana, GPT Image, and others). Persist the answer without touching any existing key:
python3 "$SKILL_DIR/scripts/illo.py" init --backend <codex|grok|grok-bot|openrouter> --no-key
, then continue. A brand-new install (no config at all) is ordinary onboarding, not this migration — it does not fire.
Prefer your own CLI when you are a subscription-CLI agent. The engine's auto-default reads host capability (Codex > Grok > OpenRouter; it can't tell which agent invoked it) — but you know which agent you are. So when you are a subscription-CLI agent and your own CLI is usable on this host, add your own backend flag to
generate
for non-cutout renders: the Grok CLI agent adds
--backend grok
, the Codex agent adds
--backend codex
. This keeps "in Grok CLI, generate with Grok" true even on a host that also has Codex, with no runtime-sniffing in the engine. Cutouts ignore this (Grok can't make them — they auto-fall back). A user's config
backend:
overrides everything. Resolution and precedence mechanics:
references/backends.md
.
For Grok Bot, the equivalent self-identify rule happens before
generate
: when backend is unset/auto, persist
backend: grok-bot
with
init --backend grok-bot --no-key
and use the native Grok image tool path above. If the user explicitly configured or requested an engine backend, honor that choice instead of silently switching to Grok Bot native.
Read the printed config path before concluding the key is missing: under Hermes, multi-profile setups can resolve
HOME
/
XDG_CONFIG_HOME
to another profile's home (e.g.
…/profiles/<name>/home/.config/illo/…
), so a key that exists looks absent. If the path points at the wrong profile, re-run with the right
HERMES_HOME
/
HOME
/
XDG_CONFIG_HOME
rather than treating the key as missing. If the key is genuinely missing, stop and ask the user to run
python3 "$SKILL_DIR/scripts/illo.py" init
themselves — do not enter the key for them. In a chat session the user can't run commands where they are, so shrink their host-side step first: run
init --no-key
yourself (allowed — it scaffolds the config with defaults and a commented
# apiKey:
placeholder, mode 600, never touching a key), then offer the user two equivalent one-time options on the machine the agent runs on (that host is theirs — it's where they installed the agent): run
python3 <resolved absolute $SKILL_DIR>/scripts/illo.py init
(hidden prompt), or open
~/.config/illo/config.yaml
and fill in the
apiKey:
line. The key must never transit the chat: never ask for it in a message, and if the user pastes it anyway, do not use it — tell them to revoke that key at openrouter.ai and set a fresh one on the host (the pasted key now lives in chat history and platform servers). Never copy a key from the environment or any other store into the config yourself — the user is the only writer of that line — with one scoped exception: an ephemeral cloud workspace (Claude Code web, Codex cloud, CI) where the user provisioned
OPENROUTER_API_KEY
through the platform's secrets mechanism. That provisioning is itself the user's deliberate, workspace-scoped consent, and there is no interactive prompt or persistent home for
init
— so there, seed the config from the workspace secret once (the "Cloud & CI" one-liner in README.md). On a personal machine an ambient env var proves nothing about intent (it may belong to other tools) — the rule stands: never copy it.
Optional pack-freshness offer (preflight, consent-first). When this run will render with an installed community pack (
doctor
lists packs; installs carry a
.version
stamp), optionally check freshness:
python3 "$SKILL_DIR/scripts/illo.py" packs list
flags stale installs (
[installed 1.0.0 — 1.0.2 available]
). The check may run here, but the offer fires once the active pack is known — after Step 2 resolves the character (or after surprise mode's character roll), immediately before the first render that uses it. If that resolved pack is stale, offer once — via the platform's blocking-question capability, as in the config migration above — to refresh it before rendering, and run
packs update <name>
only on an explicit yes (updating overwrites the local copy; the hand-edit warning and
--as
alternative are in
references/pack-sharing.md
). Never update silently, and never block on this: a "no", an offline host, a registry error, or a headless/scheduler run (e.g. surprise
--autopick
) all continue with the pinned copy — a pack without a declared
## Interaction model
still plans safely via the conservative derivation (
references/character.md
). Skip the check entirely when no community-installed pack is involved.
生成前,确认引擎已就绪:
bash
SKILL_DIR="<技能路径>";
python3 "$SKILL_DIR/scripts/illo.py" doctor
单独运行
illo.py
命令——切勿用
&&
链式调用——以便显示的退出码直接表示就绪状态(0=就绪):链式调用中其他命令的失败会导致健康检查被误判为错误。
该命令会报告Python版本、配置路径、解析后的模型/调色板默认值、是否存在自定义角色包自定义调色板文件Codex/Grok CLI检测结果和解析后的后端/传输方式、是否找到OpenRouter密钥(不会泄露密钥);退出码0表示解析后的后端已就绪。仅安装OpenRouter的环境(无订阅CLI)仍会返回退出码0——就绪状态取决于解析后的后端,而非硬性密钥检查(详见
references/backends.md
)。
Grok Bot原生路径(代理端)。若你是Grok Bot(Cursor的Grok Bot/Grok桌面助手且具备内置Grok图像工具),且用户未明确选择
backend: openrouter
backend: codex
backend: grok
,则在依赖
doctor
前初始化代理端传输:
bash
SKILL_DIR="<技能路径>";
python3 "$SKILL_DIR/scripts/illo.py" init --backend grok-bot --no-key
然后正常运行
doctor
:退出码0表示此路径已就绪。当
backend: grok-bot
时,缺失Codex CLI、Grok CLI或OpenRouter密钥不会导致失败;但损坏的资源、无法读取的自定义包、格式错误的调色板或错误的
SKILL_DIR
仍会失败,必须在渲染前修复。步骤5中通过调用Grok Bot的内置Grok图像工具生成图片,并传入当前角色的模型表作为参考。此路径仅适用于Grok Bot自身的Grok图像工具;其他具备无关图像工具的代理必须使用以下引擎后端。
Grok Bot的简化流程:当后端未设置/自动时,运行一次
init --backend grok-bot --no-key
,运行
doctor
检查资源/配置/包,正常使用
packs
命令(包括安装后运行
packs install --all
以获取本地社区角色),阅读相同的参考文档,构建相同的提示,然后调用Grok Bot的内置Grok图像工具并传入当前角色参考。除非用户明确需要OpenRouter或其他引擎后端默认值,否则跳过
illo.py init
;除非用户明确选择引擎后端,否则跳过
illo.py generate
首次成功安装Grok Bot并批量安装角色后,询问用户是否需要定期检查技能更新(
npx skills update
)和角色更新(
packs update
)。默认关闭:若用户拒绝、未回答或宿主无定时任务机制,则不进行任何设置。仅在用户明确同意时创建定期检查/提醒,且绝不静默更新——需告知用户拟进行的技能/角色更新并获得同意后再执行。
配置迁移——交互式选择后端。当你要使用
illo.py generate
时,若
doctor
报告
backend: NEEDS CHOICE
(或
generate
停止并提示配置“已过期”),说明用户的配置早于后端选择功能——他们使用的是旧版本,从未被提供订阅CLI选项。切勿静默选择。使用平台的阻塞式提问功能(Claude Code中的
AskUserQuestion
,其他平台的等效功能;若宿主无此功能——如普通聊天会话——则以简洁消息提出相同问题并等待回复,绝不静默选择)呈现交互式选择:“illo现在支持多种图像后端/传输方式——你希望使用哪种?”提供四个选项——Codex(免费,使用你的Codex订阅)、Grok CLI(免费,使用你的Grok订阅;不支持透明抠图)、Grok Bot(代理端原生工具;仅当你是Grok Bot时使用)、OpenRouter(可选择模型:Grok Imagine、Nano Banana、GPT Image等)。保存用户选择,不修改任何现有密钥:
python3 "$SKILL_DIR/scripts/illo.py" init --backend <codex|grok|grok-bot|openrouter> --no-key
,然后继续操作。全新安装(无任何配置)属于常规引导,不属于此迁移场景——不会触发该流程。
订阅CLI代理优先使用自身CLI。引擎的自动默认值会读取宿主能力(Codex > Grok > OpenRouter;无法识别调用它的代理)——但你知道自己是哪个代理。因此,当你是订阅CLI代理且自身CLI在当前宿主可用时,为非抠图渲染添加自身后端参数:Grok CLI代理添加
--backend grok
Codex代理添加
--backend codex
。这样即使宿主同时安装了Codex,也能保证“在Grok CLI中使用Grok生成”;引擎无需运行时检测。抠图请求会忽略此参数(Grok无法生成抠图——会自动回退)。用户配置中的
backend:
会覆盖所有设置。解析和优先级机制详见
references/backends.md
对于Grok Bot,等效的自识别规则在**
generate
之前**执行:当后端未设置/自动时,通过
init --backend grok-bot --no-key
保存
backend: grok-bot
,并使用上述原生Grok图像工具路径。若用户明确配置或请求引擎后端,则优先遵循用户选择,而非静默切换到Grok Bot原生路径。
在判断密钥缺失前,先读取打印的配置路径:在Hermes下,多配置文件环境可能将
HOME
/
XDG_CONFIG_HOME
解析为另一配置文件的主目录(如
…/profiles/<name>/home/.config/illo/…
),导致存在的密钥看起来缺失。若路径指向错误的配置文件,需使用正确的
HERMES_HOME
/
HOME
/
XDG_CONFIG_HOME
重新运行,而非判定密钥缺失。若密钥确实缺失,停止操作并要求用户自行运行
python3 "$SKILL_DIR/scripts/illo.py" init
——切勿自行输入密钥。在聊天会话中,用户无法在当前位置运行命令,因此先简化宿主端步骤:自行运行
init --no-key
(允许操作——会用默认值和注释的
# apiKey:
占位符生成配置文件,权限为600,不会触碰密钥),然后为用户提供两个等效的一次性选项在代理运行的机器上(该宿主属于用户——是他们安装代理的地方):运行
python3 <解析后的绝对$SKILL_DIR>/scripts/illo.py init
(隐藏提示符),或打开
~/.config/illo/config.yaml
并填写
apiKey:
行。密钥绝不能通过聊天传输:切勿在消息中索要密钥,若用户粘贴密钥,切勿使用——告知用户在openrouter.ai撤销该密钥,并在宿主上设置新密钥。切勿从环境变量或其他存储中复制密钥到配置文件——只有用户能写入该行——仅有一种例外场景:临时云工作区(Claude Code web、Codex cloud、CI),用户通过平台的机密机制配置了
OPENROUTER_API_KEY
。这种配置本身是用户明确的、工作区范围内的同意,且
init
无交互式提示符或持久主目录——因此可在此场景下从工作区机密一次性生成配置(详见README.md中的“Cloud & CI”单行命令)。在个人机器上,环境变量不能证明用户意图(可能属于其他工具)——规则依然有效:切勿复制。
可选的包新鲜度提示(预检查,需用户同意)。当本次渲染将使用已安装的社区包(
doctor
会列出包;安装的包带有
.version
标记),可选择性检查新鲜度:
python3 "$SKILL_DIR/scripts/illo.py" packs list
会标记过时的安装包(
[installed 1.0.0 — 1.0.2 available]
)。可在此处检查,但提示需在确定当前包后触发——步骤2确定角色后(或惊喜模式的角色随机选择后),首次使用该包渲染前立即触发。若确定的包已过时,仅提示一次——通过平台的阻塞式提问功能,如配置迁移中所述——询问用户是否在渲染前更新,仅在用户明确同意时运行
packs update <name>
(更新会覆盖本地副本;
references/pack-sharing.md
中有手动编辑警告和
--as
替代方案)。绝不静默更新,也不因此阻塞操作:用户拒绝、宿主离线、注册表错误或无头/调度器运行(如惊喜模式
--autopick
)均会继续使用固定版本的包——未声明
## Interaction model
的包仍会通过保守推导安全运行(详见
references/character.md
)。若未涉及社区安装包,可完全跳过此检查。

1. Read the input — and clarify a thin concept (briefly)

1. 读取输入——并澄清模糊的创意(简要)

Three kinds of input, handled differently:
  • Surprise / random ("surprise me", "random", "surprise me with art quote using bray", "surprise me --autopick", and close variants) — the ask is invent-and-render, not a supplied thesis. Stop and read
    references/surprise.md
    in full
    , run Step 0 first, then resolve character and provenance there (ignore
    defaultCharacter
    ; random character when unnamed), build three saying candidates and lock one via picker or auto-pick-best, pick register from the locked saying, then continue Steps 3–7 as one image — Steps 0 and 2 are skipped in that render pass because preflight and pack are already done. Do not enter the thin-concept Q&A path below. A prompt that already names a concrete idea ("illustrate 'you are the bottleneck'") is not surprise mode even if it also says "surprise me".
  • A URL / article / paste / long post carries its own context — but never generate from the first vivid detail. Route it first (
    references/composition.md
    , "Source routing"): classify the source's shape and genre, infer the requested artifact's job (what this image must do for its audience), separate that job from the source's most drawable mechanism, lock the main thesis in one sentence (a hero locks the source/artifact job, not its loudest evidence — the genre guardrails say what each genre heroes), then pick the coverage — hero, hero + per-section set (the full article job), set, mini-comic, or shot list first. Sets need placements: compact sources (a tweet, one concept) never yield a set — their multi-beat form is the mini-comic. Pull the load-bearing moments — the few places that turn on a judgment, a loop, an input→output, a before/after, or a trap — never one image per paragraph. The text already says what it's about, so don't interrogate the user, with one exception: a materially multi-beat source (long article, postmortem, multi-claim launch) gets a single coverage question before any multi-image spend — unless the user already named the coverage. A lone image from a multi-beat source is a hero, delivered saying so — not as coverage of the piece.
  • A bare concept or one-liner (e.g. "illustrate 'you are the bottleneck'") usually underspecifies the picture. Ask up to ~3 quick questions — only the ones that change the output — then build. Draw from:
    • the single takeaway (what should the reader conclude?),
    • where it's headed (blog / deck / X post / X article body / X Article banner → sets palette, aspect, pixel normalization, and watermark),
    • the shape: one image (the default), a mini-comic (2–4 panels in one image — only when the idea itself advances through stages), or several separate images — plus any must-include element or constraint. The shape follows the idea, never the destination (
      references/composition.md
      ).
    Keep it to one short round, then proceed. Skip the questions entirely if the user already gave enough, said "just make it" / "single shot", or the answer is obvious from context. Never block a clear request by asking.
三种输入类型,处理方式不同:
  • 惊喜/随机模式(“给我惊喜”、“随机生成”、“用Bray生成艺术名言的惊喜插画”、“惊喜模式 --autopick”及类似变体)——需求是创意+渲染,而非提供核心论点。停止操作并完整阅读
    references/surprise.md
    ,先运行步骤0,然后在该文档中确定角色和创意来源(忽略
    defaultCharacter
    ;未指定角色时随机选择),构建3种文案候选并通过选择器或自动最优选择锁定,根据锁定的文案选择呈现形式,然后执行步骤3-7生成单张图片——该渲染流程跳过步骤0和2,因为预检查和包已处理完成。不要进入下方的模糊创意问答流程。若提示已明确具体创意(如“为‘你是瓶颈’生成插画”),即使包含“给我惊喜”,也不属于惊喜模式。
  • URL/文章/粘贴内容/长帖自带上下文——但绝不要从第一个生动细节开始生成。先进行路由(详见
    references/composition.md
    中的“来源路由”):分类来源的形状和体裁,推断请求产物的作用(该图片需为受众提供什么价值),区分该作用与来源中最具画面感的机制,用一句话锁定核心论点(主图锁定来源/产物的作用,而非最显眼的证据——体裁约束会说明每种体裁的主图内容),然后选择覆盖范围——主图、主图+分节系列图(完整文章需求)、系列图、迷你漫画或优先生成镜头列表。系列图需要放置位置:紧凑来源(如推文、单个概念)绝不会生成系列图——其多节拍形式对应迷你漫画。提取核心时刻——即涉及判断、循环、输入→输出、前后对比或陷阱的少数场景——绝不为每个段落生成一张图。文本已说明内容,因此无需询问用户,仅有一种例外:内容包含多个核心节拍的来源(长篇文章、事后分析、多主张发布)在生成多图前需先询问覆盖范围——除非用户已明确覆盖范围。多节拍来源的单张图片是主图,需明确告知用户——而非作为内容的完整覆盖。
  • 单一概念或短句(如“为‘你是瓶颈’生成插画”)通常对画面描述不足。最多问约3个简短问题——仅询问会影响输出的问题——然后开始构建。问题可来自:
    • 核心结论(读者应得出什么结论?),
    • 使用场景(博客/演示文稿/X平台帖子/X平台文章正文/X平台文章横幅→决定调色板、比例、像素标准化和水印),
    • 形式:单张图(默认)、迷你漫画(单张图片内含2-4个分镜——仅当创意需分阶段推进时)或多张独立图片——外加任何必须包含的元素或约束。形式由创意决定,而非目标载体(详见
      references/composition.md
      )。
    保持一轮简短提问,然后继续操作。若用户已提供足够信息、说“直接生成”/“单张图”,或答案从上下文可明显推断,则完全跳过提问。绝不要因提问而阻塞明确的请求。

2. Resolve the character

2. 确定角色

Surprise / random mode: skip this step — character was already resolved in
references/surprise.md
(named pack, or random among installed + Blot; never
defaultCharacter
). Continue at Step 3+.
Installed packs live under
${XDG_CONFIG_HOME:-~/.config}/illo/characters/
(format and location details:
references/character.md
);
doctor
lists what's installed. A user can keep several and pick per run. First match wins:
  1. Explicit request — "use <pack name>", "as <name>": that pack (or the shipped default when asked for by name,
    blot
    ). When the word matches no pack name, resolve by approximation: match it against each installed pack's
    Aliases:
    line and subject (the
    character.md
    opening line and Locked design Body) —
    doctor
    prints names + aliases, so this needs no file reads in the common case — and against catalog
    description
    s (
    packs list
    ). So "use ox" finds a pack subtitled an ox (e.g.
    yoke
    ). On one clear match, use it and name it; on several, ask which; on none, say so before falling through.
  2. Config default
    defaultCharacter
    from the user config, if set.
  3. Shipped defaultBlot (spec in
    references/character.md
    , model sheet
    assets/character-reference.webp
    ).
Once resolved, read the pack's
character.md
and use its prompt spec, value rules, optional
Cutout chroma:
compatibility preference, and
reference.png
everywhere the default's would be used.
When rerouting an article set to a new character — especially after a weak attempt, or for a technical/platform essay — read
references/article-set-character-reroute.md
in full before planning or rendering. Do the legibility preflight there before spending renders.
If the user wants a new character, that is the character builder (
references/character-builder.md
); if they want someone else's, packs install from the community repo (
references/pack-sharing.md
). Either way, install first, then continue here.
惊喜/随机模式:跳过此步骤——角色已在
references/surprise.md
中确定(命名包,或从已安装包+Blot中随机选择;绝不使用
defaultCharacter
)。继续步骤3+。
已安装的包位于
${XDG_CONFIG_HOME:-~/.config}/illo/characters/
(格式和位置详情见
references/character.md
);
doctor
会列出已安装的包。用户可安装多个包并每次选择使用。匹配优先级如下:
  1. 明确请求——“使用<包名>”、“以<名称>风格”:使用该包(若请求默认包名
    blot
    ,则使用内置默认包)。若名称未匹配任何包名,通过近似匹配解决:将其与每个已安装包的
    Aliases:
    行和主题(
    character.md
    开头行和锁定设计的Body)匹配——
    doctor
    会打印名称+别名,因此通常无需读取文件——并与目录
    description
    匹配(
    packs list
    )。例如“使用ox”会找到副标题为ox的包(如
    yoke
    )。若匹配唯一,使用该包并告知用户;若匹配多个,询问用户选择;若无匹配,告知用户后继续下一步。
  2. 配置默认值——用户配置中的
    defaultCharacter
    (若已设置)。
  3. 内置默认值——Blot(详见
    references/character.md
    ,模型表为
    assets/character-reference.webp
    )。
确定角色后,读取该包的
character.md
,并在所有使用默认角色的地方使用其提示规范、价值规则、可选的**
Cutout chroma:
**兼容性偏好和
reference.png
当为文章系列图重新选择角色时——尤其是初次尝试效果不佳,或针对技术/平台类文章——在规划或渲染前需完整阅读
references/article-set-character-reroute.md
。在生成前需完成可读性预检查。
若用户需要角色,使用角色构建器(
references/character-builder.md
);若需要他人的角色,从社区仓库安装包(
references/pack-sharing.md
)。无论哪种情况,先安装,然后回到此步骤。

3. Plan (shot list) — when asked to plan, or for anything multi-image

3. 规划(镜头列表)——当用户要求规划,或任何多图需求

If the user wants planning ("where should this be illustrated", "shot list"), output a shot list before generating. Per image: placement, the one idea, the artifact job, the register (editorial unless the row passes the explainer gate), the staging (or structure type — pick per
references/composition.md
, "Pick the diagram type"), what the mascot is doing, the palette, and the text hierarchy — primary read/title when the artifact needs one, plus short supporting labels/callouts within the per-register budgets in
references/composition.md
. Let the anchor count drive how many (bands and the never-pad rule are in
references/composition.md
). When a stretch of the piece advances through stages in one place, plan a single mini-comic image there instead of several — the mini-comic-vs-separate routing is in
references/composition.md
.
For article-set character reroutes, add the mandatory preflight fields from
references/article-set-character-reroute.md
before any render: section claim, visual object/action, and reader mapping. Reject rows that need a private metaphor glossary or more than one conceptual substitution.
若用户需要规划(“哪里需要插画”、“镜头列表”),在生成前输出镜头列表。每张图需包含:位置、核心创意、产物作用、呈现形式(默认编辑场景,除非符合说明图条件)、场景布置(或结构类型——根据
references/composition.md
中的“选择图表类型”选择)、吉祥物的动作、调色板、文本层级——产物需要的主标题/核心文字,以及呈现形式限制内的简短支持标签/标注(详见
references/composition.md
)。锚点数量决定图片数量(
references/composition.md
中有分组和禁止填充规则)。当内容的某一部分在同一场景内分阶段推进时,规划为单张小漫画图片,而非多张——小漫画vs独立图片的路由规则详见
references/composition.md
对于重新选择角色的文章系列图,在渲染前需添加
references/article-set-character-reroute.md
中的必填预检查字段:章节论点、视觉对象/动作、读者映射。拒绝需要私有隐喻词汇表或多个概念替换的条目。

4. Resolve the palette (the style is the character's)

4. 确定调色板(风格由角色决定)

Style is not separately resolvable: the active character's pack carries it — the
Style:
line in its
character.md
names a bundled look (
references/styles/<name>.md
, riso in
visual-style.md
) or a custom one at
${XDG_CONFIG_HOME:-~/.config}/illo/styles/<name>.md
; absent line = riso. Blot is riso. For any non-riso style, read its file in full: it supplies the STYLE and LINE LANGUAGE prompt blocks, the palette mapping, the character treatment, and extra QA checks. A request for the same character in a different look is a variant-pack build (route table) — never restyle on the fly.
Palette: read
references/palettes.md
in full and resolve there — it holds the resolution order (explicit request, then destination cue via the user's palettes file, then config default, then house
ink-punch
), the named presets, custom palettes, and the derive-a-palette-from-one-color algorithm. End with concrete hex values; when the pack's style isn't riso, run them through that style's palette mapping.
风格不可单独选择:当前角色的包自带风格——其
character.md
中的
Style:
行指定内置风格(
references/styles/<name>.md
,riso风格见
visual-style.md
)或自定义风格(位于
${XDG_CONFIG_HOME:-~/.config}/illo/styles/<name>.md
);若无该行则默认riso风格。Blot为riso风格。对于非riso风格,需完整阅读其风格文件:文件提供STYLE和LINE LANGUAGE提示块、调色板映射、角色处理方式和额外QA检查。若请求同一角色使用不同风格,需构建变体包(见路由表)——绝不要临时修改风格。
调色板:完整阅读
references/palettes.md
并确定——该文档包含优先级(明确请求、用户调色板文件中的目标提示、配置默认值、内置
ink-punch
调色板)、命名预设、自定义调色板、从单一颜色衍生调色板的算法。最终需确定具体十六进制值;若包的风格非riso,需通过该风格的调色板映射处理这些值。

5. Generate — reference-locked, one metaphor per image

5. 生成——参考锁定,每张图一个隐喻

Cutout branch. When the request routed to the cutout register, read
references/cutout.md
in full first — it covers backend-aware transparency (Codex native alpha by default; chroma compatibility for OpenRouter or explicit
--chroma
), registration-locked silhouette (no ink-layer offset),
--cutout
/
--aspect 1:1
, OpenRouter
--image-config
, and manifest
cutout_alpha
disclosure. Build the prompt from
references/prompt-recipe.md
, "Cutout variant" — not the editorial template — and omit manual
BACKGROUND:
/ output-format instructions; the engine appends the contract for the backend that actually runs. Pass
--chroma
only to force a compatibility reroll. Use only the character model sheet as
--ref
(no editorial style anchor, no watermark). QA against the cutout section of
references/quality-bar.md
. Skip the editorial shot-list / thesis steps.
Editorial and explainer. When the locked type is labeled stages, run the pack-solve scratch in
references/composition.md
("Labeled stages — skeleton, then pack-solve") before writing the prompt — stage list → operator stage → contact map → bind; do not invent a look. Build a full prompt per image from
references/prompt-recipe.md
(scene + structure + communication hierarchy + style + the active character's spec + resolved palette hexes + the per-register text budget), write it to a file, and render it. Pass the active character's model sheet as
--ref
every time
— that reference conditioning is what keeps the mascot on-model; style and palette come from the prompt, so both stays swappable. A pack's sheet is born in its own style, so sheet and style always match — no cross-style reference juggling. (Under Hermes Agent, the asset-repair preflight above must have run before the first
--ref
use — a corrupted sheet conditions every render on garbage.)
Grok Bot native render. If you are Grok Bot and the native path from Step 0 applies, do not run
illo.py generate
. Use the same full prompt recipe, same aspect ratio, same character lock, same style-anchor rule for sets, and call Grok Bot's built-in Grok image tool. Attach the active character's model sheet as a reference image (
assets/character-reference.webp
for Blot, or the pack's
reference.png
); for later images in a set, also attach the accepted style anchor image. Ask the tool to save/return the generated file and treat that saved path as the engine JSON
.path
equivalent for QA and delivery. Grok Bot's image tool is the same Grok image-model class as the Grok CLI transport: no model selector, no OpenRouter billing, and no alpha channel. Transparent cutouts stay off this path; route them to a cutout-capable engine backend instead, or stop and ask for that backend to be configured.
Set
SKILL_DIR
inline (see Prerequisites), and use the bundled sheet as
REF
— or the active pack's
reference.png
for a custom character. Add
--model <id>
to override the config/default model for this image (OpenRouter backend only):
bash
SKILL_DIR="<path to this skill>";
REF="$SKILL_DIR/assets/character-reference.webp";
python3 "$SKILL_DIR/scripts/illo.py" generate --prompt-file /tmp/shot-01.txt --ref "$REF" --aspect 16:9 --out "assets/<slug>-illustrations/01-topic.png"
For engine renders,
illo.py generate
prints a JSON line per image (
{path, backend, model, id, cost, width, height, label, prompt}
;
backend
is
codex
,
grok
, or
openrouter
, and
model
/
id
/
cost
are OpenRouter-only — they are null on a CLI-served record (Codex or Grok).
cost
is null unless
--cost
is passed —
gallery
backfills it) and appends the same record to
<out-dir>/manifest.jsonl
. Read
.path
— it may differ from
--out
: the engine names the file by the actual encoding (some models return JPEG bytes, so a requested
.png
lands as
.jpg
). Use
.width/.height
to catch a square when 16:9 was requested (re-roll). A failed Codex/Grok CLI render stops by default even when an OpenRouter key is configured. Add
--allow-paid-fallback
only when the user has explicitly approved a pay-per-image retry. Direct
--backend openrouter
renders and the intentional Grok-cutout redirect remain direct routes and do not need this flag. Generate each image separately — never combine ideas into one canvas. Default aspect is 16:9; use
1:1
for square social,
9:16
/
4:5
for vertical, and
1536:640
for an X Article banner / hero. For X Article banners, the platform target is 1536 × 640 px. Generate through the normal image pipeline; do not manually composite or rebuild the scene from crops as a substitute for an illo render. Check
.width/.height
, and only do final post-processing when it is a non-distorting resize/crop that preserves normal proportions and all essential information. Never stretch or squash the art to force exact dimensions. Pass
--label
for a caption that shows in the gallery.
Sets read as one artist. For any multi-image set, the first image that passes the full quality bar (and, for a hero in a rerouted article set, passes the thesis-legibility gate in
references/article-set-character-reroute.md
; never anchor on an unvetted render — a failed anchor, e.g. an off-palette ground or illegible metaphor, would propagate its failure set-wide) becomes the set's style anchor: pass it as a second
--ref
after the character sheet for every later image in the set and for every re-roll of a set member, so line weight, halftone density, and flat-vs-dimensional treatment stay consistent throughout. The same trick locks style for a one-off: add any finished example as a second
--ref
.
Model choice (OpenRouter backend only).
--model
and config
model:
are an OpenRouter-only axis — on Codex, Grok CLI, and Grok Bot native the image model is automatic and
--model
does not apply (
references/backends.md
). For the OpenRouter path, read
references/models.md
in full before passing any
--model
(or whenever the user names a model in plain language or asks for "best quality" / "cheapest"): it holds the friendly-name → OpenRouter id map, per-model traits, the aspect-ratio caveat, and the 404/fallback handling. Resolution is
--model
> config
model
> built-in default.
Watermark / attribution (optional, off by default). The skill ships with no default watermark — the text comes only from the user's
watermark
config map (read from the config file) or an explicit request, so installers never inherit someone else's handle. The resolution order, the prompt line to append, and the two-render caveat are in
references/prompt-recipe.md
.
抠图分支。若请求路由到抠图模式,先完整阅读
references/cutout.md
——该文档包含后端感知的透明度(默认Codex原生透明通道;OpenRouter或显式
--chroma
参数使用色度键)、注册锁定轮廓(无墨层偏移)、
--cutout
/**
--aspect 1:1
参数、OpenRouter的
--image-config
参数、清单中的
cutout_alpha
**披露。使用
references/prompt-recipe.md
中的“抠图变体”构建提示——而非编辑场景模板——并省略手动
BACKGROUND:
/输出格式说明;引擎会为实际运行的后端添加约定。仅在需要强制兼容性重新生成时添加
--chroma
参数。仅使用角色模型表作为
--ref
(无编辑场景风格锚点、无水印)。根据
references/quality-bar.md
中的抠图部分进行QA。跳过编辑场景的镜头列表/核心论点步骤。
编辑场景和说明图。若锁定类型为标注阶段,在编写提示前先完成
references/composition.md
中的包适配(“标注阶段——框架,然后包适配”)——阶段列表→操作阶段→接触映射→绑定;不要发明新风格。为每张图从
references/prompt-recipe.md
构建完整提示(场景+结构+沟通层级+风格+当前角色规范+已确定的调色板十六进制值+呈现形式的文本限制),写入文件并渲染。每次都要传入当前角色的模型表作为
--ref
——参考条件是保持吉祥物风格一致的关键;风格和调色板来自提示,因此两者均可替换。包的模型表以自身风格创建,因此模型表与风格始终匹配——无需跨风格参考调整。(在Hermes Agent下,首次使用
--ref
前需运行上述资源修复预检查——损坏的模型表会导致所有渲染基于错误参考。)
Grok Bot原生渲染。若你是Grok Bot且步骤0中的原生路径适用,不要运行
illo.py generate
。使用相同的完整提示模板、相同比例、相同角色锁定、相同系列图风格锚点规则,调用Grok Bot的内置Grok图像工具。附加当前角色的模型表作为参考图片(Blot为
assets/character-reference.webp
,自定义角色为包的
reference.png
);对于系列图中的后续图片,还需附加已确认的风格锚点图片。要求工具保存/返回生成的文件,并将保存路径视为引擎JSON中的
.path
等效值,用于QA和交付。Grok Bot的图像工具与Grok CLI传输使用相同的Grok图像模型类:无模型选择器、无OpenRouter计费、无透明通道。透明抠图不使用此路径;需路由到支持抠图的引擎后端,或停止操作并要求用户配置该后端。
内联设置
SKILL_DIR
(见前置条件),使用内置模型表作为
REF
——或自定义角色的包
reference.png
。添加
--model <id>
参数可覆盖此图片的配置/默认模型(仅OpenRouter后端):
bash
SKILL_DIR="<技能路径>";
REF="$SKILL_DIR/assets/character-reference.webp";
python3 "$SKILL_DIR/scripts/illo.py" generate --prompt-file /tmp/shot-01.txt --ref "$REF" --aspect 16:9 --out "assets/<slug>-illustrations/01-topic.png"
对于引擎渲染,
illo.py generate
会为每张图打印一行JSON
{path, backend, model, id, cost, width, height, label, prompt}
backend
codex
grok
openrouter
model
/
id
/
cost
仅适用于OpenRouter——Codex或Grok的CLI服务记录中这些值为null)。
cost
仅在传入
--cost
参数时不为null——
gallery
会补充该值。同时会将相同记录追加到
<out-dir>/manifest.jsonl
。读取
.path
——可能与
--out
不同:引擎会根据实际编码命名文件(部分模型返回JPEG字节,因此请求的
.png
会保存为
.jpg
)。使用
.width/.height
检查是否在请求16:9时生成了正方形图片(需重新生成)。默认情况下,即使配置了OpenRouter密钥,Codex/Grok CLI渲染失败也会停止。仅当用户明确批准按次付费重试时,添加
--allow-paid-fallback
参数。直接指定
--backend openrouter
的渲染和Grok抠图的重定向属于直接路径,无需此参数。单独生成每张图片——绝不要将多个创意合并到一个画布。默认比例为16:9;方形社交内容使用
1:1
,垂直内容使用
9:16
/
4:5
X平台文章横幅/主图使用
1536:640
。对于X平台文章横幅,目标尺寸为1536 × 640 px。通过常规图像流水线生成;不要手动合成或从裁剪图重建场景来替代illo渲染。检查
.width/.height
,仅在需要非变形调整大小/裁剪以保持正常比例和所有关键信息时进行最终后期处理。绝不要拉伸或挤压插画以强制匹配精确尺寸。传入
--label
参数添加在gallery中显示的标题。
系列图保持统一风格。对于任何多图系列,第一张通过完整质量检查的图片(对于重新选择角色的文章系列图的主图,还需通过
references/article-set-character-reroute.md
中的核心论点可读性检查;绝不要以未经验证的渲染作为锚点——失败的锚点,如调色板错误或隐喻模糊,会导致整个系列失败)成为系列的风格锚点:在系列后续每张图的
--ref
参数中,将其作为第二个参考传入角色表之后,系列中任何图片的重新生成也需如此,以保持线条粗细、半色调密度、平面/立体处理的一致性。该技巧也可锁定单张图的风格:添加任何已完成的示例作为第二个
--ref
模型选择(仅OpenRouter后端)
--model
参数和配置中的
model:
仅OpenRouter后端的选项——Codex、Grok CLI和Grok Bot原生的图像模型是自动选择的,
--model
参数不适用(详见
references/backends.md
)。对于OpenRouter路径,在传入任何
--model
参数前(或当用户用普通语言命名模型或询问“最佳质量”/“最便宜”时)需完整阅读
references/models.md
:该文档包含友好名称→OpenRouter ID映射、每个模型的特性、比例限制、404/回退处理。优先级为
--model
> 配置
model:
> 内置默认值。
水印/署名(可选,默认关闭)。该技能默认水印——文本仅来自用户的
watermark
配置映射(从配置文件读取)或明确请求,因此安装者不会继承他人的账号。优先级、需追加的提示行、两次渲染的注意事项详见
references/prompt-recipe.md

5b. Batches & comparison (only when it helps)

5b. 批量生成与对比(仅在有帮助时使用)

Default to ONE image. Fan out only when the user asks for options/comparison or the piece is important enough to be worth it — and say first what each image costs: on the Codex backend it draws on the user's Codex quota (no per-image charge), on Grok CLI or Grok Bot native it draws on the user's Grok quota, and on the OpenRouter backend it bills their OpenRouter account (typically under ten cents per image, varying by model). Keep N small (2–4). Orchestrate the loop with the engine's primitives:
newrun
prints a fresh run dir (
/tmp/illo/<runid>
) into
RUN
. Record the user's VERBATIM request (URL, pasted text, concept) to
request.txt
— the gallery shows it as provenance so anyone can tell what the run was for. Adapt each
generate
line below to a real path and run it on its own:
bash
SKILL_DIR="<path to this skill>";
RUN=$(python3 "$SKILL_DIR/scripts/illo.py" newrun);
printf '%s' "<the verbatim request>" > "$RUN/request.txt"
默认生成一张图。仅当用户要求选项/对比,或内容足够重要值得投入时才生成多张——需先说明每张图的成本:Codex后端消耗用户的Codex配额(无按次费用),Grok CLI或Grok Bot原生消耗用户的Grok配额,OpenRouter后端从用户的OpenRouter账户扣费(通常每张图不到10美分,因模型而异)。保持数量较少(2-4张)。使用引擎原语编排循环:
newrun
会打印新的运行目录(
/tmp/illo/<runid>
)到
RUN
。将用户的原始请求(URL、粘贴文本、概念)记录到
request.txt
——gallery会将其作为来源显示,以便任何人了解本次运行的目的。将以下每个
generate
行调整为实际路径并单独运行:
bash
SKILL_DIR="<技能路径>";
RUN=$(python3 "$SKILL_DIR/scripts/illo.py" newrun);
printf '%s' "<原始请求>" > "$RUN/request.txt"

(a) VARIATIONS — same prompt+model, pick-the-best:

(a) 变体——相同提示+模型,选择最佳:

python3 .../illo.py generate --prompt-file p.txt --ref <ref> --count 4 --label "draft→ship" --out "$RUN/v.png"
python3 .../illo.py generate --prompt-file p.txt --ref <ref> --count 4 --label "draft→ship" --out "$RUN/v.png"

(b) MODEL COMPARISON — loop the SAME prompt over the chosen models

(b) 模型对比——同一提示循环测试所选模型

(full OpenRouter ids from references/models.md):

(完整OpenRouter ID见references/models.md):

for m in <model-id-1> <model-id-2>; do python3 .../illo.py generate --prompt-file p.txt --ref <ref> --model "$m" --label "$m" --out "$RUN/$(basename $m).png"; done
for m in <model-id-1> <model-id-2>; do python3 .../illo.py generate --prompt-file p.txt --ref <ref> --model "$m" --label "$m" --out "$RUN/$(basename $m).png"; done

(c) CONCEPT VARIATIONS — different prompts (different stagings) for one idea:

(c) 概念变体——同一创意的不同提示(不同场景布置):

python3 .../illo.py generate --prompt-file staging-A.txt --ref <ref> --label "as a funnel" --out "$RUN/a.png" python3 .../illo.py generate --prompt-file staging-B.txt --ref <ref> --label "as a crossing" --out "$RUN/b.png"
python3 "$SKILL_DIR/scripts/illo.py" gallery "$RUN" --title "<the piece or request>" --open
python3 .../illo.py generate --prompt-file staging-A.txt --ref <ref> --label "漏斗风格" --out "$RUN/a.png" python3 .../illo.py generate --prompt-file staging-B.txt --ref <ref> --label "交叉风格" --out "$RUN/b.png"
python3 "$SKILL_DIR/scripts/illo.py" gallery "$RUN" --title "<内容或请求>" --open

always pass --title so a saved gallery stays identifiable later;

始终传入--title,以便保存的gallery后续可识别;

add --embed for a single portable file (images inlined)

添加--embed参数生成单个可移植文件(图片内嵌)


Every `generate` self-records to `$RUN/manifest.jsonl`; `gallery` assembles them
into one page with each image's **label, model, dimensions, cost, and a
collapsible prompt** — the prompt toggle is what makes concept-variation
comparison readable (the prompt is the variable). Always present the gallery
**with a recommendation**, not a raw dump — and in a chat session, present
the labeled candidates directly in the chat instead of a gallery (delivery
routing in step 7). Multi-model failures are per-image
(an unavailable model errors that one render only); keep the rest.

每个`generate`会自动记录到`$RUN/manifest.jsonl`;`gallery`会将它们组装成一个页面,包含每张图片的**标签、模型、尺寸、成本和可折叠的提示**——提示切换功能使概念变体对比更易读(提示是变量)。始终**附带推荐**呈现gallery,而非原始输出——在聊天会话中,直接在聊天中呈现带标签的候选,而非gallery(交付路由见步骤7)。多模型失败为单张图失败(不可用模型仅导致该张渲染失败);保留其余图片。

6. QA and iterate

6. QA与迭代

Check every image against
references/quality-bar.md
. Re-roll or edit when the mascot is decorative or off its locked spec, the body is wrong-value for the palette, label text sits on a colored fill, the accent has spread past the character's accent part + 1–2 elements, an unwanted title bar appears, the composition copies an example, or text is misspelled. Subject scale varies run-to-run — re-roll if the subject is tiny (check
.width/.height
in the JSON: a square back when 16:9 was requested → re-roll). When a re-roll supersedes a render, rebuild any delivery gallery with
--exclude <superseded label>
(repeatable) so rejected rolls don't appear in the review artifact.
根据
references/quality-bar.md
检查每张图片。当吉祥物过于装饰化或偏离锁定规范、主体与调色板不匹配、标签文字位于彩色填充上、装饰超出角色的装饰部分+1-2个元素、出现多余标题栏、构图复制示例或文字拼写错误时,重新生成或编辑。主体比例每次运行可能不同——若主体过小则重新生成(检查JSON中的
.width/.height
:请求16:9却生成正方形→重新生成)。当重新生成替代原有渲染时,使用
--exclude <被替代标签>
(可重复)重建交付gallery,使被拒绝的渲染不会出现在评审产物中。

7. Deliver — match the session's medium

7. 交付——匹配会话媒介

Copy finals next to the user's work when appropriate; never overwrite existing assets without being asked. Filenames carry the role — they are the only metadata that survives a document attachment, so make them self-identifying:
00-hero-<slug>.png
for the hero, then
01-<section-slug>.png
,
02-<section-slug>.png
, … for anchors in piece order (
assets/<slug>-illustrations/
). Then report: how many images, the palette used, which are strongest vs optional — and for any multi-image job, a placement map: one line per image naming the file, its role (hero, or after which section), and the one idea it lands, so the user can drop each file where it belongs without re-deriving the plan. Deliver the images themselves the way this session can actually show them:
  • Filesystem sessions (IDE/terminal agents — Claude Code, Codex, Cursor): report each final's absolute path (the engine's JSON
    .path
    is already absolute) and present the gallery for multi-image runs. The file on disk already is the original — never emit
    [[as_document]]
    here: it's a Hermes gateway token, literal noise in any other runtime. If the runtime has its own in-chat file delivery, use that.
  • Grok Bot native sessions: deliver the file returned by Grok Bot's built-in image tool inline/as an attachment in chat, and include its saved file path in the same role that engine renders use
    .path
    . The returned file is the original for this transport; do not ask the user to configure OpenRouter just to retrieve it.
  • Chat sessions (the user is on a messaging surface — Hermes over Telegram/Discord/WhatsApp, or any chat surface with lossy media delivery — and cannot open local files): a path alone is not a complete deliverable; the image must land in the chat, and a final must arrive as the original file. Platform photo delivery recompresses images — exactly what destroys riso grain, halftone texture, ink-layer offset, and fine hand-lettering — so finals are delivered as document attachments. On Hermes, tag each final with an explicit
    MEDIA:
    attachment tag — the tag is
    MEDIA:
    immediately followed by the absolute path, no space — and the literal directive
    [[as_document]]
    in the same reply. Do not rely on a bare absolute path for a final: bare paths can pass through to the user as literal text instead of being dispatched as an attachment.
    text
    MEDIA:/absolute/path/to/final.jpg
    [[as_document]]
    Candidate/options rounds may use normal inline photo delivery when quick glances help — say so ("preview — original file to follow") — but a final is never delivered that way. Skip the HTML gallery in chat — the user has no easy way to open or host it; send the labeled finals directly with the recommendation as text, and only build
    gallery --embed
    (one self-contained file) if a portable artifact is explicitly requested, delivering it with
    [[as_document]]
    .
Before the final reply in a chat session, check:
  • every final's path came from the engine's JSON
    .path
    , not the requested
    --out
    (the actual extension may differ);
  • every final appears as an explicit
    MEDIA:/absolute/path
    attachment tag in the reply;
  • [[as_document]]
    is in the reply unless this is explicitly preview-only;
  • rejected/re-rolled candidates are excluded from delivery;
  • the text says what was made — character, palette, strongest final, and for sets the placement map (which file is the hero, which follows which section) — without implementation noise.
在合适时将最终文件复制到用户工作目录旁;未经请求绝不覆盖现有资源。文件名需体现作用——这是唯一能在文档附件中保留的元数据,因此需具备自识别性:主图命名为
00-hero-<slug>.png
,然后按内容顺序命名锚点图为
01-<section-slug>.png
02-<section-slug>.png
…(位于
assets/<slug>-illustrations/
目录)。然后报告:图片数量、使用的调色板、最佳图片与可选图片——对于多图任务,需提供放置映射:每行对应一张图,包含文件名、作用(主图或位于某章节之后)、核心创意,以便用户无需重新推导规划即可将每个文件放到对应位置。以会话可展示的方式交付图片:
  • 文件系统会话(IDE/终端代理——Claude Code、Codex、Cursor):报告每个最终文件的绝对路径(引擎JSON中的
    .path
    已为绝对路径),多图任务呈现gallery。磁盘上的文件已是原始文件——切勿在此处使用
    [[as_document]]
    :这是Hermes网关令牌,在其他运行时中是无效内容。若运行时支持聊天内文件交付,使用该功能。
  • Grok Bot原生会话:将Grok Bot内置图像工具返回的文件以内联/附件形式在聊天中交付,并将保存路径作为引擎渲染的
    .path
    使用。返回的文件是该传输方式的原始文件;切勿要求用户配置OpenRouter来获取文件。
  • 聊天会话(用户在消息界面——Hermes通过Telegram/Discord/WhatsApp,或任何有损媒体交付的聊天界面——无法打开本地文件):仅提供路径并非完整交付;图片必须在聊天中显示,且最终文件需以原始文件形式交付。平台照片交付会重新压缩图片——这会破坏riso颗粒、半色调纹理、墨层偏移和精细手写文字——因此最终文件需以文档附件形式交付。在Hermes中,为每个最终文件添加明确的
    MEDIA:
    附件标签——标签为
    MEDIA:
    后紧跟绝对路径,无空格——并在同一回复中添加文字指令
    [[as_document]]
    切勿仅依赖绝对路径交付最终文件:纯路径可能会作为纯文本传递给用户,而非作为附件发送。
    text
    MEDIA:/absolute/path/to/final.jpg
    [[as_document]]
    候选/选项轮次可使用常规内联照片交付(快速预览)——需说明(“预览——原始文件随后发送”),但最终文件绝不能以此方式交付。聊天中跳过HTML gallery——用户无法轻松打开或托管;直接发送带标签的最终文件并附带文字推荐,仅当明确请求可移植产物时才生成
    gallery --embed
    (单个自包含文件),并以
    [[as_document]]
    交付。
聊天会话的最终回复前,检查:
  • 每个最终文件的路径来自引擎JSON的
    .path
    ,而非请求的
    --out
    (实际扩展名可能不同);
  • 每个最终文件在回复中均有明确的
    MEDIA:/absolute/path
    附件标签;
  • 回复中包含
    [[as_document]]
    (除非明确为预览);
  • 被拒绝/重新生成的候选未包含在交付中;
  • 文字说明生成的内容——角色、调色板、最佳最终文件,以及系列图的放置映射(哪个文件是主图,哪个位于哪个章节之后)——无实现细节噪音。

Output discipline

输出规范

Pre-generation planning is short and concrete. Post-generation, let the images speak — report what was made and where, not style theory. Keep labels few and short; the fewer words baked into an image, the more reliably it renders.
Talk like a person doing the work, not a recap of this file. Never narrate workflow steps or jargon in chat: doctor, preflight, provenance, register, saying bar, thesis, backend, or "doctor's green." Status, if any, is ordinary speech, not a liturgy of steps. When surprise mode picks an unnamed character, introduce the character once in plain English — pack name plus what they are ("Inch, the chalk inchworm") — then show the lines and ask which one. Never status-ping with the name alone ("Inch."), say "Still Inch," or ask "which one should <name> draw?" The user hears the result — character, sayings / image, what was made — not the procedure.
生成前的规划需简短具体。生成后,让图片自己说话——报告生成的内容和位置,而非风格理论。标签要少而精;图片中嵌入的文字越少,渲染越可靠。
以实际工作者的语气交流,而非复述本文档内容。聊天中绝不要叙述工作流步骤或行话:如doctor、preflight、provenance、register、saying bar、thesis、backend或“doctor's green”。若需状态更新,使用日常语言,而非步骤罗列。当惊喜模式选择未命名角色时,用通俗易懂的英语介绍一次角色——包名加角色描述(“Inch,粉笔尺蠖”)——然后展示文案并询问选择。绝不要仅用名称进行状态提示(“Inch.”)、说“仍使用Inch”或问“<名称>应该画哪个?”用户关注的是结果——角色、文案/图片、生成的内容——而非流程。