synthetic-screen-recording
Compare original and translation side by side
🇺🇸
Original
English🇨🇳
Translation
ChineseSynthetic Screen Recording (Remotion TerminalScene)
合成屏幕录制(Remotion TerminalScene)
Decision this skill answers: When the user wants a screen-recording-looking demo of a terminal, CLI tool, or coding workflow — do I capture the real desktop (OS screen recording via , Windows-MCP, Cap, or Playwright), or do I synthesize it in Remotion with the component?
screen_recorderTerminalSceneHeuristic: If the agent can author the exact command/output sequence in advance, synthesize. Only capture live when the real behavior is unpredictable, needs a real app UI, or the user explicitly asked for a real recording.
本技能解决的决策问题: 当用户需要一个看起来像屏幕录制的终端、CLI工具或编码流程演示时,我应该捕获真实桌面(通过、Windows-MCP、Cap或Playwright进行系统屏幕录制),还是使用Remotion的组件合成录制内容?
screen_recorderTerminalScene判断准则: 如果Agent可以预先编写好精确的命令/输出序列,就选择合成录制。只有当真实行为不可预测、需要真实应用UI,或者用户明确要求真实录制时,才进行实时捕获。
Why this exists
为什么会有这个技能
v3 of the OpenMontage showcase tried to use Windows-MCP + to drive a Git-Bash window for the install walkthrough. It stalled on window positioning, focus races, and taskbar privacy concerns. We pivoted to pure Remotion rendering — a React component named that draws a fake terminal and types commands character-by-character. The output is visually indistinguishable from a real screen recording (same traffic-light window chrome, blinking cursor, scrolling output) but deterministic, privacy-safe, pixel-perfect at 1080p, and pace-controllable to the frame.
screen_recorderTerminalSceneThat component + pattern is the capability this skill makes discoverable.
OpenMontage展示项目的v3版本曾尝试使用Windows-MCP + 来驱动Git-Bash窗口完成安装流程演示,但在窗口定位、焦点竞争和任务栏隐私问题上遇到了阻碍。于是我们转向纯Remotion渲染——一个名为的React组件,它可以绘制一个模拟终端,并逐字符输入命令。输出效果与真实屏幕录制视觉上几乎无差别(拥有相同的红绿灯窗口边框、闪烁光标、滚动输出),但具有确定性、隐私安全、1080p像素级完美,并且可以精确控制到每一帧的节奏。
screen_recorderTerminalScene本技能的作用就是让大家了解这个组件和对应的使用模式。
When to use synthetic (TerminalScene)
何时使用合成录制(TerminalScene)
YES, synthesize when:
- The demo is a terminal / CLI / coding session where commands and outputs are predictable
- The user wants a polished tutorial feel (clean typography, floating pills, cursor blink)
- Install walkthroughs, setup demos, API key config, targets,
makeflowsgit clone - You need tight sync with narration — every command must land on a specific beat
- You want the result reproducible (re-render gets identical pixels)
- The user's actual desktop has private apps/windows visible you'd otherwise have to crop
NO, capture a real screen when:
- The demo is a real app UI that can't be faked (Figma, Photoshop, a web app with live state, a browser flow)
- The user explicitly asked for a recording of their actual screen
- The behavior depends on timing you can't script (streaming LLM output, real network latency)
- There's a visual quirk (a cursor effect, a plugin pop-up) that only appears in the live environment
For a browser demo → skill, not this one.
For a real desktop → tool or Cap via .
playwright-recordingscreen_recordercap_recorder适合使用合成录制的场景:
- 演示内容是终端/CLI/编码会话,且命令和输出可预测
- 用户想要精致的教程风格(清晰排版、浮动提示框、光标闪烁)
- 安装流程演示、设置演示、API密钥配置、目标、
make流程git clone - 需要与旁白严格同步——每个命令必须在特定节拍出现
- 希望结果可重现(重新渲染得到完全相同的像素)
- 用户的实际桌面有隐私应用/窗口,否则需要裁剪处理
适合捕获真实屏幕的场景:
- 演示内容是无法模拟的真实应用UI(Figma、Photoshop、带有实时状态的Web应用、浏览器流程)
- 用户明确要求录制他们的实际屏幕
- 行为依赖无法脚本化的时序(流式LLM输出、真实网络延迟)
- 存在仅在实时环境中才会出现的视觉特性(光标效果、插件弹窗)
浏览器演示 → 使用技能,而非本技能。
真实桌面录制 → 使用工具或通过调用Cap。
playwright-recordingscreen_recordercap_recorderThe component — TerminalScene
TerminalScene组件——TerminalScene
TerminalSceneLocated at:
Exported from:
Wired in dispatch: ()
remotion-composer/src/components/TerminalScene.tsxremotion-composer/src/components/index.tsremotion-composer/src/Explainer.tsxif (cut.type === "terminal_scene")Props:
ts
interface TerminalSceneProps {
title?: string; // shown in the window title bar
steps: TerminalStep[]; // the timeline
prompt?: string; // "$", ">", etc.
accentColor?: string; // pill + prompt glow
backgroundColor?: string;
}Step kinds:
ts
{ kind: "cmd", text: string, typeSpeed?: number, holdSeconds?: number }
{ kind: "out", text: string, holdSeconds?: number }
{ kind: "pause", seconds: number }
{ kind: "pill", text: string, color?: string, durationSeconds?: number }- — prints the prompt, types the text character-by-character (
cmdis seconds per character, default 0.035), then holds fortypeSpeed(default 0.3)holdSeconds - — a line of program output, reveals instantly with a short fade-in
out - — dead time. Terminal holds on last visible state. USE THIS TO SYNC WITH NARRATION.
pause - — non-blocking floating badge (top-right). Spring-in, hold, spring-out. Does NOT advance the cursor — the next step runs in parallel.
pill
位置:
导出位置:
调度位置:(处)
remotion-composer/src/components/TerminalScene.tsxremotion-composer/src/components/index.tsremotion-composer/src/Explainer.tsxif (cut.type === "terminal_scene")Props:
ts
interface TerminalSceneProps {
title?: string; // 显示在窗口标题栏中
steps: TerminalStep[]; // 时间线
prompt?: string; // "$", ">", 等
accentColor?: string; // 提示框 + 提示符高亮色
backgroundColor?: string;
}步骤类型:
ts
{ kind: "cmd", text: string, typeSpeed?: number, holdSeconds?: number }
{ kind: "out", text: string, holdSeconds?: number }
{ kind: "pause", seconds: number }
{ kind: "pill", text: string, color?: string, durationSeconds?: number }- — 显示提示符,逐字符输入文本(
cmd为每个字符的输入时长,默认0.035秒),然后保持typeSpeed时长(默认0.3秒)holdSeconds - — 一行程序输出,通过短时间淡入立即显示
out - — 停顿时间。终端保持最后可见状态。使用此步骤与旁白同步。
pause - — 非阻塞式浮动徽章(右上角)。弹入、保持、弹出。不会推进光标——下一步会并行运行。
pill
Authoring pattern
创作模式
Author a new scene by adding a cut to (or your equivalent props builder):
build_composition.pypython
install_steps = [
{"kind": "pause", "seconds": 7.0}, # wait for intro narration
{"kind": "cmd", "text": "git clone https://github.com/calesthio/OpenMontage.git",
"typeSpeed": 0.045, "holdSeconds": 0.3},
{"kind": "out", "text": "Cloning into 'OpenMontage'..."},
{"kind": "out", "text": "remote: Enumerating objects: 2847, done."},
{"kind": "pill", "text": "repo cloned", "color": "#34D399", "durationSeconds": 2.6},
{"kind": "pause", "seconds": 3.8}, # bridge to next narration cue
# ...
]
cuts.append({
"id": "install-terminal",
"type": "terminal_scene",
"terminalTitle": "bash — OpenMontage setup",
"prompt": "$",
"accentColor": "#22D3EE",
"steps": install_steps,
"in_seconds": 50.0,
"out_seconds": 110.0,
})通过向(或等效的Props构建器)添加一个片段来创作新场景:
build_composition.pypython
install_steps = [
{"kind": "pause", "seconds": 7.0}, // 等待开场旁白
{"kind": "cmd", "text": "git clone https://github.com/calesthio/OpenMontage.git",
"typeSpeed": 0.045, "holdSeconds": 0.3},
{"kind": "out", "text": "Cloning into 'OpenMontage'..."},
{"kind": "out", "text": "remote: Enumerating objects: 2847, done."},
{"kind": "pill", "text": "repo cloned", "color": "#34D399", "durationSeconds": 2.6},
{"kind": "pause", "seconds": 3.8}, // 过渡到下一个旁白提示点
// ...
]
cuts.append({
"id": "install-terminal",
"type": "terminal_scene",
"terminalTitle": "bash — OpenMontage setup",
"prompt": "$",
"accentColor": "#22D3EE",
"steps": install_steps,
"in_seconds": 50.0,
"out_seconds": 110.0,
})THE RULE: pace with narration, never ahead
规则:与旁白同步节奏,绝不超前
The #1 failure mode: steps run continuously and burn through all content in the first 40% of the scene, leaving the terminal frozen for the remaining 60%. This is what killed the v3 first pass — the capability menu rendered at t=80s but narration didn't announce it until t=92s.
Do this instead:
- Know your narration cues — for each scene, write down the exact video-time each narration segment starts.
- Start with a pause that reaches the first narration cue before any command types.
- Time each command to land with its narration line — should start typing the moment narration says its line, not before.
cmd - Put pauses between command groups that bridge to the next narration cue.
- End with a closer hold — a pause long enough that the final state is readable after narration ends.
Sanity-check your steps before rendering — every minute of Remotion render is precious. Sum the step durations and verify they equal scene duration:
python
import math
def trace(steps, scene_start, fps=30):
t = 0.0
for s in steps:
k = s["kind"]
if k == "cmd":
tf = math.ceil(len(s["text"]) * s.get("typeSpeed", 0.035) * fps)
t += tf / fps + s.get("holdSeconds", 0.3)
elif k == "out":
t += max(2, math.ceil(0.08 * fps)) / fps + s.get("holdSeconds", 0.15)
elif k == "pause":
t += s["seconds"]
# "pill" is non-blocking — does NOT advance cursor
print(f" {t + scene_start:6.2f}s {k}: {s.get('text', '')[:40]}")
trace(install_steps, 50)Look at the output column. Each narration cue's video-time must appear adjacent to the command/output it announces. If a command lands 10s before or after its cue, adjust pauses.
See for a reusable version of this script.
lib/verify_scene_pacing.py最常见的失败模式: 步骤持续运行,在场景前40%的时间内就完成所有内容,导致终端在剩余60%的时间内处于冻结状态。这正是v3版本首次尝试失败的原因——功能菜单在第80秒渲染完成,但旁白直到第92秒才提及它。
正确做法:
- 了解旁白提示点 — 针对每个场景,写下每个旁白片段开始的精确视频时间。
- 以停顿开头 — 在输入任何命令前,先添加一个停顿,直到第一个旁白提示点出现。
- 让每个命令与对应旁白同步 — 应该在旁白提及该命令的瞬间开始输入,而不是提前。
cmd - 在命令组之间添加停顿 — 过渡到下一个旁白提示点。
- 结尾保持停顿 — 添加足够长的停顿,确保旁白结束后最终状态仍可阅读。
渲染前检查步骤节奏 — Remotion的每一分钟渲染都很宝贵。计算所有步骤的总时长,验证是否等于场景时长:
python
import math
def trace(steps, scene_start, fps=30):
t = 0.0
for s in steps:
k = s["kind"]
if k == "cmd":
tf = math.ceil(len(s["text"]) * s.get("typeSpeed", 0.035) * fps)
t += tf / fps + s.get("holdSeconds", 0.3)
elif k == "out":
t += max(2, math.ceil(0.08 * fps)) / fps + s.get("holdSeconds", 0.15)
elif k == "pause":
t += s["seconds"]
# "pill" 是非阻塞的——不会推进光标
print(f" {t + scene_start:6.2f}s {k}: {s.get('text', '')[:40]}")
trace(install_steps, 50)查看输出列。每个旁白提示点的视频时间必须与它所介绍的命令/输出相邻。如果命令比提示点早或晚10秒出现,请调整停顿时长。
可复用的脚本版本请查看。
lib/verify_scene_pacing.pyDesign rules (inherited from the v3 retune)
设计规则(继承自v3版本调整)
- Intro pause — every terminal scene opens with at least 2s of empty-terminal-with-blinking-cursor before anything types. The viewer needs to register the window.
- Pill timing — a pill should fire at the exact moment its named event completes on screen (e.g., immediately after the last
repo clonedline). Pills are your substitute for real-world UI notifications.Receiving objects - Command hold after typing — keep ≥ 0.3 on every
holdSecondsso viewers register the completed command before the first output scrolls in.cmd - Output cadence — space on output lines between 0.4 and 1.0. Output that flies too fast feels like a bug; output that crawls feels boring.
holdSeconds - Auto-scroll works — the terminal holds the most recent 18 lines. Don't worry about off-screen content.
- Cursor blinks only on the latest command line while typing + a ~0.2s tail after typing completes.
- 开场停顿 — 每个终端场景在开始输入内容前,至少要有2秒的空终端+闪烁光标时间。观众需要时间识别窗口。
- 提示框时机 — 提示框应在屏幕上对应事件完成的瞬间弹出(例如,应在最后一行
repo cloned输出后立即出现)。提示框是真实世界UI通知的替代方案。Receiving objects - 命令输入后的保持时长 — 每个的
cmd应≥0.3秒,以便观众在第一行输出滚动前识别完整命令。holdSeconds - 输出节奏 — 输出行的应设置在0.4到1.0之间。输出过快会感觉像bug;过慢则会显得乏味。
holdSeconds - 自动滚动生效 — 终端保留最近的18行内容。无需担心屏幕外的内容。
- 光标仅在输入最新命令行时闪烁 + 输入完成后约0.2秒的延迟
ProviderChip
(companion component)
ProviderChipProviderChip
(配套组件)
ProviderChipThe pattern also owns — a rotating badge overlay that cycles through a list of provider names at a fixed cadence. Used in the v3 showcase to cycle through all 11 AI video-gen providers during the "generated motion" section.
.agents/skills/synthetic-screen-recordingProviderChippython
overlays.append({
"type": "provider_chip",
"providers": ["Veo 3.1", "Seedance 2.0", "Kling 2.5", ...],
"cycleSeconds": 2.5,
"position": "bottom-right",
"accentColor": "#22D3EE",
"label": "generated with",
"in_seconds": 195.0,
"out_seconds": 222.5,
})Wired in dispatch at: overlay renderer ().
remotion-composer/src/Explainer.tsxoverlay.type === "provider_chip".agents/skills/synthetic-screen-recordingProviderChippython
overlays.append({
"type": "provider_chip",
"providers": ["Veo 3.1", "Seedance 2.0", "Kling 2.5", ...],
"cycleSeconds": 2.5,
"position": "bottom-right",
"accentColor": "#22D3EE",
"label": "generated with",
"in_seconds": 195.0,
"out_seconds": 222.5,
})调度位置:的覆盖层渲染器(处)。
remotion-composer/src/Explainer.tsxoverlay.type === "provider_chip"Adding new synthetic-UI components
添加新的合成UI组件
The pattern generalizes. When you need to fake another UI surface (Claude Code chat bubbles, a Jira ticket view, a GitHub PR diff, a Slack message, a VS Code status bar):
- Copy as a template.
TerminalScene.tsx - Define a interface for the relevant timeline primitives.
steps - Render each step by interpolating against cumulative start/end times.
frame - Wire it into 's
Explainer.tsxdispatch with a newSceneRenderer.cut.type - Add the type to the interface in
Cutand toExplainer.tsx.components/index.ts - Add a section to this skill documenting it.
- Update with the new cut type.
remotion-composer/SCENE_TYPES.md
该模式可以推广。当你需要模拟其他UI界面(Claude Code聊天气泡、Jira工单视图、GitHub PR差异、Slack消息、VS Code状态栏)时:
- 复制作为模板。
TerminalScene.tsx - 为相关时间线原语定义接口。
steps - 通过将与累积开始/结束时间插值来渲染每个步骤。
frame - 在的
Explainer.tsx调度中添加新的SceneRenderer,将其接入。cut.type - 在的
Explainer.tsx接口和Cut中添加该类型。components/index.ts - 在本技能中添加一个部分对其进行文档说明。
- 更新,添加新的片段类型。
remotion-composer/SCENE_TYPES.md
Related skills
相关技能
- — general Remotion authoring (hooks, springs, sequences)
.agents/skills/remotion - — real browser-flow capture for web apps
.agents/skills/playwright-recording - — ffmpeg-based desktop capture
tools/capture/screen_recorder - — Cap.so polished desktop capture
tools/capture/cap_recorder - — chooses between synthetic and real for a screen-demo project
skills/pipelines/screen-demo/asset-director.md
- — 通用Remotion创作(hooks、弹簧动画、序列)
.agents/skills/remotion - — 针对Web应用的真实浏览器流程捕获
.agents/skills/playwright-recording - — 基于ffmpeg的桌面捕获
tools/capture/screen_recorder - — Cap.so的精致桌面捕获
tools/capture/cap_recorder - — 为屏幕演示项目选择合成或真实录制
skills/pipelines/screen-demo/asset-director.md
Provenance
来源
Introduced: OpenMontage showcase v3 render (2026-04-16). Original motivation: the v3 setup walkthrough section needed a 60-second install demo where every command aligned to Chirp 3 HD narration cues, and Windows-MCP-driven real capture was too flaky in practice. See for the reference implementation.
projects/openmontage-showcase/build_composition.py引入时间:OpenMontage展示项目v3版本渲染(2026-04-16)。最初动机:v3版本的安装流程演示部分需要一个60秒的安装演示,其中每个命令都要与Chirp 3 HD旁白提示点对齐,而Windows-MCP驱动的真实捕获在实践中过于不稳定。参考实现请查看。
projects/openmontage-showcase/build_composition.py