explainer-video

Compare original and translation side by side

🇺🇸

Original

English
🇨🇳

Translation

Chinese

Explainer Video

解释视频

Turn one message into a paced, narrated 30–90s short. Run the full pipeline: script → storyboard → scene build → narration/caption sync → edit → polish, with a consistent style system throughout.
将单一信息转化为节奏适宜、带有旁白(VO)的30-90秒短视频。执行完整流程:脚本 → 分镜 → 场景构建 → 旁白/字幕同步 → 编辑 → 打磨,全程采用统一的风格系统。

When to use

使用场景

  • Produce a 30–90s product, how-it-works, onboarding, or concept video.
  • Go from a raw idea or feature to a script, storyboard, and assembled piece.
  • Sync narration to visuals and add captions for muted playback.
  • 制作30-90秒的产品介绍、工作原理、入门引导或概念讲解视频。
  • 从原始想法或功能点出发,完成脚本、分镜及成品制作。
  • 同步旁白与视觉画面,并添加字幕以适配静音播放场景。

Story structure (decide this first)

故事结构(需先确定)

An explainer is an argument, not a feature tour. Before any pipeline step, lock the narrative so every later choice serves it.
  • One core idea. A single sentence the viewer should be able to repeat afterward. If you can't state it in one line, the video has no spine — cut scope until you can. Everything that doesn't serve that idea gets dropped, not shrunk.
  • Script-first. The VO is the spine; visuals illustrate the line being spoken, never lead it. Write and time the words before you storyboard or animate — it is far cheaper to cut a sentence than a built scene.
  • Earn the "how" with stakes. Don't jump from problem to mechanism. Make the viewer feel the cost of the problem first; that tension is what makes them watch the solution.
The explainer story arc — a beat per stage, in order:
StageNarrative jobWhat the viewer should think
ProblemName the pain in their words"That's me."
StakesShow what the pain costs (time, money, risk)"I need this fixed."
SolutionIntroduce the product/idea as the fix, in one line"Oh — that solves it."
How it worksThe mechanism in 1–3 concrete steps"I get how it does that."
Payoff / CTAThe after-state + one next action"I want that. I'll do X."
(The
Script formula
table below maps these stages to runtime shares and word budgets — this section is the why and order; that one is the how long.)
解释视频是一种论证,而非功能展示。在进入任何流程步骤前,先锁定叙事框架,确保后续所有选择都服务于核心叙事。
  • 单一核心观点:观众看完后能复述的一句话。如果无法用一句话概括,说明视频缺乏核心——需缩减范围直至可以。所有不服务于该观点的内容都应删除,而非压缩。
  • 脚本优先:旁白(VO)是核心骨架;视觉画面用于阐释旁白内容,而非主导叙事。在制作分镜或动画前,先撰写并计时旁白内容——修改一句话的成本远低于修改已完成的场景。
  • 用痛点铺垫「解决方案」:不要直接从问题跳到实现机制。先让观众切实感受到问题带来的代价;这种紧张感才会促使他们关注解决方案。
解释视频的叙事弧——每个阶段对应一个节拍,按顺序推进:
阶段叙事目标观众的心理活动
问题用观众的语言点明痛点「这说的就是我。」
痛点代价展示问题带来的损失(时间、金钱、风险)「我需要解决这个问题。」
解决方案用一句话介绍产品/理念作为解决办法「哦——这能解决问题。」
工作原理用1-3个具体步骤展示实现机制「我明白它是怎么运作的了。」
收益/行动号召(CTA)展示解决后的状态 + 明确的下一步行动「我想要这个。我会做X。」
(下方的「脚本公式」表格将这些阶段对应到时长占比和字数预算——本部分讲的是叙事逻辑与顺序,而脚本公式讲的是时长分配。)

Pick one analogy and ride it

选定一个类比并贯穿始终

Abstract mechanisms land when mapped to something the viewer already knows. Choose a single metaphor and keep it consistent across scenes — switching analogies mid-video resets comprehension.
  • Pick a metaphor from the viewer's world (a queue, a thermostat, an assembly line), not the engineering domain.
  • One metaphor per video; reuse it for both the visual grammar and the VO wording.
  • Test it: if the analogy needs its own explanation, it's the wrong one.
抽象的机制需映射到观众熟悉的事物才能被理解。选择一个单一隐喻,并在所有场景中保持一致——中途切换类比会打断观众的理解。
  • 从观众的日常场景中选择隐喻(如队列、恒温器、装配线),而非技术领域的概念。
  • 每个视频只用一个隐喻;在视觉风格和旁白措辞中重复使用。
  • 测试标准:如果该隐喻需要额外解释,说明选得不合适。

Pacing per beat

每个节拍的节奏

Pace tracks tension. Move quickly through Problem/Stakes to reach the value; slow down on the "how" so each step lands; let the payoff breathe.
BeatFeelCut rhythm
Problem / StakesBrisk, a little tenseFaster cuts, short holds
SolutionA beat of reliefOne clear hold
How it worksDeliberate, one step at a timeSlowest — hold each step to read
Payoff / CTAConfident, openHold the end card; one CTA
节奏与紧张感挂钩。快速推进「问题/痛点代价」部分以尽快触及核心价值;放慢「工作原理」部分的节奏,确保每个步骤都被理解;让「收益」部分有足够的展示空间。
节拍氛围剪辑节奏
问题 / 痛点代价轻快、略带紧张快速剪辑,短时间停留
解决方案放松的节拍清晰的长停留
工作原理从容、逐步推进最慢——每个步骤都停留足够时间让观众阅读
收益 / 行动号召(CTA)自信、开阔停留于结尾画面;单一行动号召

The pipeline

制作流程

  1. Script — one core message. Structure: problem → solution → how → payoff. Write tight voiceover (VO); time it at ~2.3 words/second (≈140 wpm). A 60s video is ~138 words.
  2. Storyboard — one idea per scene. Sketch each frame plus its transition; map each VO line to a visual.
  3. Build scenes — animate diagrams/data, kinetic text/keywords, and backgrounds (techniques inlined below).
  4. Narration & captions — record/generate VO, align scene cuts to VO beats, burn in captions (most plays are muted).
  5. Edit — cut to the narration, hold each idea long enough to land, remove any scene that doesn't advance the message.
  6. Polish — color pass, consistent easing, audio mix, end card.
  1. 脚本——单一核心信息。结构:问题 → 解决方案 → 工作原理 → 收益。撰写简洁的旁白(VO);语速控制在约2.3词/秒(≈140词/分钟)。60秒的视频约为138词。
  2. 分镜——每个场景对应一个核心想法。绘制每一帧及转场效果;将每一句旁白对应到视觉画面。
  3. 场景构建——制作动画图表/数据、动态文本/关键词及背景(下方内嵌技术实现方法)。
  4. 旁白与字幕——录制/生成旁白,将场景剪辑与旁白节拍对齐,添加内嵌字幕(大多数播放场景为静音状态)。
  5. 编辑——根据旁白剪辑视频,每个想法保留足够的展示时间,移除所有不推进核心信息的场景。
  6. 打磨——色彩调整、统一缓动效果、音频混音、结尾画面。

Script formula

脚本公式

BeatJobShare of runtime
ProblemName the pain the viewer feels~20%
SolutionIntroduce the product/idea as the fix~15%
HowShow the mechanism, 1–3 concrete steps~45%
PayoffThe outcome + a clear next step (CTA)~20%
Write VO first, in spoken language (contractions, short sentences). Read it aloud and time it before building anything.
节拍目标时长占比
问题点明观众感受到的痛点~20%
解决方案介绍产品/理念作为解决办法~15%
工作原理展示实现机制,1-3个具体步骤~45%
收益展示成果 + 明确的下一步行动(CTA)~20%
优先撰写旁白,使用口语化表达(缩略语、短句)。在制作任何内容前,先朗读并计时。

VO timing math

旁白计时公式

  • Pace: ~2.3 words/sec (140 wpm) for clear, friendly narration; slow technical lines to ~2.0.
  • Estimate scene length from its VO line:
    seconds = words / 2.3
    , then add ~0.4s breathing room.
  • A scene with no VO (pure visual beat) still needs ≥1.0s to register.
  • 语速:约2.3词/秒(140词/分钟),清晰友好;技术内容可放慢至约2.0词/秒。
  • 根据旁白内容估算场景时长:
    秒数 = 字数 / 2.3
    ,再添加约0.4秒的缓冲时间。
  • 无旁白的场景(纯视觉节拍)仍需至少1.0秒的展示时间。

Scene-build techniques (inlined)

场景构建技术(内嵌)

Diagram / data reveal (progressive disclosure: nodes → edges → labels):
js
import gsap from "gsap";
gsap.set(".edge", { strokeDasharray: 1, strokeDashoffset: 1 });
gsap.timeline()
  .from(".node", { opacity: 0, scale: .85, transformOrigin: "center",
                   duration: .45, stagger: .3, ease: "back.out(1.5)" })
  .to(".edge",  { strokeDashoffset: 0, duration: .5, stagger: .3 }, "-=0.6");
Kinetic keyword (word pops in on the stressed syllable):
css
.kw { display:inline-block; opacity:0; transform:translateY(18px); }
.kw.in { opacity:1; transform:none; transition:all .4s cubic-bezier(.22,1,.36,1); }
Count-up stat:
js
function countUp(el, to, dur=1200){ const t0=performance.now(); (function f(n){
  const k=Math.min(1,(n-t0)/dur), e=1-Math.pow(1-k,3);
  el.textContent=Math.round(to*e).toLocaleString(); if(k<1)requestAnimationFrame(f);})(t0); }
Render the assembled piece with Remotion (data-driven, deterministic) or After Effects.
图表/数据逐步展示(渐进式披露:节点 → 连线 → 标签):
js
import gsap from "gsap";
gsap.set(".edge", { strokeDasharray: 1, strokeDashoffset: 1 });
gsap.timeline()
  .from(".node", { opacity: 0, scale: .85, transformOrigin: "center",
                   duration: .45, stagger: .3, ease: "back.out(1.5)" })
  .to(".edge",  { strokeDashoffset: 0, duration: .5, stagger: .3 }, "-=0.6");
动态关键词(重读音节时弹出):
css
.kw { display:inline-block; opacity:0; transform:translateY(18px); }
.kw.in { opacity:1; transform:none; transition:all .4s cubic-bezier(.22,1,.36,1); }
数字递增动画:
js
function countUp(el, to, dur=1200){ const t0=performance.now(); (function f(n){
  const k=Math.min(1,(n-t0)/dur), e=1-Math.pow(1-k,3);
  el.textContent=Math.round(to*e).toLocaleString(); if(k<1)requestAnimationFrame(f);})(t0); }
使用Remotion(数据驱动、确定性)或After Effects渲染成品。

Caption sync

字幕同步

Captions are mandatory — assume muted autoplay. Use a
.srt
/
.vtt
cue list keyed to the VO. Keep ≤2 lines, ≤42 chars/line, on-screen ≥1.0s, gone before the next line starts.
1
00:00:00,300 --> 00:00:02,600
Managing releases by hand is slow.

2
00:00:02,800 --> 00:00:05,400
Shipfast automates the whole pipeline.
In Remotion, drive captions from the same timing array used for scene cuts so they never drift.
字幕是必需的——默认假设为静音自动播放。使用与旁白关联的
.srt
/
.vtt
提示列表。字幕最多2行,每行≤42字符,屏幕停留时间≥1.0秒,下一行出现前消失。
1
00:00:00,300 --> 00:00:02,600
Managing releases by hand is slow.

2
00:00:02,800 --> 00:00:05,400
Shipfast automates the whole pipeline.
在Remotion中,使用与场景剪辑相同的计时数组驱动字幕,确保不会出现不同步。

Style-system consistency

风格系统一致性

Lock one type scale, one color grammar (color = meaning, never reassigned), and one easing curve before building scene 2. Reuse the same enter/exit transitions across scenes so the piece feels like one object, not a montage of experiments. Visuals must illustrate the VO line currently spoken — show, don't decorate.
在制作第二个场景前,锁定一套字体层级、一套色彩规则(色彩代表特定含义,不得随意更改)和一套缓动曲线。在所有场景中重复使用相同的入场/退场转场效果,让成品看起来是一个整体,而非实验片段的拼接。视觉画面必须阐释当前播放的旁白内容——展示内容,而非装饰。

Output checklist

输出检查清单

  • Single clear message; every scene advances it.
  • VO timed (~2.3 w/s); scene lengths derived from VO.
  • Storyboard maps each VO line to a visual.
  • Captions burned/available; ≤2 lines, readable.
  • Consistent type/color/motion system across all scenes.
  • End card with one clear CTA.
  • 单一清晰的核心信息;每个场景都推进该信息。
  • 旁白计时准确(约2.3词/秒);场景时长基于旁白内容确定。
  • 分镜将每一句旁白对应到视觉画面。
  • 字幕已内嵌/可获取;最多2行,清晰可读。
  • 所有场景采用统一的字体/色彩/动效系统。
  • 结尾画面带有明确的行动号召(CTA)。

Deliver & verify (rendered stills → MP4)

交付与验证(渲染静帧 → MP4)

Packaged helper (
scripts/
): tile your stills with
scripts/contact-sheet.sh sheet.png f-hook.png f-mid.png f-end.png
, then assert the encode with
scripts/probe-mp4.sh out.mp4 [WxH] [fps]
. See
scripts/README.md
.
The assembled explainer is a Remotion composition — frame-deterministic, so any exact frame renders headlessly with no seek harness. Use this tier when the deliverable is an MP4/GIF that carries baked narration and burned captions; for a single web scene, deliver standalone HTML instead.
Output contract:
  • A Remotion project with the composition registered (
    <Composition>
    + zod
    schema
    +
    defaultProps
    ), all motion frame-driven (no timers /
    Date.now()
    /
    Math.random()
    ).
  • Deliverable = the rendered
    out/*.mp4
    (plus the project, so VO/script/data can be re-rendered).
  • Narration via
    <Audio src={staticFile()}>
    ; VO/scene timing baked into props (no realtime audio at render). Captions driven from that same timing array so they never drift.
  • Duration data-dependent? compute it in
    calculateMetadata
    , not by hand.
Verify loop — render stills → inspect → encode. Render single frames first (cheap, no encode), inspect them, encode only once the frames are right.
bash
undefined
打包辅助工具
scripts/
目录):使用
scripts/contact-sheet.sh sheet.png f-hook.png f-mid.png f-end.png
拼接静帧,然后使用
scripts/probe-mp4.sh out.mp4 [WxH] [fps]
验证编码。详见
scripts/README.md
成品解释视频是一个Remotion合成项目——帧确定性,因此任何精确帧都可无头渲染,无需搜索工具。当交付物为带有内嵌旁白和字幕的MP4/GIF时使用此层级;若为单个网页场景,则交付独立HTML文件。
输出规范:
  • 一个已注册合成的Remotion项目(
    <Composition>
    + zod
    schema
    +
    defaultProps
    ),所有动效由帧驱动(无计时器 /
    Date.now()
    /
    Math.random()
    )。
  • 交付物 = 渲染后的
    out/*.mp4
    (附带项目文件,以便重新渲染旁白/脚本/数据)。
  • 旁白通过
    <Audio src={staticFile()}>
    实现;旁白/场景计时内嵌到props中(渲染时无实时音频)。字幕由同一计时数组驱动,确保不会不同步。
  • 时长依赖数据?在
    calculateMetadata
    中计算,而非手动设置。
验证流程——渲染静帧 → 检查 → 编码。 先渲染单帧(成本低,无需编码),检查无误后再编码:
bash
undefined

Frame-exact stills at start / mid / end — render with the SHIPPED props, not just defaults

使用最终props(而非仅默认props)渲染开头/中间/结尾的精确静帧

npx remotion still Explainer out/f-start.png --frame=0 --props='{...}' npx remotion still Explainer out/f-mid.png --frame=N --props='{...}' npx remotion still Explainer out/f-end.png --frame=L --props='{...}' # L = durationInFrames - 1
npx remotion still Explainer out/f-start.png --frame=0 --props='{...}' npx remotion still Explainer out/f-mid.png --frame=N --props='{...}' npx remotion still Explainer out/f-end.png --frame=L --props='{...}' # L = durationInFrames - 1

Inspect: script→storyboard→scene maps onto frames — sample a frame inside EACH scene's hold,

检查:脚本→分镜→场景是否与帧对应——在每个场景的停留帧中采样,确认旁白内容和字幕与画面同步,无文本溢出/超出画布。

confirm narration line and caption are synced to the visual at that frame, no text overflow / off-canvas.

仅在静帧检查通过后,进行编码:

Only after the stills check out, encode:

npx remotion render Explainer out/explainer.mp4 --props='{...}'

- `npx remotion compositions` reads `durationInFrames`/`fps` to pick the end frame and per-scene hold frames.
- **README demo GIF for free**: `npx remotion render Explainer out/demo.gif --codec=gif`.

**Before you finish:**
1. `npx remotion still` renders cleanly at frame 0, a frame inside each scene's hold, and last — no errors, no missing assets/fonts.
2. At every sampled scene frame, the caption and the spoken VO line match the visual on screen (no drift).
3. Captions are inside safe areas, ≤2 lines, readable; text is exact and not clipped.
4. Frame-driven only — no `Date.now()` / `Math.random()` / timers; the **shipped** props render correctly (not just `defaultProps`).
5. Full MP4 encoded and plays with narration; (optional) GIF rendered for the README.
npx remotion render Explainer out/explainer.mp4 --props='{...}'

- `npx remotion compositions`读取`durationInFrames`/`fps`以选择结束帧和每个场景的停留帧。
- **免费生成README演示GIF**:`npx remotion render Explainer out/demo.gif --codec=gif`。

**完成前检查:**
1. `npx remotion still`可成功渲染第0帧、每个场景停留帧及最后一帧——无错误,无缺失资源/字体。
2. 在所有采样场景帧中,字幕和旁白内容与画面匹配(无不同步)。
3. 字幕位于安全区域内,最多2行,清晰可读;文本准确无截断。
4. 仅由帧驱动——无`Date.now()` / `Math.random()` / 计时器;最终props可正确渲染(而非仅`defaultProps`)。
5. 完整MP4已编码并可播放旁白;(可选)已为README渲染GIF。

Reference files

参考文件

  • references/script-to-screen-workflow.md
    — full pipeline detail: the problem→solution→how→payoff script formula with a worked example and word budget, VO timing tables, a fill-in storyboard template with timecodes, a per-scene build checklist, caption authoring guidance (VTT/SRT), and a style-system spec sheet for cross-scene consistency.
  • references/script-to-screen-workflow.md
    ——完整流程细节:包含问题→解决方案→工作原理→收益的脚本公式(附实例和字数预算)、旁白计时表、带时间码的可填写分镜模板、逐场景构建检查清单、字幕制作指南(VTT/SRT)、跨场景一致性的风格系统规范表。