niwo-render-everything
Compare original and translation side by side
🇺🇸
Original
English🇨🇳
Translation
ChineseNiwo 短视频素材包
Niwo Short-Video Asset Bundle
把和用户聊过的内容,整理成一个可以直接上传到 Niwo 渲染成片的「素材包」。
Turn the content discussed with the user into an "asset bundle" that can be directly uploaded to Niwo for rendering into a video.
先确认你能不能做
First Confirm If You Can Do This Task
这个任务需要以下三项能力,缺任何一项都要立刻告诉用户,不要硬着头皮往下做:
- 能联网访问网页并下载图片、视频文件到本地(不是只给链接)。
- 能执行 shell 命令,环境里有 和
ffmpeg。ffprobe - 能把本地目录打包成 zip 并把文件交付给用户。
如果你只能上网搜索、不能下载文件(例如大多数手机端 AI 助手),直接说明,让用户换一个工具。
This task requires the following three capabilities; if any one is missing, inform the user immediately and do not proceed forcefully:
- Ability to access web pages online and download images and video files to local disk (not just provide links).
- Ability to execute shell commands, with and
ffmpegavailable in the environment.ffprobe - Ability to package a local directory into a zip file and deliver it to the user.
If you can only search online but cannot download files (e.g., most mobile AI assistants), explain this directly and ask the user to switch to another tool.
工作流
Workflow
- 先按「第零步」把成片形态和文案要求一次问完,拿到用户回答后再往下走。
- 定稿口播文案,竖屏信息版连钩子标题一起定稿。
- 下载图片与 B-roll 视频。
- 给每段视频定好用哪几秒。
- 写 前先读
manifest.json。references/manifest-schema.md - 写 前先读
content.json。references/content-schema.md - 跑 自校验,通过后再打包。
scripts/validate_bundle.py
- First, ask all questions about the video format and script requirements in one go as per "Step Zero". Only proceed after receiving the user's answers.
- Finalize the narration script; for vertical info-style videos, finalize the hook headline along with the script.
- Download images and B-roll footage.
- Determine which seconds of each video clip to use.
- Read before writing
references/manifest-schema.md.manifest.json - Read before writing
references/content-schema.md.content.json - Run for self-check; only package after passing the validation.
scripts/validate_bundle.py
交付物
Deliverables
一个 zip,解压后结构如下:
content.json
manifest.json
images/
01-xxx.jpg
02-xxx.png
videos/
01-xxx.mp4images/videos/manifest.jsonfile你只负责内容:文案、素材、多音字读音、钩子标题。成片形态、音色、语速、字幕开关、IP 形象、背景音乐这些渲染参数由用户在 Niwo 里渲染前自己在界面上选,不要写进 。
content.jsonA zip file with the following structure after extraction:
content.json
manifest.json
images/
01-xxx.jpg
02-xxx.png
videos/
01-xxx.mp4The filenames in and can be arbitrary, but must match the field in .
images/videos/filemanifest.jsonYou are only responsible for content: script, assets, pronunciation of polyphonic characters, hook headline. Rendering parameters such as video format, voice tone, speaking speed, subtitle toggle, IP character, and background music are selected by the user on Niwo's rendering interface before rendering; do not write them into .
content.json第零步:开工前先确认几件事
Step Zero: Confirm These Matters Before Starting
动手找素材、写文案之前,先把下面该问的问清楚。不要默默替用户决定。
只有用户在这次对话里明确说过的,才算答案。跨会话记忆、以前几支片子里的选择、「用户一般偏好」这类推断,一概不算,哪怕你很确定也不算。这些只能作为选项里的推荐项递给用户确认,不能直接当结论用。
具体说,不要出现这种开场:「默认按你常用的 90 秒、面向科技商业用户、强钩子来写」然后直接甩出一版草稿。你可以把 90 秒放在选项第一个并标「推荐」,但按哪个走要用户点。
用户明确说「你定」之后(包括点了选项里那条「我没特殊要求,你帮我定」),你再自己拍板。凡是用了兜底默认值,都要在回复里明确列出用了哪几个,不要不声不响地用掉。
Before searching for assets or writing the script, clarify all the following questions. Do not make decisions on behalf of the user silently.
Only what the user explicitly states in this conversation counts as an answer. Cross-session memories, choices from previous videos, or inferences like "user's general preferences" do not count, even if you are certain. These can only be presented as recommended options for the user to confirm, not directly taken as conclusions.
Specifically, avoid opening with something like: "By default, I'll write it according to your usual 90-second format, targeting tech and business users with a strong hook" and then directly present a draft. You can place the 90-second option first and mark it as "Recommended", but the final choice must be confirmed by the user.
Only after the user explicitly says "You decide" (including selecting the option "I have no special requirements, please decide for me") can you make decisions on your own. Whenever a fallback default value is used, clearly list which values were used in your response; do not use them silently.
问什么:五个独立维度,同一轮问完
What to Ask: Five Independent Dimensions, All Asked in One Round
文案还没定稿时,成片形态、目标时长、受众、风格口吻、必讲信息点必须在同一轮一次性问完,不要先问形态、拿到回答再问文案。文案已经定稿时,这一轮只问成片形态。这一轮拿到用户回答之前,不要开始写文案草稿,也不要开始搜索下载素材。
这五项是五个独立维度,必须一项一个问题分开问,绝对不许为了凑工具的问题数量上限把它们合并成组合选项。 比如不要问「这期采用什么成片形态和目标时长?」,选项写成「竖屏信息版·90 秒」——那等于把两个独立变量强行绑死,用户想选「竖屏信息版 + 60 秒」时无从下手,只能手填。同理不许把受众和风格口吻捏成一题。
如果你用的组件确实装不下五题,就在同一轮回复里连着放多个组件分批呈现,五个问题该长什么样还是什么样,不要跨轮追问,也不要靠合并来省题数。
选项别只给光秃秃的名词。 「横屏视频」这种标签,不熟术语的人看不出选它会得到什么。每个选项后面用破折号带一句短说明,讲清它长什么样、适合投在哪,比如「横屏视频 — 16:9 满屏,适合 B 站、YouTube 和电脑上看」。说明控制在 20 个中文字内,一句话说完;组件只接受纯短标签、放不下说明时,就把这几句写在组件前面的正文里。
When the script is not finalized, the video format, target duration, audience, style/tone, and mandatory information points must all be asked in one round at once. Do not ask about the format first, then ask about the script after receiving the answer. When the script is already finalized, only ask about the video format in this round. Do not start writing the script draft or searching/downloading assets until you receive the user's answers in this round.
These five are independent dimensions; each must be asked as a separate question. Never combine them into combined options to meet the tool's question limit. For example, do not ask "What video format and target duration will be used for this episode?" with options like "Vertical Info-Style · 90 Seconds" — this forces two independent variables together, leaving the user unable to choose "Vertical Info-Style + 60 Seconds" except by typing it manually. Similarly, do not combine audience and style/tone into one question.
If the component you are using cannot fit five questions, present multiple components in the same response in batches. Keep each question as it is; do not split the questions across rounds or merge them to reduce the number of questions.
Do not provide bare options only. A label like "Horizontal Video" is unclear to users unfamiliar with the terminology. Add a short explanation after each option with an em dash, explaining what it looks like and where it is suitable for posting, e.g., "Horizontal Video — 16:9 full screen, suitable for Bilibili, YouTube, and viewing on computers". Keep the explanation within 20 Chinese characters (or equivalent in English) in one sentence; if the component only accepts short labels and cannot fit the explanation, write these sentences in the main text before the component.
开工那一轮的问题模板
Question Template for the Kickoff Round
照这个结构出五题,每题都留自由填写入口。下面的选项拿 DeepSeek 涨价这个选题举例,实际按当期选题改写,第 5 题的选项必须重写,不要照抄:
- 这期采用哪种成片形态?(单选)— 三个选项照下面「必问:成片形态」那一节的写法出,别在这里另起一套
- 目标时长是多少?(单选)
- 90 秒(推荐)— 约 540 字,能把涨幅和原因讲完整
- 60 秒 — 约 360 字,只保留核心判断
- 30 秒 — 约 180 字,突出单一爆点
- 主要讲给谁看?(单选)
- 泛科技与商业用户(推荐)— 重点讲价格变化和商业逻辑
- AI 开发者 — 重点讲 API 成本和模型选择
- 普通用户 — 重点解释是否影响自己日常使用
- 希望什么风格口吻?(单选)
- 强钩子锐评(推荐)— 观点鲜明,开场就给冲突
- 客观财经分析 — 数据完整,判断克制
- 实用解读 — 聚焦对读者的具体影响和应对
- 哪些信息点必须讲到?(多选)
- 具体数字:涨幅、时段、最新价格
- 普通用户是否受影响
- 为什么现在做这个决定
- 对从业者成本的真实影响
- 涨价后还有没有性价比
- 我没特殊要求,你帮我定
第 5 题最后固定放一条**「我没特殊要求,你帮我定」**。这题的选项是你按选题现编的,用户未必对讲哪几点有想法,得给他一键放权的出口,别逼他在一堆都行的信息点里硬勾。前四题不加这条:形态、时长、受众、口吻都有明确的推荐项,用户不点就走兜底,本来就不会卡住。
用户勾了这条,就等于他明确说了「必讲信息点你定」:你按文案需要自己挑,然后在回复里写明这项是你定的、定了哪几点。如果他既勾了这条、又勾了别的具体项,按具体项走,忽略这一条。放权只限这一项,别顺手把其他四项也一起替他决定了。
Follow this structure to create five questions, each with a free-text input option. The example below uses the topic of DeepSeek's price increase; rewrite the options according to the current topic, the options for Question 5 must be rewritten, do not copy them directly:
- What video format will be used for this episode? (Single selection) — Use the three options as written in the "Mandatory Question: Video Format" section below, do not create a new set here
- What is the target duration? (Single selection)
- 90 seconds (Recommended) — Approximately 540 words, enough to fully explain the price increase and reasons
- 60 seconds — Approximately 360 words, only core judgments are retained
- 30 seconds — Approximately 180 words, highlighting a single explosive point
- Who is the main audience? (Single selection)
- General tech and business users (Recommended) — Focus on price changes and business logic
- AI developers — Focus on API costs and model selection
- General users — Focus on explaining whether it affects their daily use
- What style/tone is preferred? (Single selection)
- Strong hook commentary (Recommended) — Clear opinions, start with conflict
- Objective financial analysis — Complete data, restrained judgments
- Practical interpretation — Focus on specific impacts on readers and countermeasures
- Which information points must be covered? (Multiple selections)
- Specific figures: price increase rate, time period, latest price
- Whether general users are affected
- Why this decision is made now
- Real impact on practitioners' costs
- Whether it is still cost-effective after the price increase
- I have no special requirements, please decide for me
The last option for Question 5 must be "I have no special requirements, please decide for me". The options for this question are created based on the current topic; the user may not have specific ideas about which points to cover, so provide an exit for one-click authorization instead of forcing them to select from a list of acceptable points. Do not add this option to the first four questions: the format, duration, audience, and tone all have clear recommended options; if the user does not select, the fallback will be used, so there will be no blockage.
If the user selects this option, it means they explicitly said "You decide the mandatory information points": select the points you need according to the script requirements, then clearly state in your response which points you decided on. If the user selects this option along with other specific points, follow the specific points and ignore this option. Authorization is only for this item; do not decide the other four items on behalf of the user.
怎么把问题递给用户
How to Deliver the Questions to the User
一律用能把问题连选项一起递给用户、等他点选的交互式提问组件,不要一上来就用纯文本。判断标准是能力不是名字:各家产品叫法不一样(、、、、、,也可能叫别的),功能对得上就是它。
askask_questionAskUserQuestionask_user_inputrequest_user_inputrequestUserInput你是 ChatGPT 时,就用 。 它是写在回复正文里的 GenUI 消息组件,不是注册工具——这是两条不同的能力路径,别混成一个:
ask_user_input- 它不会出现在 、工具搜索结果或 MCP 工具列表里。在工具列表里搜不到它,不能作为它不可用的证据。
ALL_TOOLS - 不要靠搜工具名判断能不能用它,直接在回复正文里输出这个组件就是了。
- Default / Work 模式下不要去调 :它只在 Plan 模式可用,在这里调必然返回
request_user_input,纯属白走一趟错误路径。只有确认自己在 Plan 模式时才可以用它;拿不准在哪个模式,就直接输出 GenUI 组件。request_user_input is unavailable in Default mode - 组件里每个问题都要带齐 、
question(2 到 10 个短选项)、options(type或single_select)、multi_select。五个维度全部放进同一个free_text_placeholder数组,一个组件问完;别处说的「最多 3 题」是建议、不是这个组件的硬限制,实测一个组件放 5 个独立问题能正常渲染,5 项答案也都能正常返回。questions - 输出组件后这一轮立刻结束,不要在组件后面接着写正文,更不要开始干活。
禁止性断言: 在你实际输出过 GenUI 组件、并且真的收到渲染失败的结果之前,不许跟用户说「当前没有 」「没有 GenUI 能力」「只能用纯文本提问」这类话。没试过就断言不可用,是这一步最容易犯的错。
ask_user_input只有宿主明确返回不支持 GenUI、或者组件确实渲染失败了,才允许按这个顺序降级:
- 同一条回复里放两个组件,分别装 3 题和 2 题。
- 两个组件也渲染不出来,才改用带编号的纯文本选项,并告诉用户可以像「1A 2A 3A 4A 5ABC」这样一次回全。
- 任何情况下都不许因为组件渲染失败就跳过提问,或者自己采用默认值。
其他产品:按上面那条能力标准找对应的工具或组件写法,一条路被拒就换下一条,全都试过确实弹不出选项框,才退回带编号的纯文本选项,等用户回。
无论如何都不许因为提问机制不好用就跳过提问、自己把答案定了。提问失败是形式问题,不是可以自行拍板的理由。
Always use an interactive question component that presents the questions along with options and waits for the user's selection. Do not start with plain text. Judge by capability, not name: different products have different names (e.g., , , , , , , etc.); use the one with the correct function.
askask_questionAskUserQuestionask_user_inputrequest_user_inputrequestUserInputWhen you are ChatGPT, use . It is a GenUI message component written in the response body, not a registered tool — this is a different capability path, do not confuse them:
ask_user_input- It will not appear in , tool search results, or the MCP tool list. Not finding it in the tool list is not evidence that it is unavailable.
ALL_TOOLS - Do not judge whether it can be used by searching for the tool name; directly output this component in the response body.
- Do not call in Default/Work mode: it is only available in Plan mode, and calling it here will definitely return
request_user_input, which is a useless error path. Only use it if you confirm you are in Plan mode; if you are unsure, directly output the GenUI component.request_user_input is unavailable in Default mode - Each question in the component must include ,
question(2 to 10 short options),options(typeorsingle_select), andmulti_select. Put all five dimensions into the samefree_text_placeholderarray and ask them in one component; the "maximum 3 questions" mentioned elsewhere is a recommendation, not a hard limit for this component. Tests show that putting 5 independent questions in one component renders normally, and all 5 answers are returned correctly.questions - End the round immediately after outputting the component; do not continue writing text after the component, let alone start working.
Prohibited Assertion: Do not tell the user things like " is not currently available", "No GenUI capability", "Only plain text questions can be used" until you have actually output the GenUI component and received a rendering failure result. Asserting unavailability without trying is the most common mistake in this step.
ask_user_inputOnly when the host explicitly returns that GenUI is not supported, or the component indeed fails to render, is it allowed to downgrade in the following order:
- Put two components in the same response, containing 3 questions and 2 questions respectively.
- If two components also fail to render, switch to numbered plain text options and tell the user they can reply with all answers at once like "1A 2A 3A 4A 5ABC".
- Under no circumstances should you skip the questions or use default values on your own due to component rendering failure.
Other Products: Find the corresponding tool or component writing method according to the capability standard above; if one path is rejected, switch to another. Only after all attempts fail to bring up the option box should you revert to numbered plain text options and wait for the user's reply.
Under no circumstances should you skip the questions or make decisions on your own due to poor question mechanism. Failure to ask is a form issue, not a reason to make decisions independently.
必问:成片形态
Mandatory Question: Video Format
哪怕口播文案已经定稿,这一项也要问一次;文案还没定稿时,跟下面「文案还没定稿时」那几项放在同一轮里问:
- 竖屏信息版(默认):成片是 9:16,内容按 16:9 横屏渲染后嵌进画布中部,钩子标题在上方、字幕与 IP 形象在下方
- 竖屏视频:内容铺满 9:16 画布
- 横屏视频:内容铺满 16:9 画布
上面这三条是给你自己理解形态用的。问的时候把差别和适用场景带进选项里,别把画布参数原样甩给用户,照这个写:
- 竖屏信息版(推荐)— 9:16 画布中间嵌 16:9 画面,顶部有大标题,信息量最足
- 竖屏视频 — 9:16 满屏,适合抖音、视频号、小红书
- 横屏视频 — 16:9 满屏,适合 B 站、YouTube 和电脑上看
用户没回答、或让你自己定,就按竖屏信息版,并在回复里写明「成片形态按竖屏信息版兜底」。
形态不是你要交的字段,它是用户在 Niwo 渲染界面上自己选的,写进 会直接校验失败。你问它只为两件事:
content.json- 决定素材取向:竖屏视频收竖构图,横屏视频与竖屏信息版都收横构图。
- 决定要不要写钩子标题:只有竖屏信息版有那块版面。
所以形态一确认,钩子标题的写法就跟着定死了,两者不能各写各的:
- 用户选竖屏信息版:里必须有
content.json。缺了它顶部那条标题带没有内容,成片开头会空掉整块版面,Niwo 不会自动拿标题或文案去补。它要跟口播文案一起写、一起给用户确认,做法见第一步。hook_headline - 用户选竖屏视频或横屏视频:一律不写 。这两种形态没有这块版面,写了画面上也不会出现;更麻烦的是 Niwo 会把它当成「这个包是按信息版收的」,直接把渲染界面的形态初值设成竖屏信息版,跟用户刚选的形态对着干。
hook_headline - 用户没回答形态、由你按竖屏信息版兜底时,照信息版的规矩要写,并且照第一步的做法跟文案一起给用户确认;同时在回复里写明形态用了这个默认值。
hook_headline
确认完在 里留一句,比如「素材按横构图收集,建议在 Niwo 里选竖屏信息版」,提醒用户在界面上选一致的形态。选反了视频会被满屏裁切,主体容易切掉。
notesEven if the narration script is finalized, this item must be asked once; when the script is not finalized, ask it in the same round as the items under "When the Script Is Not Finalized":
- Vertical Info-Style (Default): The finished video is 9:16, with content rendered in 16:9 horizontal format embedded in the middle of the canvas, hook headline at the top, and subtitles + IP character at the bottom
- Vertical Video: Content fills the 9:16 canvas
- Horizontal Video: Content fills the 16:9 canvas
The above three points are for your own understanding of the formats. When asking, include the differences and applicable scenarios in the options; do not directly present the canvas parameters to the user. Write them like this:
- Vertical Info-Style (Recommended) — 9:16 canvas with 16:9 video embedded in the middle, large headline at the top, highest information density
- Vertical Video — 9:16 full screen, suitable for Douyin, WeChat Channels, Xiaohongshu
- Horizontal Video — 16:9 full screen, suitable for Bilibili, YouTube, and viewing on computers
If the user does not answer or asks you to decide, use the Vertical Info-Style and clearly state in your response "Video format defaults to Vertical Info-Style".
The format is not a field you need to submit; it is selected by the user on Niwo's rendering interface. Writing it into will directly cause validation failure. You ask this only for two purposes:
content.json- Determine asset orientation: Vertical videos require vertical composition; horizontal videos and Vertical Info-Style require horizontal composition.
- Determine whether to write a hook headline: Only the Vertical Info-Style has this layout space.
Therefore, once the format is confirmed, the way to write the hook headline is fixed; they cannot be written independently:
- If the user selects Vertical Info-Style: must include
content.json. Without it, the top headline bar will have no content, leaving a blank area at the start of the video, and Niwo will not automatically fill it with the title or script. It must be written and confirmed with the user along with the narration script; see Step 1 for details.hook_headline - If the user selects Vertical Video or Horizontal Video: Do not write . These formats do not have this layout space, so it will not appear on the screen; more importantly, Niwo will treat the bundle as "prepared for Info-Style" and directly set the default format on the rendering interface to Vertical Info-Style, conflicting with the user's selected format.
hook_headline - If the user does not answer the format and you default to Vertical Info-Style: must be written according to the Info-Style rules and confirmed with the user along with the script; at the same time, clearly state in your response that the default format was used.
hook_headline
After confirmation, leave a note in , e.g., "Assets collected in horizontal composition, recommend selecting Vertical Info-Style on Niwo", to remind the user to select the same format on the interface. If the selection is reversed, the video will be cropped to fill the screen, and the subject may be easily cut off.
notes文案还没定稿时,跟形态同一轮问这些
When the Script Is Not Finalized, Ask These in the Same Round as the Format
文案还没定稿时,下面四项跟成片形态放在同一轮里问,不要另开一轮:
- 目标时长大概多长(字数怎么换算见下面「时长与字数」)
- 讲给谁看
- 什么风格口吻
- 哪些信息点必须讲到
这一轮拿到回答之前,不要开始写文案草稿,也不要开始搜索下载素材。用户回答之后再给一版文案草稿,竖屏信息版连钩子标题一起给,用户改完并明确说「就这版」之后再继续找素材。
When the script is not finalized, ask the following four items in the same round as the video format; do not start a new round:
- Approximate target duration (see "Duration vs. Word Count" below for word count conversion)
- Who the content is for
- What style/tone to use
- Which information points must be covered
Do not start writing the script draft or searching/downloading assets until you receive the user's answers in this round. After receiving the answers, provide a script draft; for Vertical Info-Style, include the hook headline along with the script. Only proceed to find assets after the user revises it and explicitly says "This version is final".
第一步:确认口播文案
Step 1: Confirm the Narration Script
content.jsonscript- 如果对话里已经有一份用户确认过的定稿,直接用它。
- 如果还没定稿,先别自己拍板。第零步那一轮已经把时长、受众、风格、必讲信息点一起问过了,按用户的回答写草稿,用户改完并明确说「就这版」之后再继续。
- 文案该怎么写、什么风格更好,按你和用户聊出来的结论来,这份 skill 不替你决定。
第零步选的是竖屏信息版时,钩子标题和口播文案是一份东西的两半,要一起走完这个流程:草稿里就带上钩子标题,跟文案一并给用户看,用户改完并明确说「就这版」才算定稿。不要文案定完就去找素材、把钩子标题留到写 时自己补一条——那条标题是开场几秒里字最大的一行,用户没点头不算定稿。
content.json哪怕用户直接甩来一份已经定稿的口播文案,钩子标题也还是要单独提炼一版给他确认,因为文案定稿不代表标题定稿。
唯一的硬要求:最终放进 的必须是能直接逐字朗读的成品正文,不要标题、分段小标、镜头说明、emoji 或者括号里的舞台提示。
scriptThe in is a finalized narration script: it will be read word-for-word as voiceover and displayed word-for-word as subtitles; Niwo will not rewrite, polish, add, or delete any words. Therefore, finalize the script before starting to find assets.
scriptcontent.json- If there is already a finalized script confirmed by the user in the conversation, use it directly.
- If it is not finalized, do not make decisions on your own. The round in Step Zero has already asked about duration, audience, style, and mandatory information points. Write a draft according to the user's answers, and only proceed after the user revises it and explicitly says "This version is final".
- How to write the script and what style is better should follow the conclusions from your conversation with the user; this skill does not make decisions for you.
When the format selected in Step Zero is Vertical Info-Style, the hook headline and narration script are two parts of the same content and must go through this process together: include the hook headline in the draft, present it to the user along with the script, and only consider it finalized after the user revises it and explicitly says "This version is final". Do not finalize the script and then go find assets, leaving the hook headline to be added later when writing — that headline is the largest text in the opening seconds, and it is not finalized until the user approves it.
content.jsonEven if the user directly provides a finalized narration script, the hook headline must still be extracted separately and confirmed with the user, as a finalized script does not mean a finalized headline.
The only hard requirement: the final must be a finished text that can be read word-for-word directly, without titles, section subheadings, shot descriptions, emojis, or stage directions in parentheses.
script时长与字数
Duration vs. Word Count
成片时长完全由配音朗读 的实际长度决定,没有别的旋钮:Niwo 不会为了凑时长增删一个字,也不会拉长或压缩画面来对齐目标秒数。所以字数写多少,片子就多长。
script中文口播原速约每秒 6 字,按这个速率换算目标时长:
- 30 秒 ≈ 180 字
- 1 分钟 ≈ 360 到 380 字
- 1 分半 ≈ 540 字
The finished video duration is entirely determined by the actual length of the voiceover reading the ; there are no other controls: Niwo will not add or delete words to meet the duration, nor will it stretch or compress the footage to align with the target seconds. Therefore, the number of words written determines the length of the video.
scriptThe original speed of Chinese voiceover is approximately 6 words per second; convert the target duration according to this rate:
- 30 seconds ≈ 180 words
- 1 minute ≈ 360 to 380 words
- 1.5 minutes ≈ 540 words
多音字读音
Pronunciation of Polyphonic Characters
文案定稿后,扫一遍里面的多音字,把配音容易读错的写进 的 。比如「调用量」的「调」该读 diào,配音默认会读成 tiáo。写法见 。
content.jsonpronunciationsreferences/content-schema.md只标真正会读错的那个字,标得越少配音的语气越自然,字幕时间戳也越准,所以别整句都注音。
After finalizing the script, scan for polyphonic characters and write those that the voiceover is likely to mispronounce into in . For example, the character "调" in "调用量" (diàoyòngliàng) should be pronounced diào, but the voiceover will default to tiáo. See for writing methods.
pronunciationscontent.jsonreferences/content-schema.mdOnly mark the character that is actually likely to be mispronounced; the fewer marks, the more natural the voiceover tone, and the more accurate the subtitle timestamps. Do not add phonetic notation to the entire sentence.
钩子标题
Hook Headline
竖屏信息版必写,另两种形态一律不写。 顶部标题带上那条钩子大标题,从口播文案里提炼,写进 的 。它不进配音、不进字幕,只负责开场抓注意力。
content.jsonhook_headline- 推荐 2 行:第一行铺垫,第二行给结果或冲突。
- 每行不超过 12 个中文字,把最有冲击力的数字或结论用 圈出来做高亮。英文与数字只占半个字宽,
[[ ]]这样的行折合 8.2 个字,不算超。超过 12 会告警,超过 14 会校验失败。DeepSeek砸1.4亿 - 跟文案一起给用户确认,别自己拍板。给 1 到 2 个候选让用户挑或改,用户点头后再往下走。
- 另两种形态没有这块版面,写了也不会出现在画面上,还会把 Niwo 渲染界面的形态初值拧成竖屏信息版,跟用户选的形态冲突,所以一律省略此字段。
示例:
json
"hook_headline": {
"lines": ["DeepSeek砸1.4亿", "一锁就是[[36个月]]"]
}写法与校验规则见 的 小节。
references/content-schema.mdhook_headlineMandatory for Vertical Info-Style, omitted for the other two formats. The top headline bar includes this hook headline, extracted from the narration script and written into in . It is not included in the voiceover or subtitles; it only serves to grab attention at the opening.
hook_headlinecontent.json- Recommended to be 2 lines: the first line sets the scene, the second line presents the result or conflict.
- Each line should not exceed 12 Chinese characters; enclose the most impactful number or conclusion in for highlighting. English words and numbers take up half a character width; a line like "DeepSeek砸1.4亿" (DeepSeek invests 140 million) is equivalent to 8.2 characters and does not exceed the limit. Exceeding 12 characters will trigger a warning, and exceeding 14 will cause validation failure.
[[ ]] - Confirm it with the user along with the script; do not make decisions on your own. Provide 1 to 2 candidates for the user to choose or revise, and only proceed after the user approves.
- The other two formats do not have this layout space; writing it will not appear on the screen, and will also set the default format on Niwo's rendering interface to Vertical Info-Style, conflicting with the user's selected format. Therefore, omit this field entirely.
Example:
json
"hook_headline": {
"lines": ["DeepSeek砸1.4亿", "一锁就是[[36个月]]"]
}See the section in for writing methods and validation rules.
hook_headlinereferences/content-schema.md资料来源
Sources
文案定稿后,把引用过的公开资料名称写进 的 。它不进配音、不进字幕,只渲染为画面底部的小字。
content.jsonsources默认都写上,不用问用户要不要。写文案时你查过哪些资料,你自己最清楚;露不露是用户在 Niwo 渲染界面上的开关,不是你要替他决定的事。
- 有引用就写,一条都没有就省略这个字段,别为了凑数编来源。
- 只写来源名称(公告名、媒体报道名等),不要写 URL,单条不超过 24 个中文字宽。
- 不要写「资料来源:」前缀,那个由 Niwo 拼。
- 免责声明是另一回事:它的开关和正文都在用户那边的渲染参数里,跟 各自独立,不要写进
sources。content.json
示例:
json
"sources": ["宇树科技招股书", "DeepSeek 官方公告", "新浪财经"]写法与校验规则见 的 小节。
references/content-schema.mdsourcesAfter finalizing the script, write the names of public references used into in . It is not included in the voiceover or subtitles; it is only rendered as small text at the bottom of the screen.
sourcescontent.jsonWrite it by default, do not ask the user if they want it. You know best which references you checked when writing the script; whether to display it is a toggle on the user's rendering interface, not something you decide on their behalf.
- Write it if there are references; omit this field if there are none, do not fabricate sources to fill the field.
- Only write the source name (e.g., announcement name, media report name), do not write URLs; each entry should not exceed 24 Chinese characters in width.
- Do not add the prefix "Source:"; that will be added by Niwo.
- Disclaimer is a separate matter: its toggle and content are in the user's rendering parameters, independent of ; do not write it into
sources.content.json
Example:
json
"sources": ["Unitree Robotics Prospectus", "DeepSeek Official Announcement", "Sina Finance"]See the section in for writing methods and validation rules.
sourcesreferences/content-schema.md第二步:收集素材
Step 2: Collect Assets
需要两类素材。
Two types of assets are needed.
B-roll 视频素材
B-Roll Video Footage
直接从网络、公开素材网站或相关网页中搜索并下载真实视频。可以包括人物、场景、产品操作、平台使用、行业画面、新闻现场,以及抽象概念对应的实拍。
Search and download real videos directly from the web, public asset websites, or related web pages. These can include people, scenes, product operations, platform usage, industry footage, news scenes, and live-action shots corresponding to abstract concepts.
网络图片素材
Web Image Assets
通过 Bing、百度等图片搜索引擎,寻找与文案内容直接相关的真实网络图片。优先收集:
- 相关产品、公司和人物图片
- App 真实界面与历史版本截图
- 新闻报道配图
- 行业报告与数据图表
- 产品对比图
- 活动、事件和历史资料图片
- Logo、宣传图和媒体资料图
- 能准确表达文案概念的网络照片或插图
Use image search engines like Bing or Baidu to find real web images directly related to the script content. Prioritize collecting:
- Images of related products, companies, and people
- Screenshots of real app interfaces and historical versions
- News report illustrations
- Industry reports and data charts
- Product comparison charts
- Images of activities, events, and historical materials
- Logos, promotional images, and media materials
- Web photos or illustrations that accurately express the concepts in the script
明确不要的素材
Assets to Explicitly Avoid
- 不要原创 Motion 动效
- 不要用前端框架渲染概念动画
- 不要把 B-roll 视频抽帧后当作图片素材
- 不要用低信息量的占位图代替真实资料
- 不要为了凑数量收集大量与文案只有弱关联的素材
你交的素材就是成片的全部画面来源:导入素材包后,Niwo 不会再联网搜图或补空镜。所以素材要覆盖整篇文案,别让某一段找不到画面可配;同时也别为了凑数塞弱关联的东西,宁缺毋滥——真配不上时会退化成纯文字卡,那比配错画面好。
一分钟到一分半的片子,通常 20 到 30 张图加 8 到 10 段视频就够。取材渠道和实际数量你自己判断。素材远超这个量时会沿 manifest 顺序均匀抽取,所以 manifest 的排列顺序最好跟着文案推进走。
素材方向尽量和第零步确认的成片形态一致:
- 视频:方向不一致时,成片会按满屏裁切()适配画布,主体靠边时容易被切掉。竖屏成片优先找竖屏视频,横屏成片优先找横屏视频。
cover - 图片:一律完整展示、不裁切;方向不一致时只是上下或左右多出模糊铺底,横竖屏混用问题不大。
- No original motion effects
- No concept animations rendered with front-end frameworks
- No frame-extracted images from B-roll videos used as image assets
- No low-information placeholder images instead of real materials
- No large amounts of assets weakly related to the script collected just to meet the quantity
The assets you submit are the entire source of footage for the finished video: after importing the asset bundle, Niwo will not search for images online or fill in empty shots. Therefore, the assets must cover the entire script; do not leave any section without matching footage. At the same time, do not stuff weakly related assets just to meet the quantity; it is better to have fewer high-quality assets than many irrelevant ones — if there is no matching footage, it will degrade to a text-only card, which is better than mismatched footage.
For a 1 to 1.5-minute video, usually 20 to 30 images plus 8 to 10 video clips are sufficient. You decide the source channels and actual quantity. When there are far more assets than this amount, Niwo will extract them evenly according to the manifest order, so the manifest order should preferably follow the script's progression.
Try to align the asset orientation with the video format confirmed in Step Zero:
- Videos: If the orientation does not match, the finished video will be cropped to fill the screen (), and the subject may be easily cut off if it is near the edge. Prioritize vertical videos for vertical finished videos, and horizontal videos for horizontal finished videos.
cover - Images: All are displayed in full without cropping; if the orientation does not match, only blurred padding will be added above/below or left/right, so mixing vertical and horizontal images is not a major issue.
第三步:给每个视频定好用哪几秒
Step 3: Determine Which Seconds of Each Video to Use
成片里一个镜头只用几秒,所以别丢一段几十秒的原片过来就完事。两种做法任选:
- 自己裁好(推荐):直接交裁完的短片段,一个文件就是一个镜头,8 到 12 秒最好用。宁可长一点:镜头时长由配音决定,最长可能到 9 秒,片段比镜头短时只能把它慢放补足,短得太多画面就会明显发滞。
bash
ffmpeg -ss <起点秒> -to <终点秒> -i 原视频.mp4 \
-map 0:v:0 -c:v libx264 -pix_fmt yuv420p -preset veryfast -movflags +faststart \
videos/01-xxx.mp4- 交原片并在 manifest 里标出选段:写上 /
clip_start_seconds,Niwo 按这个窗口裁。clip_end_seconds
裁或者标之前,用 确认时长,别落在黑屏、片头台标或转场糊帧上,挑画面稳定、主体清晰的一段。
ffprobe -v error -show_entries format=duration -of csv=p=0 原视频.mp4有个坑要避开:如果你交的是长原片,manifest 里又写了 和 、却没写选段时间,Niwo 会认为你已经挑好了,直接从第 0 秒开始取——大概率取到片头。要么裁好,要么把选段时间写上,要么干脆别写 和 ,让 Niwo 自己看片挑镜头。
summarytagssummarytags音轨和编码不用管:Niwo 会统一重编码成静音的 H.264 mp4,成片用自己的配音和 BGM。文件是常见的 mp4 / mov / mkv / webm 就行,带不带原声都无所谓。
Each shot in the finished video only uses a few seconds, so do not just submit a raw video of tens of seconds. Choose one of the two methods:
- Trim it yourself (Recommended): Submit the trimmed short clips directly; one file is one shot, 8 to 12 seconds is most suitable. It is better to make it a bit longer: the shot duration is determined by the voiceover, which can be up to 9 seconds. If the clip is shorter than the shot, it will only be slowed down to fill the duration; if it is too short, the footage will obviously become sluggish.
bash
ffmpeg -ss <start_seconds> -to <end_seconds> -i original_video.mp4 \
-map 0:v:0 -c:v libx264 -pix_fmt yuv420p -preset veryfast -movflags +faststart \
videos/01-xxx.mp4- Submit the raw video and mark the selected segment in the manifest: Write /
clip_start_seconds, and Niwo will trim it according to this window.clip_end_seconds
Before trimming or marking, use to confirm the duration. Do not select segments with black screens, opening logos, or blurry transition frames; choose a segment with stable footage and clear subject.
ffprobe -v error -show_entries format=duration -of csv=p=0 original_video.mp4Avoid this pitfall: if you submit a long raw video, write and in the manifest but do not write the segment time, Niwo will assume you have already selected it and start from the 0th second — which will likely capture the opening logo. Either trim it, write the segment time, or do not write and at all, letting Niwo select the shot automatically.
summarytagssummarytagsDo not worry about audio tracks or encoding: Niwo will re-encode all videos into silent H.264 mp4, and the finished video will use its own voiceover and BGM. The file can be in common formats like mp4 / mov / mkv / webm; it does not matter if it has original audio or not.
第四步:写 manifest.json
Step 4: Write manifest.json
读 ,按里面的字段说明写。
references/manifest-schema.md要点: 是单行中文、只描述画面里看得见的东西; 2 到 5 个中文短标签;已经裁好的视频不要写 / 。
summarytagsclip_start_secondsclip_end_secondsRead and write according to the field descriptions inside.
references/manifest-schema.mdKey points: is a single line of Chinese, only describing what is visible in the footage; are 2 to 5 short Chinese labels; do not write / for trimmed videos.
summarytagsclip_start_secondsclip_end_seconds第五步:写 content.json
Step 5: Write content.json
读 ,按里面的模板写。
references/content-schema.md只有 、、、、、 和 这几项: 固定填 ,是必填的协议版本号,漏了会直接校验失败;其余全是内容。除这几项之外的字段一律不接受,把渲染参数(包括成片形态)写进去同样会校验失败。
schema_versiontitlescripthook_headlinesourcespronunciationsnotesschema_version1写之前再对一遍第零步确认的形态:竖屏信息版必须带上用户确认过的 ,竖屏视频与横屏视频必须没有这个字段。校验脚本不知道用户选了哪个形态,这一项对不上它查不出来,只能你自己核。
hook_headlineRead and write according to the template inside.
references/content-schema.mdOnly include these fields: , , , , , , and : must be filled with , which is a required protocol version number; omitting it will directly cause validation failure. All other fields are content. No other fields are accepted; writing rendering parameters (including video format) into it will also cause validation failure.
schema_versiontitlescripthook_headlinesourcespronunciationsnotesschema_version1Before writing, double-check the format confirmed in Step Zero: Vertical Info-Style must include the user-confirmed , while Vertical Video and Horizontal Video must not have this field. The validation script does not know which format the user selected, so it cannot detect this mismatch; you must check it yourself.
hook_headline第六步:自校验并打包
Step 6: Self-Check and Package
打包前先跑自校验,它会检查字段、文件对应关系和视频选段窗口。脚本在这个 skill 目录的 下,用绝对路径调用:
scripts/bash
python3 <skill 目录>/scripts/validate_bundle.py <素材包目录>有报错就改到通过为止,不要带着报错打包。如果你拿到的是一份复制粘贴的提示词、手上没有这个脚本,就按文末的自查清单逐项人工核对。校验通过后:
bash
cd <素材包目录>
zip -r ../bundle.zip content.json manifest.json images videos -x '.*' -x '__MACOSX/*'zip 里不要放软链接,也不要放上面结构之外的其他文件。
Run the self-check before packaging; it will check fields, file correspondence, and video segment windows. The script is in the directory of this skill; call it with the absolute path:
scripts/bash
python3 <skill_directory>/scripts/validate_bundle.py <asset_bundle_directory>Fix any errors until it passes; do not package with errors. If you only have a copied prompt and do not have this script, manually check each item according to the self-check list at the end. After passing the validation:
bash
cd <asset_bundle_directory>
zip -r ../bundle.zip content.json manifest.json images videos -x '.*' -x '__MACOSX/*'Do not include soft links in the zip, nor any other files outside the above structure.
交付
Delivery
把 zip 交给用户,并告诉他上传到 Niwo 生成视频:
<!-- TODO: 替换成正式的上传入口地址与产品名 -->
上传前用户可以自己再筛一遍素材:不满意的直接删文件,想加的直接放进 或 ——多出来的文件不写进 manifest 也能用,Niwo 会自动补描述。
images/videos/Deliver the zip file to the user and tell them to upload it to Niwo to generate the video:
<!-- TODO: Replace with official upload entry address and product name -->
Before uploading, the user can re-screen the assets: delete unsatisfactory files directly, or add new ones directly to or — extra files not written in the manifest can also be used, and Niwo will automatically add descriptions.
images/videos/交付前自查
Pre-Delivery Self-Check
- 开工前已把成片形态与文案要求一次问完,答案来自用户本次的回复,不是从记忆或历史偏好推断的;用了兜底默认值的已在回复里说明
- 跑通、无报错
scripts/validate_bundle.py - 与
content.json都写了manifest.json"schema_version": 1 - 素材取向与第零步确认的成片形态一致,并在 里提醒用户在 Niwo 里选同一个形态
notes - 是用户确认过的定稿口播,逐字可读,没有标题和舞台提示
script - 形态是竖屏信息版时 存在且用户确认过;是竖屏视频或横屏视频时没有这个字段
hook_headline - 每个素材都在 里有条目,
manifest.json路径与实际文件一一对应file - 每个视频要么已经裁成几秒的短镜头,要么在 manifest 里标了选段时间
- 都是单行中文,描述的是画面而不是含义
summary - 素材都和文案有实打实的关联,没有占位图、没有抽帧图、没有渲染动画
- 里没有渲染参数,只有内容
content.json - (如有)为 1 到 3 行、每行 12 个中文字内(英文数字按半个字算)、至少一处
hook_headline高亮[[ ]] - 文案引用过公开资料时 已默认写上,只写来源名称,不写「资料来源:」前缀,也不写免责声明
sources
- Before starting, all questions about video format and script requirements were asked in one go, and the answers came from the user's current response, not inferred from memory or historical preferences; any fallback default values used were clearly stated in the response
- ran successfully with no errors
scripts/validate_bundle.py - Both and
content.jsonhavemanifest.jsonwritten"schema_version": 1 - Asset orientation matches the video format confirmed in Step Zero, and a note in reminds the user to select the same format on Niwo
notes - is the finalized narration confirmed by the user, readable word-for-word, with no titles or stage directions
script - exists and is confirmed by the user when the format is Vertical Info-Style; this field is omitted when the format is Vertical Video or Horizontal Video
hook_headline - Each asset has an entry in , and the
manifest.jsonpath corresponds to the actual file one-to-onefile - Each video is either trimmed into a short shot of a few seconds, or the segment time is marked in the manifest
- All are single-line Chinese, describing the footage rather than its meaning
summary - All assets are closely related to the script, with no placeholder images, frame-extracted images, or rendered animations
- contains no rendering parameters, only content
content.json - (if present) has 1 to 3 lines, each within 12 Chinese characters (English/numbers count as half a character), with at least one
hook_headlinehighlight[[ ]] - is written by default when public references are used in the script, only including source names, no "Source:" prefix, and no disclaimer
sources