4k-vfx
Compare original and translation side by side
🇺🇸
Original
English🇨🇳
Translation
Chinese4k-vfx
4K VFX
Maps a video-to-video VFX move onto the Pika MCP's Seedance 4K reference-to-video. The user hands you a clip + the change they want; you read EVERY frame (locally extracted with ffmpeg and tiled into contact sheets) and understand the audio (speech + music + SFX + ambience), author a Seedance-faithful prompt that locks the original (face, gestures, camera move) and time-codes the change to the right beat, then re-render the same shot in 4K with the VFX baked in. One run produces one 4K clip for one requested change.
将视频转视频的VFX效果映射到Pika MCP的Seedance 4K参考转视频功能上。用户提供一段视频片段以及想要的修改效果;你需要读取每一帧(通过ffmpeg本地提取并拼接成联系表)并理解音频(语音、音乐、音效和环境音),编写符合Seedance规范的提示词,锁定原始画面的(面部、手势、镜头运动),并将修改效果与正确的节拍时间码对齐,然后重新渲染同一镜头的4K版本,将VFX效果内置其中。一次运行针对一项修改请求生成一个4K视频片段。
How it works — why reading EVERY frame (and the audio) is the mechanism (do not skip the reading pass)
工作原理——为何必须读取每一帧(和音频)(请勿跳过读取步骤)
Seedance treats the reference video as a motion/style anchor, NOT a pixel-locked base — it does not copy the input frame-for-frame, it re-generates the shot guided by your prompt. So fidelity comes from the PROMPT: the more exhaustively the prompt describes the original (subject identity, exact gestures, camera move, framing, lighting, palette, wardrobe) and the audio (dialogue + lip-sync, music/SFX beats), the more the 4K output reads as the same shot with the change layered on. Reading EVERY frame — not a sample — is how you build that exhaustive description: you locally extract all frames with ffmpeg, tile them 25-to-a-contact-sheet (5×5), and analyze every sheet so the prompt is built from the entire timeline rather than a handful of stills. Skip the read (or read only a sample) and the output drifts — a different face, a different camera move, a different room, a missed motion beat. The reading pass — all frames + the audio — IS the skill.
Seedance将参考视频视为运动/风格锚点,而非像素锁定的基础——它不会逐帧复制输入内容,而是根据你的提示词重新生成镜头。因此,还原度取决于提示词:提示词越详尽地描述原始内容(主体身份、精确手势、镜头运动、构图、光线、色调、服装)以及音频(对话与唇形同步、音乐/音效节拍),4K输出就越像是在原始镜头基础上叠加了修改效果。读取每一帧而非采样帧,是构建详尽描述的关键:你需要通过本地ffmpeg提取所有帧,将每25帧拼接成一张联系表(5×5网格),并分析每张联系表,这样提示词就能基于整个时间线而非少量静态帧构建。跳过读取步骤(或仅读取采样帧)会导致输出内容偏离——比如面部不同、镜头运动不同、场景不同、错过运动节拍。读取步骤——所有帧+音频——就是这项技能的核心。
Prerequisites
前置条件
pika MCP available, and local in the run environment (used to extract every frame and build the contact sheets). Two inputs from the user: (1) a video clip (a local file or a public URL) and (2) the VFX change they want made to it. Tools used: , , , , , , (frame extraction + tiling is done locally with ffmpeg, not an MCP tool).
ffmpegupload_assetprobe_mediaanalyze_mediatranscribe_audioestimate_costgenerate_reference_videotask_status需具备Pika MCP,且运行环境中拥有本地(用于提取所有帧并构建联系表)。需要用户提供两个输入:(1) 视频片段(本地文件或公共URL)和**(2) 想要添加的VFX修改效果**。使用的工具包括:、、、、、、(帧提取与拼接通过本地ffmpeg完成,而非MCP工具)。
ffmpegupload_assetprobe_mediaanalyze_mediatranscribe_audioestimate_costgenerate_reference_videotask_statusStage 0 — Gather inputs (settle this first)
阶段0——收集输入(先完成此步骤)
You need exactly two things before generating:
- The video — the original clip to transform. If they gave a local file, you will upload it (Step 1); if they gave a public URL, keep it.
- The change / VFX — what should be different in the output (e.g. "turn the plaza into an open desert on a finger snap", "make it snow", "set the room on fire behind me", "morph my jacket into glowing armor").
If either is missing, ask for it before doing anything else — do not invent a change, and do not proceed on a clip you cannot reach. Once both are in hand, say in one line what you're about to do ("Reading every frame of your clip to build the desert-on-snap 4K prompt…") and run Steps 1–5 without further check-ins — but stop at the Step 6 agreement gate before generating: the 4K render is the expensive, irreversible step, and it only fires after the user approves the prompt and cost.
生成前你需要明确获取以下两项内容:
- 视频——需要转换的原始片段。如果用户提供的是本地文件,你需要上传它(步骤1);如果是公共URL,直接使用即可。
- 修改/VFX效果——输出内容中需要改变的部分(例如:“打响指时将广场变成开阔沙漠”、“添加下雪效果”、“在我身后点燃房间”、“将我的夹克变形为发光盔甲”)。
如果任意一项缺失,请在进行其他操作前向用户索要——请勿自行编造修改效果,也不要处理无法获取的视频片段。获取两项内容后,用一句话告知用户你即将执行的操作(例如:“正在读取你视频的每一帧,以构建打响指变沙漠的4K提示词……”),然后无需进一步确认即可执行步骤1-5,但在执行步骤6生成前必须经过确认环节:4K渲染成本高昂且不可逆转,只有在用户批准提示词和成本后才能启动。
Step 1 — Upload the clip (upload_asset
)
upload_asset步骤1——上传视频片段(upload_asset
)
upload_assetReal video must be uploaded — inline base64 only works for tiny assets (<~3MB), so any actual clip goes through the presigned-upload path.
- Call with
upload_asset,filename(mime_typefor mp4), andvideo/mp4(the file's byte size).size_bytes - PUT the raw bytes to the returned .
presigned_url - Keep the returned →
public_url.state.clip_url
If the user already gave a public URL, skip this step and set to it.
state.clip_url真实视频必须上传——内嵌base64仅适用于极小资源(<约3MB),因此任何实际视频片段都需通过预签名上传路径处理。
- 调用,传入
upload_asset、filename(MP4文件为mime_type)和video/mp4(文件的字节大小)。size_bytes - 将原始字节内容PUT到返回的中。
presigned_url - 保存返回的至
public_url。state.clip_url
如果用户已提供公共URL,跳过此步骤,直接将设为该URL。
state.clip_urlStep 2 — Probe (probe_media
)
probe_media步骤2——探测媒体信息(probe_media
)
probe_mediaCall on to read duration, fps, dimensions, and aspect ratio. Use these to: (a) sanity-check the all-frame extraction in Step 3 (duration × fps ≈ total frame count ≈ 25 × number of sheets), and (b) choose the output and here, matched to the source — the Step-6 gate quotes this choice to the user and Step 7 fires with it.
probe_mediastate.clip_urlaspect_ratioduration调用读取的时长、帧率、分辨率和宽高比。这些信息用于:(a) 验证步骤3中全帧提取的合理性(时长×帧率≈总帧数≈25×联系表数量),(b) 在此选择与源视频匹配的输出和——步骤6的确认环节会向用户说明此选择,步骤7将按此参数执行。
probe_mediastate.clip_urlaspect_ratiodurationStep 3 — Extract EVERY frame locally and tile into contact sheets (ffmpeg)
步骤3——本地提取每一帧并拼接成联系表(ffmpeg)
Read the whole video, not a sample. Use local ffmpeg to extract every frame at the source rate (no drops, no exceptions), then tile them into legible contact sheets at 25 frames per sheet (5×5 grid) so each call covers 25 consecutive frames at once.
analyze_mediaWork in a relative working directory (do not use an absolute system temp path):
mkdir -p frames sheets
ffmpeg -i INPUT -vsync 0 frames/f_%05d.png # every frame, no dropsThen build the 25-per-sheet contact sheets (5×5), scaling each frame down so the sheet is legible and stamping the frame index onto each tile:
ffmpeg -i INPUT -vf "scale=480:-1,drawtext=text='%{n}':x=4:y=4:fontsize=20:fontcolor=yellow:box=1:boxcolor=black@0.5,tile=5x5" sheets/sheet_%03d.pngdrawtext=text='%{n}'tile=5x5drawtextscale=480:-1,tile=5x5INPUT读取完整视频,而非采样帧。使用本地ffmpeg以源帧率提取每一帧(无丢帧、无例外),然后将每25帧拼接成一张清晰的联系表(5×5网格),这样每次调用就能一次性分析25帧连续画面。
analyze_media在相对工作目录中操作(请勿使用绝对系统临时路径):
mkdir -p frames sheets
ffmpeg -i INPUT -vsync 0 frames/f_%05d.png # 提取每一帧,无丢帧然后构建每25帧一张的联系表(5×5),缩小每帧尺寸以保证联系表清晰,并在每个帧块上标记帧索引:
ffmpeg -i INPUT -vf "scale=480:-1,drawtext=text='%{n}':x=4:y=4:fontsize=20:fontcolor=yellow:box=1:boxcolor=black@0.5,tile=5x5" sheets/sheet_%03d.pngdrawtext=text='%{n}'tile=5x5drawtextscale=480:-1,tile=5x5INPUTStep 4 — Read EVERY sheet + understand the audio (upload_asset
, analyze_media
, transcribe_audio
)
upload_assetanalyze_mediatranscribe_audio步骤4——读取所有联系表+理解音频(upload_asset
、analyze_media
、transcribe_audio
)
upload_assetanalyze_mediatranscribe_audioThis is the fidelity step, and it has two mandatory halves: read all the frames and understand the audio. Neither is optional.
这是保证还原度的步骤,包含两个强制部分:读取所有帧和理解音频。两者缺一不可。
4a — Read all the frames (every sheet)
4a——读取所有帧(所有联系表)
For each contact sheet from Step 3 (cover ALL of them — every frame is on a sheet, no exceptions): upload it with (the sheets are images, so upload to get a URL), then call on that sheet URL with a that reads the full progression across those 25 frames — pulling out everything the prompt must lock:
upload_assetanalyze_mediaquery- Subject / identity — who/what is in frame; face, hair, build, distinguishing features.
- Exact motion & gestures — what the subject does, beat by beat, referencing the stamped frame indices (e.g. "raises right hand around frame 0048, snaps fingers around frame 0072").
- Camera move — static / pan / push-in / handheld / orbit, and its timing.
- Framing — shot size (close/medium/wide), subject position, headroom.
- Lighting — direction, hardness, color temperature, time of day.
- Palette — dominant colors, mood.
- Wardrobe / props / setting — clothing, objects, background.
Carry the read across all sheets so the entire timeline is understood start-to-end, not just a sampled middle.
针对步骤3生成的每张联系表(覆盖全部——每一帧都在联系表中,无例外):使用上传(联系表为图片,需上传以获取URL),然后调用分析该联系表URL,需包含这25帧的完整变化过程——提取提示词中需要锁定的所有信息:
upload_assetanalyze_mediaquery- 主体/身份——画面中的人物/物体;面部、发型、体型、显著特征。
- 精确动作与手势——主体的动作,逐节拍描述,并引用标记的帧索引(例如:“约第0048帧抬起右手,约第0072帧打响指”)。
- 镜头运动——静态/摇镜头/推镜头/手持镜头/环绕镜头,及其时间点。
- 构图——景别(近景/中景/远景)、主体位置、头部空间。
- 光线——方向、硬度、色温、时段。
- 色调——主导颜色、氛围。
- 服装/道具/场景——衣物、物品、背景。
遍历所有联系表进行读取,确保整个时间线从开始到结束都被理解,而非仅采样中间部分。
4b — Understand the audio (mandatory, not just speech)
4b——理解音频(强制要求,不仅是语音)
Always read the audio — it is required, not optional. Two passes:
- for speech + timestamps (the words and when they're said), so the prompt can preserve dialogue and lip-sync beats.
transcribe_audio - on the audio (or the video URL) to understand the NON-speech audio too — music, sound effects, ambience, tone, and rhythm/beats.
analyze_media
Fold the audio understanding into the prompt where it matters — preserve dialogue + lip-sync if anyone speaks, and time the VFX change to an audio beat / sound cue when the change should land on the music or a sound (e.g. snap, hit, downbeat) rather than a bare timestamp.
Assemble the all-frame read and the audio read into a single faithful description of the original — that description is the raw material for Step 5.
必须读取音频——这是必填项,而非可选项。分为两个步骤:
- 使用获取语音+时间戳(内容及其出现时间),以便提示词能保留对话和唇形同步节拍。
transcribe_audio - 调用分析音频(或视频URL),理解非语音音频——音乐、音效、环境音、音调、节奏/节拍。
analyze_media
将音频理解融入提示词的相关部分——如果有人物对话则保留对话+唇形同步,并且当修改效果需要配合音乐或声音(例如响指、撞击、重拍)时,将VFX修改效果与音频节拍/声音提示时间码对齐,而非仅使用裸时间戳。
将全帧读取结果和音频读取结果整合为对原始内容的完整描述——该描述是步骤5的素材。
Step 5 — Write the Seedance prompt (the authoring step, NEVER skip it)
步骤5——编写Seedance提示词(创作步骤,请勿跳过)
Compose ONE prompt = a faithful, exhaustive description of the original (the locked elements) + the user's requested change, time-coded + a reference to the input clip via the token. Be thorough: every detail you read in Step 4 that you omit is a detail Seedance is free to change.
@Video1Template — fill the from your Step-4 read:
[SLOTS]Re-render this exact shot @Video1 in cinematic 4K. LOCK the original: [SUBJECT/IDENTITY — face, hair, build, wardrobe], performing [EXACT GESTURES, beat by beat]. Keep the SAME camera move ([CAMERA: static / pan / push-in / handheld], [timing]), the SAME framing ([SHOT SIZE + subject position]), the SAME lighting ([direction, hardness, color temp, time of day]) and the SAME palette ([dominant colors / mood]). CHANGE: [the VFX], time-coded — at [t0]–[t1]s [what the scene looks like before the change], then at [t2]s [the change triggers / VFX appears] and [how the scene reads after]. Everything not described by the change stays identical to @Video1. Photorealistic, high detail, consistent identity throughout.Time-code the change so Seedance knows when it happens relative to the locked motion (e.g. "at 0–3s the plaza is unchanged as the subject raises their hand; at 3s on the finger snap the plaza dissolves into an open desert, sand and heat-haze replacing the buildings while the subject, camera move, and framing stay identical"). The stronger and more specific the time-coded change clause, the less Seedance ignores it.
撰写一个提示词 = 对原始内容的完整详尽描述(锁定元素)+ 用户请求的修改效果(带时间码) + 通过****标记引用输入视频。务必详尽:你在步骤4中读取但未写入提示词的任何细节,Seedance都可能随意修改。
@Video1模板——用步骤4的读取内容填充:
[占位符]以电影质感4K分辨率重新渲染此镜头 @Video1。锁定原始内容:[主体/身份——面部、发型、体型、服装],做出[精确手势,逐节拍描述]。保持相同的镜头运动([镜头类型:静态/摇镜头/推镜头/手持镜头],[时间点])、相同的构图([景别 + 主体位置])、相同的光线([方向、硬度、色温、时段])和相同的色调([主导颜色/氛围])。修改:[VFX效果],带时间码——[t0]–[t1]秒 [修改前的场景状态],然后在[t2]秒 [修改触发/VFX出现],之后[修改后的场景状态]。未被修改描述覆盖的所有内容均与@Video1完全一致。照片级真实感,高细节,全程身份一致。为修改效果添加时间码,让Seedance知道修改相对于锁定运动的发生时间(例如:“0–3秒广场保持不变,主体抬起手;3秒打响指时,广场溶解为开阔沙漠,沙子和热浪取代建筑,同时主体、镜头运动和构图保持不变”)。时间码修改条款越明确具体,Seedance就越不容易忽略它。
Step 6 — Agreement gate (estimate_cost
) — MANDATORY before generating
estimate_cost步骤6——确认环节(estimate_cost
)——生成前必须执行
estimate_costThe 4K render is the expensive step, and once fired it can't be un-spent. Never call without explicit user approval in this session. Present, in one message:
generate_reference_video- The full Step-5 prompt — verbatim, so the user can catch a wrong detail (face, gesture, timing) before it costs money.
- The generation recipe — provider , model
seedance, resolutionstandard, plus the4kandaspect_ratioyou chose from the Step-2 probe.duration - The estimated cost — call for the Step-7
estimate_costcall and quote the result. If the estimate errors, say the cost is unknown and flag that 4K standard is the top-price tier — do not silently skip the number.generate_reference_video
Then wait. Three outcomes:
- Approve → proceed to Step 7 unchanged.
- Revise — the user corrects a detail or the change clause → update the Step-5 prompt (re-reading specific sheets if the correction demands it) and re-present the gate.
- Abort / no answer → do not generate. Never treat silence as approval.
Only skip this gate if the user explicitly pre-authorized the spend in this session ("just run it, don't ask", a standing instruction to render without confirmation). Having provided the clip and the change is NOT pre-authorization — that's just Stage 0 input.
4K渲染成本高昂,一旦启动就无法撤销。在本次会话中,未经用户明确批准,请勿调用。用一条消息展示以下内容:
generate_reference_video- 步骤5的完整提示词——原文展示,以便用户在产生费用前发现错误细节(面部、手势、时间点)。
- 生成配置——服务商,模型
seedance,分辨率standard,以及步骤2探测后选择的4k和aspect_ratio。duration - 预估成本——调用获取步骤7
estimate_cost的成本并告知用户。如果预估出错,说明成本未知,并标注4K标准档为最高价格层级——请勿跳过成本说明。generate_reference_video
然后等待用户回应。可能的三种结果:
- 批准——直接执行步骤7。
- 修改——用户修正细节或修改条款→更新步骤5的提示词(如需修正则重新读取特定联系表),并重新展示确认环节。
- 终止/无回应——请勿生成。永远不要将沉默视为批准。
仅当用户在本次会话中明确预先授权费用(例如:“直接运行,不用问”、“无需确认直接渲染”的长期指令)时,才可跳过此确认环节。用户提供视频和修改效果并不等同于预先授权——这只是阶段0的输入。
Step 7 — Generate (Seedance 4K reference-to-video) (generate_reference_video
)
generate_reference_video步骤7——生成(Seedance 4K参考转视频)(generate_reference_video
)
generate_reference_videoCall with the exact 4K recipe:
generate_reference_video- :
providerseedance - :
seedance_model— REQUIRED for 4k (fast/mini cap at 720p).standard - :
resolution4k - :
reference_videos— the original footage as the motion/style anchor (Seedance accepts up to 3 videos).[state.clip_url] - : optional — style/creature/identity anchors (e.g. a generated reference image), only if you have one.
reference_images - : the Step-5 prompt (uses the
prompttoken to reference the input clip; no max length, so be exhaustive).@Video1 - :
aspect_ratio(or match the source from Step 2).auto - : matched to the source, OR set
duration:auto_duration.true
Do NOT pass (not supported on seedance — the call is rejected) and do NOT pass any fps param (it doesn't exist).
negative_prompt调用,使用精确的4K配置:
generate_reference_video- :
providerseedance - :
seedance_model——4K分辨率必填(fast/mini模型最高支持720p)。standard - :
resolution4k - :
reference_videos——原始素材作为运动/风格锚点(Seedance最多支持3个视频)。[state.clip_url] - : 可选——风格/生物/身份锚点(例如生成的参考图),仅当你有相关素材时使用。
reference_images - : 步骤5的提示词(使用**
prompt**标记引用输入视频;无长度限制,请尽量详尽)。@Video1 - :
aspect_ratio(或匹配步骤2中源视频的参数)。auto - : 匹配源视频,或设置
duration:auto_duration。true
请勿传入(Seedance不支持——调用会被拒绝),且请勿传入任何fps参数(该参数不存在)。
negative_promptStep 8 — Deliver
步骤8——交付
If the call returns inline, you have the 4K video URL. If it returns , poll in a tight loop until / / . On completion, return the 4K video URL to the user. Note: the output CDN may be egress-blocked for in-session download — so deliver the URL itself rather than trying to fetch the bytes locally.
{ task_id, status: running/background }task_status(task_id)completedfailedcancelled如果调用直接返回结果,你将获得4K视频URL。如果返回,则循环调用直到状态变为//。完成后,将4K视频URL返回给用户。注意:输出CDN可能限制会话内下载——因此直接交付URL即可,无需尝试本地获取字节内容。
{ task_id, status: running/background }task_status(task_id)completedfailedcancelledFailure modes
故障排查
| symptom | cause → fix |
|---|---|
| Clip too long / too large | Upload still works for large files; if Seedance rejects on duration, trim the source to the segment that matters, or set |
| No face / subject found in the read | A sheet may not show the subject clearly — re-run |
ffmpeg missing / | The run environment needs local |
| Identity / shot drift (different face, room, or camera move in the output) | The prompt description was not exhaustive enough. Go back to Step 4–5 and add the missing locked details (more on face, gestures, camera move, lighting) from the contact-sheet read — Seedance only preserves what the prompt names. Re-render = new spend: re-pass the Step-6 gate with the revised prompt before firing again. |
| 4k rejected / output downgraded | Ensure |
| Speech / audio not preserved or change off-beat | Run BOTH |
| Seedance ignores the requested change | Strengthen the time-coded change clause in Step 5 — make it more specific and explicitly time-anchored ("at Ns the change triggers"). The revised prompt is a new, never-approved prompt and the re-render is new spend: re-pass the Step-6 gate before firing again. |
| Job fails after a long run | If it ran several minutes before failing, the content already cleared the safety pass — the failure is infrastructural; re-fire the same call (the user's Step-6 approval covers retrying the identical, already-approved call — no need to re-gate). |
| 症状 | 原因→解决方法 |
|---|---|
| 视频过长/过大 | 大文件仍可上传;如果Seedance因时长拒绝,将源视频裁剪至关键片段,或设置 |
| 读取时未检测到面部/主体 | 某张联系表可能未清晰展示主体——重新调用 |
ffmpeg缺失/ | 运行环境需要本地ffmpeg;如果版本不支持 |
| 身份/镜头偏离(输出中面部、场景或镜头运动不同) | 提示词描述不够详尽。回到步骤4-5,从联系表读取结果中添加缺失的锁定细节(更多关于面部、手势、镜头运动、光线的描述)——Seedance仅保留提示词中明确提及的内容。重新渲染会产生新费用:重新执行步骤6的确认环节,获得批准后再启动。 |
| 4K被拒绝/输出降级 | 确保设置 |
| 语音/音频未保留或修改效果节拍错位 | 同时运行 |
| Seedance忽略请求的修改效果 | 强化步骤5中的时间码修改条款——使其更具体并明确时间锚点(“在N秒时触发修改”)。修改后的提示词属于新的未批准提示词,重新渲染会产生新费用:重新执行步骤6的确认环节后再启动。 |
| 长时间运行后任务失败 | 如果运行数分钟后失败,说明内容已通过安全审核——故障源于基础设施;重新发起相同调用(用户在步骤6的批准涵盖重试相同的已批准调用——无需重新确认)。 |
One-liner recap
一句话总结
Extract EVERY frame locally with ffmpeg → tile into 25-per-sheet (5×5) contact sheets and analyze every sheet → understand the audio (transcribe speech + analyze music/SFX/ambience) → write an exhaustive Seedance-faithful prompt that locks the original and time-codes the change to the right audio beat → agreement gate: show the prompt + recipe + and wait for approval → re-render the same shot in 4K via ( / / , original as , in the prompt, no ).
estimate_costgenerate_reference_videoseedancestandard4kreference_videos@Video1negative_prompt通过本地ffmpeg提取每一帧→拼接成每25帧一张(5×5)的联系表并分析所有联系表→理解音频(转录语音+分析音乐/音效/环境音)→编写详尽的符合Seedance规范的提示词,锁定原始内容并将修改效果与正确音频节拍时间码对齐→确认环节:展示提示词+配置+并等待批准→通过重新渲染4K镜头(//,原始视频作为,提示词中使用,不传入)。
estimate_costgenerate_reference_videoseedancestandard4kreference_videos@Video1negative_prompt