gemini-omni

Compare original and translation side by side

🇺🇸

Original

English
🇨🇳

Translation

Chinese

Gemini Omni Flash (Google DeepMind)

Gemini Omni Flash (Google DeepMind)

Gemini Omni is Google DeepMind's video generation and editing model family, announced at I/O 2026. The first model, Gemini Omni Flash (
gemini-omni-flash-preview
, developer access since June 30, 2026), generates 3-10 second clips at 720p/24fps with synthesized audio via the Gemini Interactions API. Its differentiator in the OpenMontage fleet is stateful conversational editing: each generation returns an
interaction_id
, and a follow-up call with
previous_interaction_id
edits that video in place — no other wrapped provider can refine a clip without regenerating it.
OpenMontage wraps it as
gemini_omni_video
(native Gemini API, no gateway). It shares
GOOGLE_API_KEY
/
GEMINI_API_KEY
with
google_imagen
and
google_tts
— one key, three capabilities. Paid tier only: ~$0.10 per second of output video (billed as 5,792 output tokens/sec at $17.50/1M).
Gemini Omni是Google DeepMind于2026年I/O大会上发布的视频生成及编辑模型系列。首款模型Gemini Omni Flash
gemini-omni-flash-preview
,自2026年6月30日起开放开发者访问)可通过Gemini Interactions API生成3-10秒、720p/24fps的剪辑,并附带合成音频。它在OpenMontage工具集中的差异化优势是有状态对话式编辑:每次生成都会返回一个
interaction_id
,后续调用时传入
previous_interaction_id
即可原地编辑该视频——其他封装提供商无法在不重新生成的情况下优化剪辑。
OpenMontage将其封装为
gemini_omni_video
(原生Gemini API,无需网关)。它与
google_imagen
google_tts
共享
GOOGLE_API_KEY
/
GEMINI_API_KEY
——一个密钥即可实现三项功能。仅支持付费 tier:每输出一秒视频约0.10美元(按每秒5792个输出令牌计费,单价为17.50美元/百万令牌)。

When to pick it (and when not)

适用场景与不适用场景

Use it forPrefer another provider for
Iterative refinement — generate, review, then edit the same clip in layersOne-shot cinematic hero clips (→ Seedance 2.0, see
seedance-2-0
)
Editing an existing/uploaded clip (restyle, add/remove objects, change text)Clips longer than 10s or above 720p
On-screen rendered text and word-by-word text beatsSeed-reproducible generations (no seed support)
Reference-image-bound subjects/styles via prompt tagsFirst/last-frame interpolation (→
veo_video
)
Timecode-scheduled multi-beat clips from one promptNon-English narration (English only fully supported)
Route through
video_selector
for generation operations. Editing (
edit_video
) is a direct-tool operation
— call
gemini_omni_video
from the registry, because the multi-turn interaction state lives outside the selector's model.
适用场景更适合其他提供商的场景
迭代优化——生成、审阅后分层编辑同一剪辑一次性生成电影级主剪辑(→ Seedance 2.0,详见
seedance-2-0
编辑现有/已上传的剪辑(重新风格化、添加/移除对象、修改文本)时长超过10秒或分辨率高于720p的剪辑
屏幕渲染文本及逐字文本节拍可通过种子复现的生成内容(不支持种子功能)
通过提示标签绑定参考图像的主体/风格首尾帧插值(→
veo_video
通过单个提示生成带时间码调度的多节拍剪辑非英语旁白(仅全面支持英语)
生成操作需通过
video_selector
路由。编辑(
edit_video
)是直接工具操作
——需从注册表调用
gemini_omni_video
,因为多轮交互状态存储在选择器模型之外。

Generation prompting

生成提示

Describe scene + camera + lighting + motion + audio. Official example:
Continuous, unbroken handheld shot of a fluffy tabby cat sitting on a sunny windowsill, looking out into a leafy garden. The cat's tail twitches slowly, and its ears rotate slightly toward ambient noises. Sunbeams illuminate dust motes in the air.
  • Force a single shot explicitly: "In a single continuous shot," / "No scene cuts." Otherwise the model may cut between scenes.
  • Negatives go in prose — there is no
    negative_prompt
    parameter: "No dialogue," "No extra sound effects."
  • No sampler controls: system instructions, temperature, top_p, and seeds are all unsupported. The prompt is the only lever.
  • Meta-prompt for quality: "Consider micro-detail, expression and timing to create a very rich, detailed but entirely natural scene."
描述场景+镜头+灯光+运动+音频。官方示例:
手持设备拍摄的连续无间断镜头,展示一只蓬松的虎斑猫坐在洒满阳光的窗台上,望向绿植繁茂的花园。猫的尾巴缓慢抽动,耳朵随环境噪音轻微转动。阳光照亮空气中的尘埃。
  • 明确强制单镜头:“采用单个连续镜头拍摄,” / “无场景切换。”否则模型可能会切换场景。
  • 否定描述用散文形式——没有
    negative_prompt
    参数:“无对话,”“无额外音效。”
  • 无采样器控制:不支持系统指令、temperature、top_p和种子。提示是唯一的调节手段。
  • 提升质量的元提示:“考虑微细节、表情和时机,打造一个丰富、细致且完全自然的场景。”

Timecode syntax

时间码语法

Schedule beats with bracketed ranges or natural language — this maps directly onto OpenMontage scene-plan timings:
[0-3s] A person is walking [3-6s] They stop and turn around
"After 3 seconds, a woman enters the scene." / "At 5s the chorus starts in the background audio."
用括号范围或自然语言调度节拍——这直接映射到OpenMontage的场景计划时间:
[0-3s] 一个人正在行走 [3-6s] 他们停下并转身
“3秒后,一名女子进入画面。” / “5秒时,背景音频开始播放副歌。”

Audio and on-screen text

音频与屏幕文本

Audio is synthesized automatically; direct it in the prompt: "Include calm background music," "The audio is a low tinny radio broadcast in the background." Rendered text works and can be timed:
One word on the screen at a time: 'did, you, know, that, Omni, can, do, awesome, text?' Each word appears for 1s.
音频会自动合成;可在提示中指定:“加入舒缓的背景音乐,”“音频为背景中微弱的收音机广播声。”渲染文本功能可用且可定时:
屏幕上一次显示一个单词:'did, you, know, that, Omni, can, do, awesome, text?' 每个单词显示1秒。

Reference images (
<FIRST_FRAME>
/
<IMAGE_REF_N>
tags)

参考图像(
<FIRST_FRAME>
/
<IMAGE_REF_N>
标签)

Pass local images via
reference_image_paths
(they are sent in order), then bind them to roles inside the prompt with tags.
<IMAGE_REF_N>
indexes from 0 in the order supplied:
in the style of <IMAGE_REF_0> a woman <IMAGE_REF_1> is walking
[0-3s] A studio fashion sequence. Starting with woman <IMAGE_REF_0>, she is
holding <IMAGE_REF_1> [3-6s] Then we see the man <IMAGE_REF_2> holding <IMAGE_REF_3>
  • <FIRST_FRAME>
    makes an image the opening frame:
    <FIRST_FRAME> a woman is walking
    .
  • Use high-resolution images; describe the intended motion specifically rather than "make it move."
  • Say what each image is (product / character / style / background reference) — the model decides usage from context.
通过
reference_image_paths
传入本地图像(按顺序发送),然后在提示内用标签将它们绑定到角色。
<IMAGE_REF_N>
按提供顺序从0开始索引:
采用<IMAGE_REF_0>的风格,一名女子<IMAGE_REF_1>正在行走
[0-3s] 一段工作室时尚序列。以女子<IMAGE_REF_0>开场,她手持<IMAGE_REF_1> [3-6s] 随后出现男子<IMAGE_REF_2>,他手持<IMAGE_REF_3>
  • <FIRST_FRAME>
    将图像设为开场帧:
    <FIRST_FRAME> 一名女子正在行走
  • 使用高分辨率图像;具体描述预期动作,而非“让它动起来”。
  • 说明每张图像的用途(产品/角色/风格/背景参考)——模型会根据上下文决定如何使用。

Conversational editing (the differentiator)

对话式编辑(差异化优势)

Editing prompts are the opposite of generation prompts: short and surgical. Overly descriptive edit prompts cause unintended changes.
  1. Generate the base clip (subject + scene + motion). The tool returns
    interaction_id
    in its result data.
  2. Pass it back as
    previous_interaction_id
    with
    operation="edit_video"
    and describe only the delta.
  3. Append "Keep everything else the same." to pin unmentioned elements.
  4. Refine in layers — one turn for lighting, one for camera, one for action, one for audio.
Official good/bad pairs:
AvoidInstead
"In the video of the man sitting on the sofa, please add a small black cat...""Add a cat that jumps onto his lap, he begins to pet it. Keep everything else the same."
"Please remove the cell phone... and fill in the background so it looks like...""Make the phone invisible. Keep everything else the same."
Other working edit prompts: "Make this video anime" / "Put a fashionable hat on this person" / "Change the lighting to be more dramatic" / "Change the text on the sign to say 'Omni Flash'".
Gotcha —
store
:
editing via
previous_interaction_id
only works if the prior call kept the interaction server-side (
store
defaults to true in
gemini_omni_video
). Set
store=false
only for one-shot generations you will never edit.
Editing uploaded videos: pass
input_video_path
instead of
previous_interaction_id
; the tool uploads it via the Files API. Unavailable in the EEA, Switzerland, and the UK (editing generated videos works everywhere).
**编辑提示与生成提示相反:简短且精准。**过于描述性的编辑提示会导致意外更改。
  1. 生成基础剪辑(主体+场景+动作)。工具会在结果数据中返回
    interaction_id
  2. 将其作为
    previous_interaction_id
    传入,设置
    operation="edit_video"
    ,并仅描述变更部分
  3. 追加**“其他内容保持不变。”**以固定未提及的元素。
  4. 分层优化——一轮调整灯光,一轮调整镜头,一轮调整动作,一轮调整音频。
官方正反示例对比:
避免写法推荐写法
“在男子坐在沙发上的视频中,请添加一只小黑猫...”“添加一只跳到他腿上的猫,他开始抚摸它。其他内容保持不变。”
“请移除手机...并填充背景使其看起来像...”“让手机消失。其他内容保持不变。”
其他可行的编辑提示:“将此视频改为动漫风格” / “给这个人戴一顶时尚的帽子” / “将灯光调整得更具戏剧性” / “将标识上的文字改为'Omni Flash'”。
注意——
store
参数
:通过
previous_interaction_id
进行编辑仅当上一次调用在服务器端保留了交互状态时才有效(
gemini_omni_video
store
默认值为true)。仅当你永远不会编辑一次性生成的内容时,才设置
store=false
编辑上传的视频:传入
input_video_path
而非
previous_interaction_id
;工具会通过Files API上传视频。该功能在欧洲经济区、瑞士和英国不可用(编辑生成的视频在所有地区都可用)。

Hard limitations (preview)

硬性限制(预览版)

  • Output: 3-10s, 720p, 24fps, MP4 with audio; aspect ratio
    16:9
    or
    9:16
    . All output carries an invisible SynthID watermark.
  • No seed, negative prompt, temperature, top_p, or system instructions.
  • No video extension or first/last-frame interpolation; no voice editing.
  • Audio reference inputs unsupported. Video references ≤3s are accepted by the schema but not processed correctly — don't rely on them.
  • Multi-video prompting unsupported; may degrade output.
  • English fully supported; other languages untested.
  • Images of minors (EEA/CH/UK) and certain recognizable people are blocked for upload/editing.
  • 输出:3-10秒,720p,24fps,带音频的MP4;宽高比为
    16:9
    9:16
    。所有输出都带有不可见的SynthID水印。
  • 不支持种子、否定提示、temperature、top_p或系统指令。
  • 不支持视频扩展或首尾帧插值;不支持语音编辑。
  • 不支持音频参考输入。视频参考≤3秒虽能通过 schema 验证,但无法正确处理——请勿依赖此功能。
  • 不支持多视频提示;可能会降低输出质量。
  • 全面支持英语;其他语言未测试。
  • 禁止上传/编辑未成年人图像(欧洲经济区/瑞士/英国)和某些可识别人物的图像。

Sources

参考资料