Published: August 27, 2026
Gemini Omni 1.1
Video Model
UPDATE — 2026-08-27 |
|---|
|
New in this release:
Still deferred: longer clips (targeting a future release) and audio references Note: all outputs carry invisible SynthID + C2PA metadata |
Gemini Omni Flash is Google's multimodal video generator that turns text, images, audio, and video into short clips with physics-aware motion.
It is meant to be the replacement for Veo — specializing in video editing but also capable of generating videos from text and/or references.
Omni is first-and-foremost a video editor — it excels at video-to-video generations.
Key Features
Audio & Output
- Native audio: Omni natively creates audio (such as perfectly synced lip movement and environmental footsteps) alongside the video.
- Resolution range: Generate at 360p, 720p, 1080p, or 4K.
Inputs
- Three input modes: Text-to-video; image-to-video; reference-to-video (note: audio-to-video coming soon).
- Multimodal references: Combine up to 5 images and 3 video clips in one generation; reference them in-prompt as @Image1 / @Video1 or conversationally as "use the first image as..".
- Frame interpolation: Give the model an explicit first and last frame and it generates the motion between them.
Video Edit
- Director-level camera control: Switching camera style to handheld movement, changing camera positions and POV all controlled via prompt.
- Sketch-to-Direct: The model excels at understanding annotations and sketch lines. You can draw a camera path, or cross out an object and the model understands you want to remove the object from the video.
- Editing & extension: Omni excels at editing existing video footage. You can tell the model to remove an object, replace a character, change the lighting, whatever you can think of.
- Clip extension: Extend an existing video by up to 10 seconds per pass; source clips up to 30 seconds.
- Camera Controls: The model excels at spatial understanding. For example, you can use an existing video and prompt the model to "rotate the camera so we see the subject from the left".
Technical Capabilities
Details |
|
|---|---|
| Modalities | Text-to-video, image-to-video, reference-to-video |
| Duration | 3-10 seconds |
| Resolution / quality | 360p, 720p, 1080p, 4K |
| Aspect ratios | 9:16, 16:9 |
| Audio | Native, synchronized — SFX, ambient, and lip-synced dialogue; on by default |
| Reference inputs | Up to 5 reference images and 3 reference videos |
| Video to video inputs | 1 video up to 30 seconds |
| Clip extension | Adds up to 10 seconds per pass |
| Frame control | Explicit first and last frame |
Limitations
- Maximum generated clip length is 10 seconds — longer clips are deferred to a future release.
- Clip extension adds up to 10 seconds at a time, and the source video cannot exceed 30 seconds.
- Audio references are not supported.
Prompting Tips
- Be specific — describe camera movement, lighting, mood, and the exact actions you want.
- Label reference assets explicitly in the prompt: @Image1, @Video1 or "in the first image..".
- For edits, state both what to change and what to preserve.
- Start with 5-second generations to nail the style, then increase the duration.
- When using first and last frame, keep the two frames close in framing and lighting — large jumps produce unstable motion.
Artlist — AI Model Reference · Internal Use Only