Published: June 30, 2026
Gemini Omni
Video Model
Gemini Omni Flash is Google's multimodal video generator that turns text, images, audio, and video into short clips with physics-aware motion.
It is meant to be the replacement for Veo — specializing in video editing but also capable of generating videos from text and/or references.
Omni is first-and-foremost a video editor — it excels at video-to-video generations.
Variants at a Glance
Type |
Modalities |
Max Res. |
Talking Points |
|
|---|---|---|---|---|
| Google Omni Flash | Video | Text / Image / Reference → Video | 720p | Accepts video, image, and audio files as inputs |
Note: upscaling to 1080p and 4K is TBD, not yet available.
Key Features
Audio & Output
- Native audio: Omni natively creates audio (such as perfectly synced lip movement and environmental footsteps) alongside the video.
Inputs
- Four input modes: Text-to-video; image-to-video; audio-to-video; and reference-to-video.
- Multimodal references: Combine up to 3 images, 1 video clip, and 3 audio files in one generation; reference them in-prompt as @Image1 / @Video1 / @Audio1 or conversationally as "use the first image as..".
Video Edit
- Director-level camera control: Switching camera style to handheld movement, changing camera positions and POV all controlled via prompt.
- Sketch-to-Direct: The model excels at understanding annotations and sketch lines. You can draw a camera path, or cross out an object and the model understands you want to remove the object from the video.
- Editing & extension: Omni excels at editing existing video footage. You can tell the model to remove an object, replace a character, change the lighting, whatever you can think of.
- Camera Controls: The model excels at spatial understanding. For example, you can use an existing video and prompt the model to "rotate the camera so we see the subject from the left".
Technical Capabilities
Details |
|
|---|---|
| Modalities | Text-to-video, image-to-video, reference-to-video (images + video + audio) |
| Duration | 3-10 seconds |
| Resolution / quality | Standard, up to 720p (1080p and 4K upscale TBD) |
| Aspect ratios | 9:16, 16:9 |
| Audio | Native, synchronized — SFX, ambient, and lip-synced dialogue; on by default |
| Reference inputs | Up to 3 images, 1 video (TBD for final limits) |
| Input video formats | MP4 (TBD for final supported formats) |
Limitations
- The model is capped at 720p.
- Maximum clip length is 10 seconds.
- Voice-editing on someone else's footage is deliberately held back over deepfake risk (TBD).
Prompting Tips
- Be specific — describe camera movement, lighting, mood, and the exact actions you want.
- Label reference assets explicitly in the prompt: @Image1, @Video1, @Audio1 or "in the first image..".
- For edits, state both what to change and what to preserve.
- Start with 5-second generations to nail the style, then increase the duration.
Artlist — AI Model Reference · Internal Use Only