Updated: October 05, 2026
Video Model — generation
Grok Imagine 1.5 is xAI's latest video generation model, released May 2026. It generates video with native audio — dialogue, ambient sounds, and sound effects — in a single pass.
It also lets you prompt multiple shots in one generation, and supports both text-to-video and image-to-video.
Variants at a Glance
Variant |
Type |
Modalities |
Max Res. |
Highlights |
|---|---|---|---|---|
| Variant | ||||
| Grok Imagine 1.5 | Video | Text-to-Video, Image-to-Video | 1080p | Full model |
| Grok Imagine 1.5 Lite | Video | Text-to-Video, Image-to-Video | 1080p | Quicker, cheaper version |
Key Features
Video Generation
- Multi-shot: Prompt for multiple shots in a single generation.
- Native audio generation: Produces synchronized audio (dialogue, ambient sounds, effects) alongside the video in one generation. No separate audio step needed.
- Text-to-video: Generate a video with audio from a text prompt alone.
- Image-to-video: Animate a still image with a text prompt. The output preserves the look of the original image while adding motion and audio.
- Configurable duration & resolution: Set duration from 1–15 seconds and choose between 480p, 720p and 1080p per request.
Technical Capabilities
Details |
|
|---|---|
| Resolution | 480p, 720p, 1080p |
| Duration | Up to 15 seconds |
| Audio | Natively generated with the video |
Limitations
- 15-second max duration: Fine for social clips, limiting for longer-form content.
- Aspect ratio: In image-to-video, it's determined by the uploaded image.
Prompting Tips
- Name the audio you want. Since audio is generated natively, include sound cues: "birds chirping," "crowd murmuring," "footsteps on gravel."
- Start at 480p for iteration, then switch to 720p for final output. 480p is faster and cheaper for testing prompts.