Wan 3.0 is Alibaba's flagship video generation model. Generates native 4K video up to 30 seconds per clip. Builds up to six connected shots inside a single generation, giving a full sequence with cuts rather than one continuous take. Keeps the same character consistent across shots and generations, and produces audio in the same pass as the picture.
Key Features
Generation Modes
- Text-to-Video — Generate a full video sequence from a text prompt alone.
- Image-to-Video — Animate a still image into motion.
- Reference-to-Video — Drive a generation from reference images, videos, or audio clips to lock in a character, product, location, or style.
Multi-Shot & Consistency
- Six connected shots — Builds up to six shots inside a single 30-second generation for a full sequence with cuts.
- Identity Lock — Keeps a character's face and features consistent from one generation to the next.
- Camera controls — Direct angles and movement directly — zoom in/out, pan, orbit, follow shot, crane shot, and more.
Audio
- Native audio — Produces audio in the same pass as the picture, rather than as a separate step.
Technical Capabilities
Details | |
|---|---|
| Inputs | Text-to-Video · Image-to-Video · Reference-to-Video (up to 10 images, 5 videos, or 5 audio clips) |
| Resolution | 480p, 720p, 1080p; clips up to 30 seconds per generation |
| Aspect Ratios | 16:9 · 9:16 · 1:1 · 4:3 · 3:4 |
Limitations
- Generation time is slow overall. A clip can take several minutes to generate, regardless of length, which limits fast iteration.
Prompting Tips
- Write one line per shot rather than one long paragraph — a multi-shot sequence follows a shot list more reliably.
- Describe lighting and framing directly instead of mood words; “low afternoon sun, shot from waist height” gives a repeatable result, “cinematic” does not.
- Keep the reference set focused — four or five images pointing at the same look guide a generation better than a dozen that pull in different directions.