Published: August 6, 2026
Video Model
MiniMax H3 (Hailuo 3) is MiniMax's current flagship video model. It generates 2K video with native stereo audio in the same pass, and reads text, images, video and audio as references in one request.
The reference handling — twelve files, each addressed by position in the prompt — makes it a strong pick for character and product consistency. It caps at 15 seconds and 2K, where Kling v3 and Seedance reach 4K.
Variants at a Glance
Type |
Modalities |
Max Res. |
Talking Points |
|
|---|---|---|---|---|
| Variant | ||||
| Text to Video | Video | T2V | 2K | Prompt only, six aspect ratios. |
| Image to Video | Video | I2V | 2K | Animates a still, or pairs first and last frame. Ratio follows the image. |
| Reference to Video | Video | Multi-ref to V | 2K | Up to 12 reference files. Also the editing route. |
Key Features
Generation
- Native audio: Stereo score, dialogue, foley and room tone, timed to picture.
- Voice transfer: Carries a voice from a reference recording onto a character.
- Multi-reference input: 9 images, 3 video clips and 3 audio tracks per request, 12 files total.
Direction & Editing
- Reference addressing: Assets cited by position — Image 1, Video 2, Audio 1 — each given a job.
- Localized editing: Swap a product, rewrite signage, relight day to night, replace a line.
- Typography: Legible type, credits and brand marks, plus animated UI, HUDs and posters.
Technical Capabilities
Details |
|
|---|---|
| Quality / Resolution | 2K only — 1440px on the short edge, roughly 3.7MP on wider ratios |
| Duration | 5–15 seconds at 24 fps |
| Aspect Ratios | 21:9, 16:9, 4:3, 1:1, 3:4, 9:16, plus adaptive on Reference to Video. Image to Video follows the input image. |
| Input Modalities | Text, image, video, audio — stereo audio out with the picture |
| Reference Limits | 9 images, 3 clips, 3 audio tracks; 12 files max, clips 2–15s each |
| Prompt Length | Up to 7,000 characters |
Limitations
- Locked at 2K and 15 seconds — wrong pick when delivery needs a 4K master.
- Faces degrade at distance; keep frontal faces in close-up, rear angles in wides.
- Adds subtitles, watermarks and stray non-Latin characters unless told not to.
- Drifts into slideshow pacing and soft dissolves without a timed beat structure.
Prompting Tips
- Give every reference a job: "Image 1 for location and texture, Image 2 for the talent."
- Storyboard in the prompt with timecoded blocks — [0-2s], [2-4s] — to hold pacing.
- Art-direct the audio like a shot: instrumentation, where the beat lands, specific foley.
- Negative direction lands hard: "no soft dissolves," "no subtitles or watermarks."
- Lock identity by listing what to preserve — hair, garment, prop — not the character's name.
- For edits, pair each change with what must stay stable, or you get a regenerated shot.