Wan 2.7 is Alibaba's latest flagship video model that generates up to 1080p clips from text, images, or multi-reference inputs for up to 15 seconds for text or image-to-video, and up to 10 seconds for video or reference-to-video.
Wan 2.7 Modes
- Image-to-Video
- Text-to-Video
- Reference-to-Video
- Edit Video
Key Features
Reference-to-Video
- Multi-Reference Consistency: Feed multiple reference images or videos to lock subject identity.
- Multi-Shot Segmentation: Generate multiple shots in one pass from a single prompt.
Image-to-Video
- First & Last Frame: Set start and end frames; model fills the transition.
- Video Continuation: Continue from an existing 2–10s video clip.
Video Editing
- Text-Guided Video Edit: Edit existing videos with plain-language instructions.
- Audio Preservation: Keep original audio or let the model regenerate it.
- Reference-Based Edit: Supply a reference image to guide the video edit style.
Audio
- Audio Sync: Drive video with WAV/MP3 audio input.
- Auto Background Audio: Model generates matching background music when no audio is provided.
Image Generation & Editing
- Text-to-Image: Generate images from text.
- Multi-Image Editing: Edit up to 4 reference images with ordered prompt references.
Technical Capabilities
Details | |
|---|---|
| Inputs | Text-to-Video · Image-to-Video · Reference-to-Video · Video Editing · Audio input (WAV, MP3) |
| Resolution | 720p · 1080p |
| Aspect Ratios | 16:9 · 9:16 · 1:1 · 4:3 · 3:4 |
| Duration | 2–15s (T2V, I2V) · 2–10s (Reference-to-Video, Video Editing) |
| Limits | Up to 5 reference images or videos per reference-to-video generation · Up to 4 reference images for image editing · Source video max 10s / 100 MB for video editing · Audio input max 30s / 15 MB |
Limitations
- Reference-to-Video capped at 10s: Reference-to-video and video editing modes max out at 10 seconds, vs 15 seconds for T2V and I2V.
- No aspect ratio on I2V: Image-to-video endpoint does not expose an aspect ratio parameter; output matches input image.
- Video edit input limit: Source video for editing must be 2–10s and under 100 MB (MP4 or MOV only).
Prompting Tips
Prompt Formula: Subject + Action + Environment + Camera + Lighting + Style + Motion + Output Intent
- Lead with camera instructions and explicit sequence logic.
- Use concrete details: materials, surfaces, props, specific motion verbs. Avoid vague terms like ‘dynamic’.
- Separate primary action from secondary environmental motion.
- One scene per prompt; prioritize one visual goal over spectacle.
- For multi-reference: assign clear roles (identity, style, composition) to each image. Don't overload one.
- Number scenes for multi-shot; keep constants explicit, vary one element per shot.
- Specify exact on-screen text in quotes; use hex color codes for precision.
- Prompt expansion is on by default — disable it if your prompt is already detailed and you want literal control.
- Use negative_prompt to suppress blur, low quality, or unwanted artifacts (max 500 chars).
- Start broad, then iterate: preserve working elements and vary one thing at a time (camera, lighting, etc.).
- Convert strong prompts into reusable templates with fixed subjects, cameras, and lighting.