Published: October 4, 2026
Video Model — Lipsync
Sync Lipsync v3 is Sync's flagship lipsync model. It re-renders the mouth to match a new audio track, on existing footage or a single still image. Audio is mandatory in both modes.
It processes the whole shot at once, so it holds sync on profile angles and over obstructed mouths, where Lipsync v2 Pro slips.
Variants at a Glance
Type |
Modalities |
Max Res. |
Talking Points |
|
|---|---|---|---|---|
| Variant | ||||
| Video-to-Video | Video | Video-to-Video + audio | 4K | Re-syncs existing footage to new audio. |
| Image-to-Video | Video | Image-to-Video + audio | — | Makes one still image speak the audio. |
Key Features
Video-to-Video
- Extreme Angles: Profile, over-the-shoulder and other non-frontal shots.
- Obstruction Handling: Works around hands, mics or objects in front of the mouth.
Image-to-Video
- Any Style: Photos, illustrations and animated frames; output runs as long as the audio.
Technical Capabilities
Details |
|
|---|---|
| Resolution / Aspect Ratio | Up to 4K; video-to-video output follows the input video's dimensions |
| Inputs / Output | Video, image (JPEG, PNG, WebP) and audio (mp3, wav, m4a, aac, ogg) in; MP4 out |
Limitations
- No resolution or aspect ratio control: video-to-video output matches the input video.
- Needs a visible mouth. Tiny faces and cluttered frames hurt quality.
- If lengths differ, video-to-video defaults to the shorter one. Audio-driven only, no text prompt.
Prompting Tips
- Use clean audio. Music and background noise muddy the sync.
- Match audio and video length to avoid the trim-to-shortest default.