LTX 2.5 Fast is a text-to-video and image-to-video model built for quick iteration — generate up to 4K resolution, up to 20 seconds, with native synchronized audio and start/end frame control.
Modes
- Text-to-Video
- Image-to-Video
Key Features
- Native audio: Generates synchronized audio along with the video, no separate voiceover/music step needed.
- Start/end frame control: Set the first and last frame and let the model fill in the transition.
- Built for speed: Optimized for fast turnaround — best when you want to iterate quickly on a concept before committing to a slower, higher-fidelity pass (see LTX 2.5 Pro).
- Up to 4K: Supports resolutions up to 4K.
- Up to 20s: Supports durations up to 20 seconds.
Technical Capabilities
Spec | Details |
|---|---|
| Inputs | Text-to-Video · Image-to-Video |
| Resolution | Up to 4K |
| Duration | Up to 20s |
| Audio | Native synchronized audio generation |
Limitations
- Frame rate / duration / resolution interaction: 48/50fps output is capped at 720p/1080p and videos of 10 seconds or less. Videos longer than 10 seconds need to be generated at 24 or 25fps.
Where can I use it?
Available under Video Generation in the AI Toolkit, open to all users with no plan restriction.