
Generates high-quality video directly from text descriptions. Utilizing the LTX-2.5 22B distilled Transformer and Gemma 4 12B text encoder, this workflow accurately renders scene details and motion logic from prompts. It features a Latent Upscale module to enhance visual fidelity and supports native audio-video synchronization. Output is at 24fps, and generating a 10-second video takes approximately 1 minute and 50 seconds.
No creation yet
