GitHub radar
LTX-2.5: Open Video + Audio Generation
Lightricks announced LTX-2.5, an open-weight model capable of generating synced videos & audios simultaneously, with multi-scene synthesis & diffusion decoder for higher quality output.
LTX-2.5 is an open weight video synthesis model developed by Lightricks. It can generate videos and corresponding audios based on text, image or video prompts. With version 2.5, it supports multi-scene synthesis within a single forward pass while keeping consistent characters, lights, and voices among scenes. It uses a diffusion decoder instead of vanilla VAE to produce higher quality images and Gemma 4 12B as its text encoder to increase prompt accuracy for complicated prompts. Also, there’s an automatic duration prediction module to estimate the length of clips based on textual prompts. Free for non-commercial use (<$10M annual revenue); commercial license required for others.
Why a vibe-coder should care
It’s currently the best open source solution if you want to deploy a local model to process both video and audio data simultaneously, without requiring multiple processing pipelines. Useful for those who create contents like video trailers, social media short videos, b-roll videos with background noises etc., but don’t want to pay monthly fees for cloud-based solutions. A high end GPU is required to use it. Integrates well with existing video editing processes via ComfyUI.
How to install
Copy this and send it to your agent — Claude Code, Codex, any of them:
Set up LTX-2.5 from Lightricks for video generation: https://huggingface.co/Lightricks/LTX-2.5 — install via Diffusers or ComfyUI per the README, generate a 5-second clip from 'a cat walking through a forest at sunset', and show me the result
A powerful GPU is required — 24 GB of VRAM or more recommended; CPU offloading optimizations are available for smaller setups.
Open on Hugging Face▌ More finds