GitHub radar
IndexTTS 2.5 Voice cloning with a single audio file.
The Index Team from Bilibili has launched IndexTTS-2.5,an open source TTS model that can clone any voice with one reference sample and support 8 emotions.
IndexTTS-2.5 is an open source TTS model developed by the Index Team of Bilibili that can synthesize a voice with only one reference audio sample to capture the tone and color of the voice. It provides 8 emotion control modes (such as happiness, sadness, anger, etc.) controlled by prompts, speech rate control and phonetic accent control, and is applicable to Chinese, English, Japanese, Spanish and Arabic, and does not support Russian. It needs to use Python 3.10–3.11 and a GPU card with about 6GB of memory, Clone the GitHub code and execute uv sync to install.
Why a vibe-coder should care
If you do English video/podcasts, you need to generate your own voice or fix the speaker identity in different episodes. With IndexTTS-2.5, you can achieve this goal with only one reference sample, without using any cloud services. It can be used for dubbing, storytelling and constructing a local TTS system with emotional expression. Please note that IndexTTS-2.5 does not support Russian, if you need to synthesize Russian voice, please use other models.
How to install
Copy this and send it to your agent — Claude Code, Codex, any of them:
Install IndexTTS-2.5 from Index Team (Bilibili): https://huggingface.co/IndexTeam/IndexTTS-2.5 — follow the GitHub README, download the weights, and synthesize 'Hello, this is a test of voice cloning' with zero-shot voice cloning. Ask me if you need a reference audio file.
Requires an NVIDIA GPU with 6 GB of VRAM or more — a MacBook without a discrete NVIDIA GPU won't work.
Open on Hugging Face▌ More finds