Gemini Learns to Design Voices From Scratch by Description
Google DeepMind released the Gemini 3.8 Flash TTS models to build a voice from scratch by text description, and on the Hume AI benchmark they took first place for voice quality.
Google DeepMind released two new speech synthesis models - Gemini 3.8 Flash TTS and a lighter Flash-Lite TTS. Now a voice is built from scratch by a verbal description of accent and character. To copy someone else's voice exactly, a 30-second recording is enough (but only with confirmed consent from the author). On the Hume AI voice design benchmark the bigger model took first place with a score of 71.4. Both versions are already available through the Gemini API and Google AI Studio.
A voice from text description removes the need for a studio and an actor if a product needs voiceover. I connect voice models into my own projects through the API, and one limitation is interesting to me. Voice cloning is not available in several regions of the world and needs explicit consent from the author (I think it's more protection against forgery than a loophole). A first place in one benchmark still doesn't mean live sound in a long dialogue - worth checking by ear before putting the model into a product.
Source: blog.google
Free course
Stop reading about AI — start building with it
The free Claude Code course: your first site, tool or game — no coding. No upsells, no cross-sells — nothing to buy here.
Start free →
Author
Evgenii Arsentev
PhD · Chief Executive Officer, digital health
Articles · Latest articles