Nvidia Open-Sources a Model That Tells Voices Apart in Conversation
Nvidia has open-sourced Nemotron 3 Diarization, a 100-million-parameter model that can tell apart up to 8 voices and currently ranks first for accuracy among comparable systems.
Nvidia has released Nemotron 3 Diarization, a model with about 100 million parameters, and put it up on Hugging Face for anyone to use. The model figures out which participant is speaking at any given moment. It can tell apart up to eight voices at once. It also handles cases where people talk over each other. On the VoiceArena Diarization Benchmark v1, it currently ranks first, with an error rate of 14.72 percent versus 19.3 percent for the next-best system. Compared with Nvidia's previous version, the error rate dropped by an average of 41 percent across the eight scenarios tested.
What interests me here is that diarization is a narrow, technical task, not the job of a general chat agent. A 100-million-parameter model for a task like this does not look like a quality compromise, the way small general-purpose models usually do. The weights are open on Hugging Face. That means you can download the model and run it yourself, with no subscription and no calls to someone else's API. And 100 million parameters is a size an ordinary server can handle, not tens of thousands of dollars in hardware. Pair it with a speech-recognition model like Parakeet, and a voice agent could accurately label who said what in a conversation.
Source: the-decoder.com
Free course
Stop reading about AI — start building with it
The free Claude Code course: your first site, tool or game — no coding. No upsells, no cross-sells — nothing to buy here.
Start free →
Author
Evgenii Arsentev
PhD · Chief Executive Officer, digital health
Articles · Latest articles