GitHub radar
NVIDIA model tags who spoke when in any recording
NVIDIA's third speaker diarization model tags up to eight voices in a recording at once and cuts errors by almost a third compared to the previous version. It runs locally, no cloud or subscription needed.
Nemotron 3 Diarization is the third version of NVIDIA's speaker diarization model: it listens to an audio recording and works out who spoke and when, without transcribing the actual words. The model tells apart up to eight speakers at once and works both in streaming mode, live during a call, and on an already recorded file. Its fastest mode reacts to a voice after just 0.32 seconds of audio, so it can be used for live speaker labeling, for example in a captioning system. On the open DIHARD III benchmark the full-dataset error rate dropped by almost a third compared to NVIDIA's previous version, from 19.09% down to 12.73%. The model is about 100 megabytes and ships as a file for local use through the lightweight NeMo-Speech.cpp engine, run with one command in a terminal. It is free to use, including in commercial projects.
Why a vibe-coder should care
If you edit podcasts or videos with several people talking, this model sorts the recording by voice on its own - who spoke and when, without listening through it by hand. You can then feed those labels into any speech-to-text model and get a transcript with speaker names attached. One honest limit: the model itself does not tell you the words, only the voices and the timing, so a full transcript still needs that extra step.
How to install
Copy this and send it to your agent — Claude Code, Codex, any of them:
Set up the local speaker diarization model Nemotron 3 Diarization from NVIDIA: https://huggingface.co/nvidia/Nemotron-3-Diarization - install it via NeMo-Speech.cpp, run it on a sample multi-speaker recording and show me who spoke when.
A regular laptop with 8 GB of RAM or more - the model is about 100 MB and needs no GPU.
Open on Hugging Face▌ More finds