GitHub radar

A free model that works out who's speaking, and when

Pyannote released a free model that figures out who spoke when in an audio recording, without manual labeling. On public benchmarks, it is noticeably more accurate than the previous 3.1 version.

01pyannote/speaker-diarization-community-1 2.4k5503k downloads/mospeaker-diarization

Pyannote builds open models for speech analysis and is already known to teams who produce podcasts and video. The new Community-1 model listens to a recording and marks who is speaking and when, sorting the conversation by voice before any transcription happens. Compared to the previous 3.1 version, errors dropped noticeably: on the AISHELL-4 benchmark the error rate fell from 12.2% to 11.7%, and on the AliMeeting recording it fell from 24.5% to 20.3%. The model can take the number of speakers as a hint if you already know it, and it produces results without an internet connection, running entirely on your own machine. Access to the weights is free under the CC-BY-4.0 license, but the download is gated: before downloading, you need to fill out a short form on the page and get a Hugging Face token.

Why a vibe-coder should care

For a podcast or a video with several people, this saves the most tedious part of editing - manually marking who speaks where. The gap with the previous version is not dramatic but consistent: on recordings with several participants, the model confuses voices noticeably less often. There is a small catch: the download is gated - before running it, you need to accept the terms on the model page and set up a Hugging Face token, even though the weights themselves are free.

How to install

Copy this and send it to your agent — Claude Code, Codex, any of them:

Set up Pyannote's speaker-labeling model for me: https://huggingface.co/pyannote/speaker-diarization-community-1 - get a Hugging Face token, install pyannote.audio, and run it on this recording (I'll give you the file), showing who spoke when.

Runs on a regular computer without a GPU - it processes on the CPU by default, a GPU just speeds it up. You need Python and the pyannote.audio library.

Open on Hugging Face