Google Announces Gemini 3.5 Transcribe
Reaches 2.6% word error rate on recorded speech, 4% word error rate in streaming speech, supports 85+ languages, up to 3 speaker diarization, and is 70% faster than the previous model (Chirp 3).
Google announces Gemini 3.5 Transcribe, their new speech recognition model that reaches 2.6% word error rate on recorded speech, 4% word error rate in streaming speech, supports 85+ languages, up to 3 speaker diarization, and is 70% faster than the previous model (Chirp 3). The model automatically adds punctuation and removes filler words, with sub-second latency.
Already used by LangChain, Vercel, LiveKit, and Pipecat, Gemini 3.5 Transcribe is now being deployed in Gboard and Chrome. The model is made accessible via two separate APIs: gemini-3.5-transcribe-live for live speech transcription, and gemini-3.5-transcribe for pre-recorded speech. Now anyone can access professional quality real-time speech transcription, without having to pay exorbitant fees for a specialty API.
Source: blog.google
Free course
Stop reading about AI — start building with it
The free Claude Code course: your first site, tool or game — no coding. No upsells, no cross-sells — nothing to buy here.
Start free →
Author
Evgenii Arsentev
PhD · Chief Executive Officer, digital health
Articles · Latest articles