Nvidia releases free Nemotron 3 speaker diarization model
Nvidia has released Nemotron 3 Diarization, a free 100-million-parameter model that tracks up to eight speakers in real time, significantly advancing open-source audio transcription.
Nvidia has launched Nemotron 3 Diarization, a new open-source artificial intelligence model designed to identify individual speakers in an audio stream. The model features approximately 100 million parameters, and Nvidia has made its weights freely available to developers. Capable of processing both pre-recorded files and live audio streams, the system can distinguish between up to eight distinct speakers and successfully identify moments when multiple people are talking simultaneously.
For developers building transcription pipelines, the model can be paired with speech recognition systems such as Parakeet. This combination allows teams to generate fully transcribed text complete with anonymous speaker labels, like "speaker_2." However, the model does face limitations in challenging acoustic environments, as heavy background noise, excessive reverb, or conversations with more than eight participants will increase the system's error rates.
In terms of performance, Nemotron 3 Diarization has claimed the top spot on the VoiceArena Diarization-Bench v1. It achieved a diarization error rate of 14.72 percent, outperforming the runner-up system which sits at a 19.3 percent error rate. This benchmark is notoriously difficult because it penalizes minor misalignments during speaker transitions and counts overlapping speech. Compared to its predecessor, Streaming Sortformer, the new model reduces the error rate by an average of 41 percent across eight different test scenarios when configured with a 1.04-second buffer.
Practitioners can adjust the model's audio buffer across four distinct levels, ranging from 30.4 seconds down to 0.32 seconds. While shorter buffers enable faster real-time processing, they generally result in lower accuracy. By offering a highly accurate, lightweight, and free model, Nvidia provides developers with a powerful tool to build responsive, multi-speaker transcription services without relying on expensive proprietary APIs.
This is our own summary of reporting by The Decoder



