Exhibitor login
AI Insider 23 September 2026

NVIDIA Nemotron 3: Leadership in Diarization Technology

NVIDIA Nemotron 3: Leadership in Diarization Technology

NVIDIA's new speech recognition model, Nemotron 3, has claimed its spot at the top of the Voice Arena, with a Diarization Error Rate (DER) of 14.72%. This evaluation included both live and recorded conversations and took place within a comparative study of 12 systems across 139 English-language conversations, totaling approximately 22 hours of audio. By utilizing timestamps per speaker channel, the model enhances the usability of search functions and summaries while supporting up to eight speakers in a single recording.

With results that are significantly better than most other systems, Nemotron 3 shows a relative improvement of 24 percent compared to the second-best system, which achieved a DER of 19.3 percent. When tested under various conditions, such as in-person and online recordings, the model demonstrates consistent performance, with an average drop in DER of 41 percent across eight benchmarks. Despite some nuances in the two-speaker subset, where the baseline sometimes performs better, it overall exceeds the expected results.

The model employs different configurations with input buffers ranging from 0.32 to 30.4 seconds, optimizing performance for real-world applications. With an impressive throughput of 15,113x at a batch size of 32 on an NVIDIA RTX PRO 5000, Nemotron 3 is a strong candidate for advanced speech recognition and processing. The technology converts audio into Mel-spectrogram features and supports common formats like .wav and .mp3, and is compatible with the latest NVIDIA GPU generations.

Additionally, NVIDIA is introducing a new pre-diarized transcription API in Argmax Pro SDK 3, making it easier for developers to pre-separate speakers before speech recognition occurs. This is particularly useful in more complex conversations with overlap, but the model also has its limitations, such as reduced performance in longer recordings or in high ambient noise. The training data was diverse, with a combination of public and licensed conversations, contributing to the broader applicability of the model.

Read the full article from AI Insider.