Nvidia open-sources Nemotron 3 for real-time speaker tracking

2 hours ago 2



Nvidia just made it a lot easier to figure out who’s talking. The company released Nemotron 3 Diarization, an open-weight model built specifically to solve one of the most annoying problems in audio processing: accurately tracking which speaker said what, and when they said it, in real time. The model launched on September 23 under Nvidia’s OpenMDW 1.1 license, making it freely available on Hugging Face and through the company’s NeMo framework. It’s also accessible via inference providers like Baseten and DeepInfra, with deployment costs reportedly as low as $0.01 per audio hour in select configurations. What Nemotron 3 Diarization actually does Nemotron 3 tackles this with roughly 100 million parameters and what Nvidia calls its Streaming Sortformer architecture. The model handles up to eight speakers simultaneously, producing speaker-activity probabilities at 10 millisecond resolution. That’s granular enough to catch the kind of rapid-fire crosstalk that typically turns transcription software into a confused mess. The system operates in both streaming and offline modes. In streaming mode, it can run with configurable latency profiles that go as low as approximately 320 millisecon...

Read Entire Article