Voice agents still don’t understand who’s speaking to them. That’s a huge gap compared with humans, hidden by all the “phone-call” demos. But that changes today!
NVIDIA is open-sourcing Nemotron 3 Diarization: a model that can reliably track speakers in live conversations, under a commercial-friendly license! In my tests, the quality is really good with one-second speech chunks. So we can use it for voice agents!
I tested it with Reachy Mini and speech-to-speech running on a DGX Spark. It’s super fun to see the robot notice a new voice, ask for a name, and remember it.
The model has day-zero integration with Transformers!
Kudos to the NVIDIA team for shipping useful tools for the whole community!