NVIDIA Nemotron 3 Diarization is now available on ModelScope—adding live speaker attribution to existing ASR workflows without replacing the transcription model. 🎙️
🤖
👥 Processes streaming audio and returns speaker labels and timestamps for up to eight speaker slots in a single conversation.
⚡ Its end-to-end streaming architecture avoids separately combining voice activity detection, speaker embeddings, clustering, and post-processing.
🧠 The 99.2M-parameter model uses a 31-layer Transformer encoder with RoPE and builds on NVIDIA’s Streaming Sortformer architecture.
🔌 Pair it with Nemotron ASR, Parakeet, Canary, Whisper, or another ASR system to create speaker-attributed transcripts.
🏢 Designed for meetings, contact centers, clinical conversations, live captioning, media analysis, and multi-party voice agents.
🖥️ Supports NVIDIA Ampere, Hopper, and Blackwell GPUs, with inference through NeMo Speech C++.
顯示更多