NVIDIA open-sourced a 600M model that transcribes 40 languages in real-time at 80ms latency and it costs $0.
that's faster than you can blink. across mandarin, arabic, hindi, portuguese, tagalog, whatever,from a SINGLE checkpoint.
→ 17x more concurrent streams than buffered ASR on the same H100.
→ punctuation + capitalization built-in. no post-processing.
→ runs on your own GPU. no API bill
100% Open Source.