On First Partial Transcription, Gemini 3.5 Transcribe Live achieves 5.8% WER at 0.25s after detected end of speech. GPT Live Transcribe achieves 6.3% WER at 0.26s, making Gemini slightly faster and more accurate at this measurement point.
New leaderboard for audio transcription just launched and our apache 2.0 Cohere-Transcribe is at the top. This eval didn't exist when we trained the model, so its nice to see us do so well on it.
Deleted the circular financing comment from Hock Tan because Quartr's transcription picked it up wrong. I still think he could've worded it better but he said:
"that is not circular finance where we are using financing to create demand. Demand is there."
What happens when speech, transcription, and coding models work together?
This prototype demo, built using a VS Code fork, showcases how MAI-Transcribe, MAI-Voice, and MAI-Code-1-Flash can work together in a unified workflow to transform spoken instructions into working code.
Grok Voice Transcribe 2.0 just launched
It's already #1# in overall transcription accuracy at 97.4%, ahead of ElevenLabs, Gemini, and Deepgram
And it starts at just $0.10/hour
Highest accuracy at the lowest price
Cohere Transcribe is now available on @superwhisper 💬
Push to talk and get your transcription back almost instantly. With Superwhisper, you can use Transcribe offline, integrate it into your favorite apps, and recall specialist vocabulary.
NVIDIA Nemotron 3 Diarization is now available on ModelScope—adding live speaker attribution to existing ASR workflows without replacing the transcription model. 🎙️
🤖
👥 Processes streaming audio and returns speaker labels and timestamps for up to eight speaker slots in a single conversation.
⚡ Its end-to-end streaming architecture avoids separately combining voice activity detection, speaker embeddings, clustering, and post-processing.
🧠 The 99.2M-parameter model uses a 31-layer Transformer encoder with RoPE and builds on NVIDIA’s Streaming Sortformer architecture.
🔌 Pair it with Nemotron ASR, Parakeet, Canary, Whisper, or another ASR system to create speaker-attributed transcripts.
🏢 Designed for meetings, contact centers, clinical conversations, live captioning, media analysis, and multi-party voice agents.
🖥️ Supports NVIDIA Ampere, Hopper, and Blackwell GPUs, with inference through NeMo Speech C++.
A few weeks ago, I vibe-coded a live transcription tool (using Swift, Parakeet model) and it's working so amazingly well.
Allows me to spin it up from anywhere with a shortcut, talk, I see the live transcription, and hitting ENTER ends it and inserts the text into which ever input (including terminal) I may have selected.
Total game changer. This just wasn't possible a few years ago.