Earlier today we introduced Gemini 3.5 Transcribe, our latest text-to-speech model. But, what does this actually mean for your projects?
We built this app in
@GoogleAIStudio to demonstrate just how much smarter 3.5 Transcribe is. When streaming live audio simultaneously through two parallel transcription pipeline modes, you can see how smart transcription automatically strips out disfluencies (ums, ahs) and condenses long, rambling thoughts while verbatim keeps exact match transcription live: