Register and share your invite link to earn from video plays and referrals.

kwindla
@kwindla
Infrastructure and developer tools for real-time voice, video, and AI. @trydaily // ᓚᘏᗢ // @pipecat_ai
Joined September 2008
3.9K Following    15.9K Followers
Deepgram has been shipping realtime speech models that deliver both very high accuracy and very low latency since before we had LLMs we could use to build realtime voice agents! Scott played me clips of their new Flux TTS model while they were training it, and I've talked a lot to the Deepgram team about what we want a next-generation text-to-speech model to be good at. For enterprise use cases, we need our models to vocalize text very, very reliably. The model can't skip numbers in an alpha-numeric string, can't mispronounce common words when they appear in an unusual sequence like a brand name or a medical term, and can't stumble over acronyms. (All of these are widespread issues with many text-to-speech models.) We also want voice models to be highly expressive and steerable. And voice agents sessions are long, multi-turn conversations. We need both pronunciation choices and expressiveness to be consistent across conversation turns. Deepgram delivered on all of these must-haves with their new Flux TTS model. I'm really impressed.
Show more
Congrats to @DeepgramAI on the GA launch of Flux TTS, now live in Pipecat 🎉 @JonPTaylor looks at how Flux TTS delivers a more consistent voice experience. Flux reads the whole conversation, not just the next line — adaptive tone, consistent pronunciation, clean interruption handling natively (no SSML markup, no style tags!) all at sub-200ms latency. ➡️
Show more