Register and share your invite link to earn from video plays and referrals.

Albert Gu
@_albertgu
assistant prof @mldcmu. chief scientist @cartesia. leading the ssm revolution.
78 Following    21.8K Followers
the team continues cooking 👩‍🍳 this is now an unprecedented gap on both the super competitive Provider Voices leaderboard as well as the newer Controlled Voices leaderboard (a stronger benchmark that can't be benchmaxxed, requiring truly better algorithms) Cartesia's TTS model is not only the best, it continues improving at a faster rate than the competition - Sonic 3.5 was released only 3 months ago. ig Sonic now means both the fastest model and the fastest team 💀
Show more
Within the span of a week, we launched streaming TTS (text-to-speech) and STT (speech-to-text) models that topped the leaderboards. I'm incredibly proud of the research team for their relentless pursuit of improvement, which have unlocked new state-of-the-art audio models on the Pareto frontier of speed and quality. As a research problem, speech requires fusing both text and audio and is the gateway to general multimodal models. We built Sonic-3.5 and Ink-2 from the ground up, developing multiple innovations along the way in a direction that will scale to general real-time intelligence. I've personally been deeply involved in building these models and more; it's been a blast working with the incredibly talented research team here @cartesia, and I can't wait to show the world what's coming next :)
Show more