Register and share your invite link to earn from video plays and referrals.

Karan Goel
@krandiash
founder ceo @cartesia, derp learned @stanfordailab @mldcmu @iitdelhi
1K Following    20.9K Followers
Introducing Multilingual Voices We partnered with Baklavastory, our neighbor in the Mission, to show what it sounds like when a business's character and warmth stay consistent, no matter who calls or what language they speak. It's easy to translate speech, but much harder to preserve a single identity across languages. Every language has different rhythms, tones, and emphasis. When you change the language, identity tends to get lost in that shift, and your voice ends up sounding like someone else. With Multilingual Voices, pick a voice and keep one brand identity in every market you serve:
Show more
0
179
1.4K
277
Forward to community
We regret to inform our AI researcher Ronald that AI has replaced him. Sorry Ronald. Here's our team trying to tell his real voice from his cloned one. They were not good at it.
Show more
I’m excited to announce that @harvey has acquired @guardrails_ai ! Back when we started Guardrails, our mission was to build infra that made the AI transformation of work more reliable and secure. Our open source framework helped establish the guardrails infrastructure category, and Snowglobe, our simulation software, helped our customers save millions of dollars in operating costs. Today, legal AI is at the frontier of the knowledge work transformation, and I’m excited to carry our mission forward by bringing everything we’ve built to Harvey. I want to thank our amazing customers, the Guardrails open source community, our team, investors and my partner-in-crime @zaydsimjee for their support throughout our journey. I’m super proud of what we built together, and can’t wait to work on the frontier of legal AI with @winstonweinberg, @gabepereyra and the incredible team they built!
Show more
0
132
558
20
Forward to community
At Cartesia, there’s no day quite like launch day 🎉 Tuning in from our new office in the Mission. Join us for the next one:
Super cool to see we're #2# on this new benchmark out of the box (#1# soon) This makes Ink-2 a great fit for consumer apps that are aimed at everyday interactions (and it runs real-time)
VoiceCodeBench is now live on Vals, and GPT Live Transcribe takes the #1# spot. @besimple_ai created VoiceCodeBench, which measures how well speech-to-text models handle structured values in English workplace speech.
Show more
We’ve been testing @cartesia's new Sonic 3.6 model at @elise_ai and the early results are incredible!! Renters are rating it higher on voice naturalness than any other TTS model we've tested. At our scale this makes a big difference - better voice quality means people stay engaged longer, giving our voice agent a chance to resolve their request. Excited to keep pushing on this with the Cartesia team, we'll be publishing more results soon!
Show more
Our new TTS model Sonic-3.6 is now GA We went into preview last week as the #1# model on AA and it's been great to see early users back this up Overwhelmingly positive feedback, so much so that we graduated 3.6 to GA within a week And yes, the next model is cooking 🧑‍🍳
Show more
Sonic-3.6 is now generally available. In January we made a bet: stop tuning the existing paradigm, rebuild from the architecture up. How we topped our own best model in two months →
Show more
We've spent the last two months cooking. Keep your sound on tomorrow ↓
Ramp's the underdog when it comes to LLM routing, but their story makes much more sense than Stripe's. The leaked Stripe memo talks about OpenRouter as an entry into "intelligence pipelines," parallel to its "revenue pipeline" products. That obfuscates the basic fact: Token routing (what they call the intelligence pipeline) is about saving money—if it wasn't you would always route to Fable—while Stripe's revenue pipelines are about, well, making money. One's money out, the other money in. Stripe is just not where teams go to understand how much they're spending. That's pretty squarely Ramp territory, and I think Stripe is swimming against the current with the OpenRouter acquisition.
Show more
Cooking. What I loved most about reading the results with @_albertgu - the confidence intervals don’t even overlap 🤯
the team continues cooking 👩‍🍳 this is now an unprecedented gap on both the super competitive Provider Voices leaderboard as well as the newer Controlled Voices leaderboard (a stronger benchmark that can't be benchmaxxed, requiring truly better algorithms) Cartesia's TTS model is not only the best, it continues improving at a faster rate than the competition - Sonic 3.5 was released only 3 months ago. ig Sonic now means both the fastest model and the fastest team 💀
Show more
Cartesia's Sonic 3.6 takes the #1# spot on both the Provider Voice and Controlled Voice Artificial Analysis Speech Arena leaderboards, surpassing Speechify AI's Simba 3.2 and Alibaba's Qwen-Audio-3.0-TTS-Plus, with Sonic 3.5 holding #2# on Controlled Voice Sonic 3.6 is the latest TTS model from @cartesia, supporting 40+ languages including English, Hindi, Spanish, French, German and Japanese. In the TTS Arena, the model has been preferred by voters and has demonstrated natural speech and adaptive pacing across conversations. Key takeaways: ➤ Controlled Voice: Sonic 3.6 takes #1# on the Controlled Voice Arena with an Elo of 1,144 (+18/-18) across 1,339 appearances, ahead of Cartesia's own Sonic 3.5 at 1,099 and ElevenLabs' Eleven v3 at 1,060 ➤ Provider Voice: Sonic 3.6 also takes #1# on the Provider Voice Arena with an Elo score of 1,286 (+19/-19) based on 1,300 arena appearances, placing it ahead of Simba 3.2 at 1,238 and Qwen-Audio-3.0-TTS-Plus at 1,237 ➤ Pricing: At $49 per 1M characters via the Cartesia platform, Sonic 3.6 is the most expensive of the top three models on quality, roughly 5x Speechify AI's Simba 3.2 at $10 and nearly double Alibaba's Qwen-Audio-3.0-TTS-Plus at $27.59, though still half the price of ElevenLabs' Eleven v3 at $100 ➤ Speed: Sonic 3.6 processes 136.1 characters per second of generation time, compared to 46.7 for ElevenLabs' Eleven v3 and 24.2 for Google's Gemini 3.1 Flash TTS
Show more
At Cartesia, our #1# value is "Get the fundamentals right". This value has been with us since Day 0. It shaped the research @_albertgu led during our PhDs: rather than using the same algorithms & architectures for every problem, he asked whether they were the right tool for the job. New algorithms will help us process the raw data in domains like audio, video, health, and robotics, and SSMs were born from that observation. At @cartesia, the same value is core to how we build models. We've obsessed over every detail in building our 10-layer cake (parfait?): data, systems, algorithms, post-training & RL, evals, inference, serving, platform, product surfaces, and customer feedback loops. These foundations enable us not only to produce world-class models, but to rapidly turn research advances into production-ready systems that scale reliably across millions of interactions. As a researcher myself, it's exciting to see that we're #1# on AA. From my perspective, the most exciting result is our leading position on the Controlled Voices benchmark. It's hard to game that benchmark because @ArtificialAnlys clones the same 8 voices on every model. It truly tests the model's ability to generalize. We're both #1# and #2# on that benchmark now with Sonic-3.6 and Sonic-3.5. We're just scratching the surface. Our latest algorithms haven't fully made their way into these models yet. The work is advancing faster than ever before, and we have exciting new results in model intelligence & data efficiency coming soon. Audio is the first step - it’s one of the simplest physical signals of the world. We’re tackling multimodal models next - models that will bridge the gap between knowledge & reasoning (text), and the physical world (other domains). Our core belief is that a fundamentally sound approach that uses the right algorithms will generalize to all the data arising from the dance of atoms in our universe. From audio to video, robotics, biology, and everything else. Expect many more exciting releases from us in the coming months.
Show more
Introducing Sonic-3.6: our most lifelike TTS yet, and another step change in naturalness across 44 languages. Just three months after our Sonic-3.5 launch, we've made fundamental model improvements based on feedback from the teams building on Sonic. We're #1# on @ArtificialAnlys across both the provider and controlled voice streaming leaderboards — best-in-class voices + the best underlying model. Available in beta today. Hear it for yourself 👇
Show more
Happy 80th to 🇮🇳 Narrated by our best model yet
A very happy 80th Independence Day to India, from team Cartesia! 🇮🇳 Coming closer to home with Sonic 3.6, we're making every one of India's languages more natural, expressive, and accurate: - Hindi that flows the way you speak it - Hinglish that switches without missing a beat - Naturalness gains across all 9 existing Indic languages - Plus two new languages: Urdu and Odia Every voice in this video generated by Sonic.
Show more
Give your agent a face and a voice worth talking to, powered by @HeyGen LiveAvatar + Cartesia. Expressive, real-time, and live in minutes via web app or API →
Show more
We're hosting a fun group of tech folks for drinks w @cartesia and @LightspeedIndia in Bellandur this Friday - come by if you're around! Register at
We're partnering with leading AI natives, and it's great to see our models directly impact their business outcomes. @elise_ai and @minnasong are no exception, with an amazing mission to improve how we live using AI and voice systems, including right here in San Francisco!
Show more
We partnered with @a16z to look at data from the millions of renter conversations we handle every month. We found that AI-handled calls now last as long as human ones. Part of how we got there is by partnering with teams like @cartesia to improve latency and voice quality. We switched to Sonic 3.5, Cartesia's newest text-to-speech model, and saw nearly immediate results: a 2.9% lift in conversion and 12.2% increase in customer engagement.
Show more
A lot of custom post-training work being done today will be deprecated in favor of smarter models and continual learning 10-100x intelligence with continual learning is around the corner Compute, time and effort to heavily customize models is not a great use of resources while these capabilities are emerging. They will disrupt the assumptions people are building around today. It analogizes to adapting open source BERT variants to build models for language in 2019. LLMs disrupted this by solving the same problem through in-context learning with much less data and effort. 1. The base model will become fundamentally smarter, there are 10-100x improvements to be made in pre-training alone. 2. Continual learning algorithms fundamentally change the interface by which models are updated, how easy they are to maintain and use, and how much effort is required to adapt them to your domain and context. 3. The context and data sources from the domain, and the judgment of what good looks like will of course continue to matter. But even there, the level of data curation and feedback needed will shrink dramatically. Of course, it's not possible for folks to sit around and do nothing, but it's good to understand the challenges the current stack might be up against in the near future.
Show more