Register and share your invite link to earn from video plays and referrals.

Cartesia
@cartesia
Architecting AI that learns and interacts like humans
4 Following    15.4K Followers
Introducing Multilingual Voices: Baklavastory has a brand voice you feel in every call. Multilingual Voices keeps it that way in every language. Born in SF, built for the world.
Introducing Multilingual Voices We partnered with Baklavastory, our neighbor in the Mission, to show what it sounds like when a business's character and warmth stay consistent, no matter who calls or what language they speak. It's easy to translate speech, but much harder to preserve a single identity across languages. Every language has different rhythms, tones, and emphasis. When you change the language, identity tends to get lost in that shift, and your voice ends up sounding like someone else. With Multilingual Voices, pick a voice and keep one brand identity in every market you serve:
Show more
Been talking with @krandiash about evaling voice for 2 years now: some wisdom from what they’ve invested in it 👇
We regret to inform our AI researcher Ronald that AI has replaced him. Sorry Ronald. Here's our team trying to tell his real voice from his cloned one. They were not good at it.
Show more
We're hosting a workshop with @dograh_hq on how to build, host, and scale low-latency multilingual voice agents on open source infra. Join @ZubinPratap and @a6kme on Tuesday 15 Sept at 5pm PT. Register here:
Show more
Come join @cartesia and Dograh AI for a a practical, hands-on session on how to build, host, and scale low-latency multilingual AI voice agents using open-source infrastructure. ​We’ll show you the analyses and trade-offs needed for production-grade architecture. You'll learn about ultra-fast multilingual TTS, STT with built-in turn detection and flexible workflow orchestration. Tuesday, 15 Sept at 5pm Pacific Time. There are still a few free spots left so register ASAP at: @a6kme and I look forward to seeing you there!
Show more
At Cartesia, there’s no day quite like launch day 🎉 Tuning in from our new office in the Mission. Join us for the next one:
Our latest model continues to top benchmarks: Sonic-3.6 is #1# on @datapointai's TTS bench across all eight customer support categories including empathy, spell-outs, repairs and disfluency. Ranked by 42K blind votes:
Show more
today, we're releasing the largest open-source human audio preferences dataset, focused on the customer support use-case - 300K+ annotations by real people - 15 SOTA TTS models ranked (Sonic 3.6, Grok TTS, Simba 3.2, Eleven Labs v3) - 8 categories (IVR menus, empathy, escalations, refunds etc) dataset + benchmark + frontier plot below:
Show more
Sonic-3.6 is #1# on @voicearena_ai for US English and Hindi. Benchmarks rarely agree on one leader when their methodologies differ. @voicearena_ai and @ArtificialAnlys measure quality differently, and Sonic-3.6 is at the top of both. Two methodologies converging on the same answer is a stronger signal than any single score. Sonic-3.6 is the current frontier. Cartesia's research is dedicated to continuously moving it with realtime models that have no quality-latency tradeoff at all.
Show more
Cartesia Sonic-3.6 is the #1# TTS model on the Voice Arena US English 🇺🇸 Leaderboard. It debuts at 1085 Elo, in a statistical tie with Google DeepMind's Gemini 3.1 Flash TTS, and ahead of its own Sonic-3.5, Simba 3.2 and Grok TTS. It holds the top spot as a real-time model. Sonic-3.6 is the latest TTS model from @cartesia. The model has been highly preferred among raters on @voicearena_ai. Results backed by 1,400+ blind, head-to-head listener votes across two languages 🧵
Show more
Sonic-3.6 is now generally available. In January we made a bet: stop tuning the existing paradigm, rebuild from the architecture up. How we topped our own best model in two months →
Show more
0
64
1.6K
330
Forward to community
We've spent the last two months cooking. Keep your sound on tomorrow ↓
Introducing Sonic-3.6: our most lifelike TTS yet, and another step change in naturalness across 44 languages. Just three months after our Sonic-3.5 launch, we've made fundamental model improvements based on feedback from the teams building on Sonic. We're #1# on @ArtificialAnlys across both the provider and controlled voice streaming leaderboards — best-in-class voices + the best underlying model. Available in beta today. Hear it for yourself 👇
Show more
A very happy 80th Independence Day to India, from team Cartesia! 🇮🇳 Coming closer to home with Sonic 3.6, we're making every one of India's languages more natural, expressive, and accurate: - Hindi that flows the way you speak it - Hinglish that switches without missing a beat - Naturalness gains across all 9 existing Indic languages - Plus two new languages: Urdu and Odia Every voice in this video generated by Sonic.
Show more
0
90
517
218
Forward to community
Give your agent a face and a voice worth talking to, powered by @HeyGen LiveAvatar + Cartesia. Expressive, real-time, and live in minutes via web app or API →
Show more
For voice agents, STT has to nail three things - accuracy, turn detection, and latency. If any one falls short, the experience breaks down: the agent misunderstands, interrupts, or just feels slow. We built Ink-2 to lead on all three. Here’s how it stacks up against other providers: link in comments.
Show more