Announcing Multilingual Text to Speech Arena Leaderboards, comparing leading TTS models across 9 languages beyond English ๐ฏ๐ต ๐จ๐ณ ๐ฎ๐ณ ๐ช๐ธ ๐ฉ๐ช ๐ซ๐ท ๐ต๐น ๐ป๐ณ ๐ธ๐ฆ
Text to Speech performance varies significantly across languages, with models that perform well in English not necessarily delivering the same pronunciation, pacing, tone, and naturalness in other languages.
We have extended the Artificial Analysis Controlled Voice Arena to 9 new languages: Japanese ๐ฏ๐ต, Mandarin Chinese ๐จ๐ณ, Hindi ๐ฎ๐ณ, Spanish ๐ช๐ธ, German ๐ฉ๐ช, French ๐ซ๐ท, Portuguese ๐ต๐น, Vietnamese ๐ป๐ณ, and Arabic ๐ธ๐ฆ.
Each model is evaluated using standardized cloned voices, with prompts written natively in each language and preference votes collected from first-language speakers. Elo scores are calculated independently for each language, allowing each leaderboard to reflect model preferences among speakers of that language.
Key results:
โค
@Cartesia's Sonic family leads 8 of the 9 new language leaderboards, with Sonic 3.6 ranking #
1# in 7 languages and Sonic 3.5 leading Portuguese.
@InworldAI's Realtime TTS-2 takes the #
1# spot in Mandarin.
โค
@ElevenLabs' Eleven v3 family ranks in the top 3 across 7 of the 9 new languages, through Eleven v3 and Eleven v3 Conversational.
โค Inworld's Realtime TTS-2 leads Mandarin at 1,185 Elo, ahead of Sonic 3.6 at 1,146 and StepFun's StepAudio 2.5 TTS at 1,130. Realtime TTS-2 also ranks #
1# in English (US).
โค More than 100,000 human preference votes have been collected across the 9 new language leaderboards, with 15 to 24 public models ranked per language.
See the per-language results below โฌ๏ธ