ElevenLabs’ Eleven v4 ranks #1 in Artificial Analysis TTS leaderboard

Summary

ElevenLabs’ Eleven v4 has achieved the top position in the Artificial Analysis Provider Voice TTS Arena Leaderboard, leading the Provider Voice category and scoring highly in Pronunciation Robustness. This latest Text to Speech model supports over 90 languages, significantly up from the previous version's 70+. In benchmark testing, Eleven v4 earned an Elo score of 1,319, surpassing competitors like Cartesia’s Sonic 3.6 and Google’s Gemini 3.8 Flash TTS, while also demonstrating the highest pronunciation robustness score of 91.7%. As leading TTS models are now evaluated through blind user votes and robust human reviews, Eleven v4 stands out for its multilingual capabilities and superior performance across multiple categories, furthering its appeal in global voice generation applications.

Analysis

Google: Google advances AI capabilities through its Gemini model family, which includes dedicated text-to-speech variants like Gemini 3.8 Flash TTS. This model achieves high rankings in controlled voice and pronunciation benchmarks, competing directly with specialized TTS providers. Google's efforts integrate voice technology into broader multimodal AI systems for diverse use cases. Alibaba: Alibaba develops AI models under the Qwen series, including audio and speech-focused variants such as Qwen-Audio-3.1-TTS-Plus. Its TTS offering leads in controlled voice evaluations where models use identical custom voices for comparison. The company contributes to the competitive landscape of high-performance speech synthesis tools. Cartesia: Cartesia builds AI-powered voice models including the Sonic series for text-to-speech generation. Its Sonic 3.6 model ranks as a top competitor in provider voice evaluations but places behind Eleven v4 in overall leaderboard standings. The company focuses on creating natural-sounding voices suitable for real-world applications such as entertainment and interactive systems. ElevenLabs: ElevenLabs develops advanced AI voice synthesis and text-to-speech technologies. Its latest Eleven v4 model has secured the top position on the Artificial Analysis Provider Voice TTS Arena Leaderboard and Pronunciation Robustness benchmark while ranking second in Controlled Voice. This release expands language support to over 90 languages and demonstrates strong performance across customer service, assistants, knowledge sharing, and entertainment categories. Benchmark Leadership: Leading TTS models are now ranked through blind user votes and human-reviewed pronunciation tests across categories like customer service and entertainment. Evaluation Standards: Pronunciation robustness assessments measure accuracy on challenging text elements including shorthand, context, standalone terms, and exact sequences. Multilingual Expansion: Recent TTS releases have broadened language coverage to support global applications in voice generation.

Categories

aitechmachine_learning
View Original Tweet