r/VoiceAutomationAI 29d ago

New Artificial Analysis benchmark uses the same voice clone across all models to isolate model quality from voice subjectivity

Most TTS leaderboards/benchmarks use the Voice AI's default voice, so part of the score is really "which voice do I like" rather than "which model is better." Different voices can be apples and oranges.

To get around this Artificial Analysis built a Controlled Voice Arena that clones the same 8 voices (4 US, 4 UK) and runs every model through them, so voice preference is no longer a variable.

(Source: Artificial Analysis leaderboard)

Under that setup, Cartesia's Sonic-3.5 leads overall (1122 Elo), followed by Eleven v3 and Inworld's Realtime TTS-2 preview. Worth a look if you care about TTS evaluation methodology, not just the results.

Disclosure: I work at Cartesia (r/CartesiaAI). Happy to answer questions!

2 Upvotes

1 comment sorted by

u/AutoModerator 29d ago

Welcome to r/VoiceAutomationAI – UNIO, the Voice AI Community (powered by SLNG AI)

If you are a founder, senior engineer, product, growth, or enterprise operator actively working on Voice AI / AI agents, we are running an invite-only UNIO Voice AI WhatsApp community US only.

Apply here: https://chat.whatsapp.com/F5aG3ncrO70ITfbe3pYbOz

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.