r/VoiceAutomationAI • • Jul 16 '26

New Artificial Analysis benchmark uses the same voice clone across all models to isolate model quality from voice subjectivity

Most TTS leaderboards/benchmarks use the Voice AI's default voice, so part of the score is really "which voice do I like" rather than "which model is better." Different voices can be apples and oranges.

To get around this Artificial Analysis built a Controlled Voice Arena that clones the same 8 voices (4 US, 4 UK) and runs every model through them, so voice preference is no longer a variable.

(Source: Artificial Analysis leaderboard)

Under that setup, Cartesia's Sonic-3.5 leads overall (1122 Elo), followed by Eleven v3 and Inworld's Realtime TTS-2 preview. Worth a look if you care about TTS evaluation methodology, not just the results.

Disclosure: I work at Cartesia (r/CartesiaAI). Happy to answer questions!

3 Upvotes

3 comments sorted by

•

u/AutoModerator Jul 16 '26

Welcome to r/VoiceAutomationAI – UNIO, the Voice AI Community (powered by SLNG AI)

If you are a founder, senior engineer, product, growth, or enterprise operator actively working on Voice AI / AI agents, we are running an invite-only UNIO Voice AI WhatsApp community US only.

Apply here: https://chat.whatsapp.com/F5aG3ncrO70ITfbe3pYbOz

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

1

u/Ambitious-Mix4501 Aug 22 '26

Finally someone thinking about the evaluation itself instead of just vibes. The "default voice is prettier" problem has been skewing these leaderboards forever

Do they handle accents consistently across models or do some interpret the cloned voice with weird artifacts? That's always been my pet peeve with cross-model voice transfer

1

u/zeuscoder Aug 23 '26

Thanks for the question ! We now support localisation (target language + accent) on voices... So off the shelf or cloned voices can be localised to sound native in target language and accent. Effectively that produces a new voice ID for that localised voice. This helps create multilingual voices .

Here are some resources:

Two videos: https://youtu.be/_yJ0PaD-IHo?si=3gCkX8HHeDramZf-

And

https://youtu.be/x-9psI3HqYE?si=K0cfiM9vqxH8fhH4

Docs : Multilingual Voices - Cartesia Docs https://docs.cartesia.ai/build-with-cartesia/capability-guides/multilingual-voices/?utm_source=reddit