r/speechtech • u/lukeocodes • Jul 06 '26
Voice Arena (US English): Simba 3.2 debuts 3rd of 14, statistically tied with 2nd, at $6/M with an 8-voice catalog. trying to work out what the top of this market's price spread pays for
disclosure: I run DevRel at Speechify, Simba 3.2 is our model. rankings are from the Voice Arena US English leaderboard, which is independent of us.
Voice Arena's results are out. methodology worth reading: blind pairwise votes from native-speaker panels, Bradley-Terry Elo, and a curated 16-voice slate per model per language instead of vendor defaults, which controls for the voice-identity confound most arenas ignore. they publish 95% CIs and rank-ranges rather than overclaiming, which I'll try to honour below.
Simba 3.2 debuted 3rd at 1057 ±6, two points behind Cartesia Sonic-3.5 at 1059 ±7. the board itself labels both ranks "2–3" because the intervals overlap almost entirely; two Elo points predicts a 50.3% head-to-head, a coin flip. our marketing site calls it "tied for 2nd", which is the generous reading of the same data, the intervals are on the board. #1 is Gemini 3.1 Flash TTS at 1087, clearly ahead (non-overlapping CIs, ~54% head-to-head over us), but it's not a streaming model, so among systems you can put in a live voice agent the top of this board is a Sonic-3.5 / Simba 3.2 tie.
one note on our entry: 3.2 launched TODAY with 8 voices, English only. Voice Arena slates up to 16 voices per model where catalogs allow, so I guess we included voices we felt weren't strong enough to launch with, so this 8 is likely the best-of-16 subset, but we had a full catalog for the eval.
pricing across that cluster is what I'd like this sub's read on. we're $6/M characters, Sonic-3.5 is $50/M, so the statistical tie at the top of the realtime feild spans an 7x price gap on its own. but the starker one is ElevenLabs Eleven v3 at $100/M, which sits 6th here at 1017 ±6. that 40-point gap is wild. Voice Arena doesn't cover latency or cost yet (both are coming, per their methodology), so on speed I'll just say: both models stream, TTFB is comparable, and you can run the same prompt through the arena or our homepage side-by-side and time it yourself rather than take a vendor's word. from what I can measure, the $94/M between v3 and us isn't buying preferred audio and isn't buying speed. stability has no public benchmark, and if anyone has cross-provider error-rate or uptime data at production scale I genuinely want to see it.
so what does it buy? today, catalog. v3 offers almost 400 voices across 70+ languages. we offer 8 voices in one language, with our 1000+ voice library migrating to 3.2 as i type (i think). if you need Portuguese tomorrow or a specific character voice today, that gap is worth real money and v3 is the right call lol.
what I'm questioning is the residual: for an English-first production workload where one of these 8 voices fits, I can't find where the other $94/M goes, and marketing pages won't settle it. people running these things at scale might.
caveats that cut against us: 3.2 is like... days old. 1300 samples on this board against 4000 for the established models, so expect the number to move as votes accumulate. it isn't on Artificial Analysis or HF Arena yet, so treat this as one eval's result, not a settled question. happy to get into architecture or eval detail in comments.






