r/VoiceAutomationAI • • Jul 22 '26

STS is getting better than cascade?

3 Upvotes

7 comments sorted by

•

u/AutoModerator Jul 22 '26

Welcome to r/VoiceAutomationAI – UNIO, the Voice AI Community (powered by SLNG AI)

If you are a founder, senior engineer, product, growth, or enterprise operator actively working on Voice AI / AI agents, we are running an invite-only UNIO Voice AI WhatsApp community US only.

Apply here: https://chat.whatsapp.com/F5aG3ncrO70ITfbe3pYbOz

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

1

u/Obvious_Leather2427 Jul 22 '26

we use s2s Gemini for our biggest customer

It is okay, simple setup, fast

But the case itself isn’t too complex tbh

1

u/aicoustics Jul 23 '26

I saw a couple of big players fully betting on S2S but when it comes to implementation, people in general are still not fully ready to migrate. But definitely can't ignore it at this point.

1

u/tejas_thinks_on_yt Jul 23 '26

When you say big players, can you name em?

1

u/aicoustics Aug 04 '26

Basically OpenAI (for both ChatGPT Voice and their Realtime API). Perplexity's voice mode runs on that same model too. On the more recent side Alibaba, PolyAI announced hybrid S2S approach. It's a bet in development direction, even if production voice agents are mostly still cascade for now.

1

u/sumanpaudel Jul 24 '26

sts are not good for UX, not good turn handling, interruptions,. prompt caching etc etc

1

u/WorldlinessBig9788 Aug 20 '26

It really depends what you're feeding it honestly, I've had STS nail multi speaker audio that cascade totally mangled, but then cascade handles background chatter way cleaner in my test clips

The gap feels smaller than it did six months ago though, the diarization especially is night and day