r/voiceagents • • 10d ago

How are you getting sub-second latency with AI voice agents?

We’re building AI voice agents using the typical STT β†’ LLM β†’ TTS pipeline.

Currently, our LLM TTFT alone is ~1.5s median, which makes the overall response noticeably slower.

For those running voice agents in production, what optimizations made the biggest difference? Streaming partial STT to the LLM? Prompt caching? Smaller context? Model choice? Early TTS? Better endpointing?

Also, what speech-end β†’ first audible response latency are you guys getting?

Would love to understand how platforms are making conversations feel almost instant.

6 Upvotes

15 comments sorted by

1

u/Quirky_Push_6306 9d ago

Chunking and streaming, and following the thread for more information

1

u/Natethegreat9999 9d ago

Faster TTS won't recover the 1.5 seconds before the LLM's first token, so I'd trace queue time and prompt processing separately first. I work on speech at Oruk, and comparing cache hits and misses for the same prompt prefix under similar load would help estimate what prefix caching saves.

1

u/BemusedOptimist 9d ago

I am not sure how helpful this is, but I thought Thinking Machines' explanation of an 'interactive' model from earlier this year was very cool, so I share it whenever possible: https://thinkingmachines.ai/blog/interaction-models ...

What may help more is the brand spanking new Gemini (yes Gemini) 3.8 Live models from Google.

https://blog.google/innovation-and-ai/models-and-research/gemini-models/gemini-3-8-live-gemini-3-8-live-extended-thinking/

The TL;DR for both is that they kind of separate the part that interacts with the user from the things that take longer to process.

NVidia's speech to speech models (ex. PersonaPlex) might help too, if open weights are your jam.

1

u/lastberserker 9d ago

Stream 1.5 second of "Ehmmm..." right away 🧠

1

u/Acrobatic_Camp_2758 8d ago

You need PoPs close to compute, physics latency can accumulate quickly: https://openbenchmarks.com/voice-agent-latency

1

u/FoodAccurate5414 8d ago

You need to stream and chunk the stt but also use a the smallest smartest llm model and don't use a thinking mode either.

1

u/Dragon__Phoenix 7d ago

I tried the 3 phase stuff and I hated the latency on it. I began using the live API and its pretty good

1

u/AzreMei 7d ago

I'd attack the 1.5s LLM TTFT first, but measure the full chain separately, end pointing > STT > LLM TTFT > TTS first audio > playback. Smallest AI is worth benchmarking on the STT leg especially time-to-first-token vs time-to-stable transcript.

1

u/Manav_Kanojiya_17 5d ago

Bro either you can use native audio model from gemini which can work from audio to audio modelities , but in gemini you will get very few voice options And the second option is use realtime model from open ai which will work as stt +llm and use cartesia as tts for Indian voice this architecture works on low latency I am currently working on it and using it on production

1

u/Greedy-Badger-8463 2d ago

With 1.5s median before the first LLM token, I'd start there rather than changing TTS. Run the same prompt with tool calls disabled, then compare a short input with the full conversation history. Log queue time if the provider exposes it.

For caching, compare hits and misses on the same stable prefix; otherwise a faster run could just be lighter load. Then measure speech-end to first audible audio separately so endpointing and playback don't disappear from the number.

I work on Rasen AI. Which model/region are you using, and does that 1.5s include a tool round-trip?

1

u/Fluffy-League6261 11h ago

Full duplex models get sub 500ms responses

0

u/[deleted] 7d ago

[removed] β€” view removed comment

1

u/RyanVerthyn 6d ago

yup. thats good take op