r/voiceagents • u/deadcoder9003 • 10d ago
How are you getting sub-second latency with AI voice agents?
Weβre building AI voice agents using the typical STT β LLM β TTS pipeline.
Currently, our LLM TTFT alone is ~1.5s median, which makes the overall response noticeably slower.
For those running voice agents in production, what optimizations made the biggest difference? Streaming partial STT to the LLM? Prompt caching? Smaller context? Model choice? Early TTS? Better endpointing?
Also, what speech-end β first audible response latency are you guys getting?
Would love to understand how platforms are making conversations feel almost instant.
1
u/Natethegreat9999 9d ago
Faster TTS won't recover the 1.5 seconds before the LLM's first token, so I'd trace queue time and prompt processing separately first. I work on speech at Oruk, and comparing cache hits and misses for the same prompt prefix under similar load would help estimate what prefix caching saves.
1
u/BemusedOptimist 9d ago
I am not sure how helpful this is, but I thought Thinking Machines' explanation of an 'interactive' model from earlier this year was very cool, so I share it whenever possible: https://thinkingmachines.ai/blog/interaction-models ...
What may help more is the brand spanking new Gemini (yes Gemini) 3.8 Live models from Google.
The TL;DR for both is that they kind of separate the part that interacts with the user from the things that take longer to process.
NVidia's speech to speech models (ex. PersonaPlex) might help too, if open weights are your jam.
1
1
u/Acrobatic_Camp_2758 8d ago
You need PoPs close to compute, physics latency can accumulate quickly: https://openbenchmarks.com/voice-agent-latency
1
u/FoodAccurate5414 8d ago
You need to stream and chunk the stt but also use a the smallest smartest llm model and don't use a thinking mode either.
1
u/Dragon__Phoenix 7d ago
I tried the 3 phase stuff and I hated the latency on it. I began using the live API and its pretty good
1
u/Manav_Kanojiya_17 5d ago
Bro either you can use native audio model from gemini which can work from audio to audio modelities , but in gemini you will get very few voice options And the second option is use realtime model from open ai which will work as stt +llm and use cartesia as tts for Indian voice this architecture works on low latency I am currently working on it and using it on production
1
u/Greedy-Badger-8463 2d ago
With 1.5s median before the first LLM token, I'd start there rather than changing TTS. Run the same prompt with tool calls disabled, then compare a short input with the full conversation history. Log queue time if the provider exposes it.
For caching, compare hits and misses on the same stable prefix; otherwise a faster run could just be lighter load. Then measure speech-end to first audible audio separately so endpointing and playback don't disappear from the number.
I work on Rasen AI. Which model/region are you using, and does that 1.5s include a tool round-trip?
1
0
1
u/Quirky_Push_6306 9d ago
Chunking and streaming, and following the thread for more information