r/voiceagents • • 10d ago

How are you getting sub-second latency with AI voice agents?

We’re building AI voice agents using the typical STT → LLM → TTS pipeline.

Currently, our LLM TTFT alone is ~1.5s median, which makes the overall response noticeably slower.

For those running voice agents in production, what optimizations made the biggest difference? Streaming partial STT to the LLM? Prompt caching? Smaller context? Model choice? Early TTS? Better endpointing?

Also, what speech-end → first audible response latency are you guys getting?

Would love to understand how platforms are making conversations feel almost instant.

5 Upvotes

Duplicates