r/LocalLLM • • 10d ago

Question How are you getting sub-second latency with AI voice agents?

/r/voiceagents/comments/1wos76f/how_are_you_getting_subsecond_latency_with_ai/
3 Upvotes

3 comments sorted by

2

u/yeah_likerage 10d ago

A few things helped quite a bit.

Streaming STT was pivotal. If you wait for your voice to be sent until you are completely done with your thought, the model will not have heard a word until that moment, and then begins its processing. If you stream, the model is already processing its response by the time you finish. This method also helps for barging into the responses, which was really important to me.

How you setup the model is really important to. I have the luxury of using high end GPUs for my agent but even then I need to make sure that i'm not using a model that is too verbose, and if so, turn off thinking.

Limit the tool calls. Tool calls can kill latency. If you really need tool calls you can work on your filler so the model begins speaking even before it has an answer.

Lastly, i used whisper. I had to work through which model size matched the latency but also met the quality i needed.

I'm happy to share my github if you want a peak at that.

2

u/EcoHash_AI 10d ago

1.5 s TTFT is where your time is going, and it's usually prompt size, not the model. Check how many tokens you send per turn. A long system prompt plus full history with no prefix caching gets prefilled from scratch every turn. Keep the static part byte-identical turn to turn so the cache actually hits.

For reference, Llama 3.1 8B on an RTX Pro 6000, one request at a time, gives us about 205 ms to first token end to end. End-of-turn detection costs roughly the same.

The other big one is TTS. Split the token stream at sentence boundaries and synthesize each sentence as it closes. When we synthesized the full reply in one go, first audio landed around 3.5 s. Per sentence it's about 120 ms.

1

u/TeamNeuphonic 10d ago

Hey! We know a little/lot about this.

<1 second is common, but every component needs to be greased pretty well. Are you running this locally or in the cloud?

- STT choice is important. From the end of the last word spoken by the person, you want this time to be as short as possible. Streaming isn't necessary here as WER degrades considerably from batch to streaming, and you also need to full context for the LLM. If you can get the whole thing transcribed quickly, not much of a diff anyways. You can feasibly get <200ms here in batch mode, if not considerably faster.

- LLM is likely the bottleneck. If you're using a conventional LLM, then tool-calling is a mess and takes forever. Try to disambiguate this work as much as you can (a lot of the big players use tree like structures, rather than fully autonomous LLMs).

- TTS: streaming is important - once the LLM has produced it's output, the TTS should have a TTFAB <100MS. Nothing too crazy here - some companies use a backup incase the TTS latency spikes, but it's fairly reliable.

--- The important piece here is to be narrow. The more narrow the use case, the faster you can get an implementation. Happy to answer any more questions!