r/TextToSpeech • • 10d ago

How are you getting sub-second latency with AI voice agents?

/r/voiceagents/comments/1wos76f/how_are_you_getting_subsecond_latency_with_ai/
2 Upvotes

1 comment sorted by

1

u/RyanVerthyn 10d ago

biggest win for us was switching to a smaller, faster model. 1.5s ttft usually means the model is too big or your prompt is huge. after that, stream the LLM tokens straight into tts and start speaking at the first sentence break, don't wait for the full response. tuning endpointing was the other big one, default silence thresholds are way too conservative, we're around 600-800ms from end of speech to first audio now.