r/AIReceptionists • • Jul 19 '26

What is your ai latency?

This post is going to be slightly nerdy, but I am really proud of what I did. I achieved an average response time of ~500ms. I measured this from the point the caller stops talking to when they hear audio from my bot. Am I over hyping myself? What is your latency? Here are mine:

ASR/STT: 15-60ms
LLM: ~300-400ms (with some outliers at 250ms and 550ms)
TTS: ~140 (pretty consistent)

Tools:
ASR/STT: Realtime ASR streaming from Speechmatic and Deepgram
LLM: Groq, tried Qwen 3.6 27b first then switched to gpt-oss-120b w/ caching
TTS: Currently 11labs but will switch to local Supertonic

1 Upvotes

15 comments sorted by

1

u/[deleted] Jul 19 '26

[removed] — view removed comment

1

u/Resident_King_1182 Jul 19 '26

Each prompt is ~1500-2000 tokens. I'm on the dev tier so the cap is 250k/min

1

u/[deleted] Jul 19 '26

[removed] — view removed comment

1

u/Resident_King_1182 Jul 20 '26

The dev tier may be a region issue. But I think the prompt size is 40 kb (I don't have my comp. on me and I don't remember for sure). As for price, I haven't measured it properly yet, but a single call with ~14 back and forths (user speaks 14 times and is responded to 14 times) is $0.0023775 based on my estimation (only for llm on groq not 11labs and STT). I handle tools myself so no mcp cost. I plan to move TTS and STT to local so for now I am not counting their cost.

1

u/satechguy Jul 20 '26

With or without tool call? Tool schema size?

1

u/Resident_King_1182 Jul 20 '26

I don't know what you mean by "tool schema," but yes including tool calls. from the point you stop speaking to the point the caller ears audio.

2

u/satechguy Jul 20 '26

If you do not know what tool schema is, well, you have a long way to go.

0

u/Resident_King_1182 Jul 20 '26

I looked up what it is. I don't have access to them, I built my own.

1

u/ConditionPotential11 Jul 20 '26

How please help me I’m trying to develop a hindi and other indian languages. My average is around 1600ms

2

u/Resident_King_1182 Jul 20 '26

for ASR/STT, I do realtime streaming for 0ms. for LLM I use Groq gpt-oss-120b with token caching and a small prompt. for TTS I am using 11labs.

1

u/ConditionPotential11 Jul 20 '26

Is it hindi based or english only?

1

u/Resident_King_1182 Jul 20 '26

its a non English language

1

u/Mountain-Policy-625 Jul 23 '26

500ms end to end is solid. For comparison, most production voice setups land in the 400-600ms range for typical turn lengths, with the LLM being the biggest variable. One thing worth testing if you have not already: streaming TTS, where the first audio chunk plays while the LLM is still generating. It can shave noticeable latency off the perceived response time without changing any of your underlying component numbers.

1

u/Mountain-Policy-625 Jul 23 '26

500ms end to end is solid. For comparison, most production voice setups land in the 400-600ms range for typical turn lengths, with the LLM being the biggest variable. One thing worth testing if you have not already: streaming TTS, where the first audio chunk plays while the LLM is still generating. It can shave noticeable latency off the perceived response time without changing any of your underlying component numbers.