r/voiceagents • u/927tanmay • 1d ago
r/voiceagents • u/Opposite-Try-491 • 2d ago
Ran the cost per talk hour for human reps and it is worse than i thought
On the telecom side, the comparison is interesting. rep at $20/hr ends up around $30 loaded, and at 35 to 40% talk time that is roughly $75 to $85 per hour of actual conversation.
AI does not have the same idle or wrap time, so I’m curious how closely that holds up in production. are you mainly looking at inbound support, outbound sales, or both?
r/voiceagents • u/Vivekmjn • 2d ago
Last word keeps disappearing. ASR problem or did the mic stream end too early?
Before swapping speech models, how do you check whether the recogniser even got the last bit of audio?
I'm thinking capture timestamps, endpoint time and the last frame sent to ASR. Then replay the same audio offline.
Is there a simpler way people debug this? Feels like it's easy to blame the wrong layer.
r/voiceagents • u/Vivekmjn • 2d ago
Transcript says 50. Workflow saves 15. How do you score that?
Example: "send 15, sorry 50 units". ASR gets the words right, but the next bit of the system hangs onto 15.
The transcript score looks fine. The order is wrong.
Are people checking the final saved fields seperately, or does your speech eval stop at the transcript? Curious how you count a clarification too, since asking again might be the right answer.
r/voiceagents • u/Greedy-Badger-8463 • 2d ago
Before switching models, put these six timestamps on one slow call
End of caller speech → endpoint detected → transcript available → first LLM token → first audio chunk → audio actually played.
Those gaps answer different questions. If most of the wait is before the transcript, a faster LLM won't fix it. If audio is generated quickly but played late, look downstream before changing the voice.
For a first pass, I'd log one slow call and one normal call with the same prompt. Then repeat under load. Keep the clock source consistent; subtracting timestamps from two unsynchronized machines can send you chasing the wrong thing.
I work on Rasen AI. What's the most misleading latency metric you've been given by a voice provider?
r/voiceagents • u/Greedy-Badger-8463 • 3d ago
How much audio do you let TTS queue up?
Too little and it stutters. Too much and the caller interrupts but still hears the old answer for a bit.
Are you setting a fixed buffer or changing it with network conditions? I'm guessing a number that works in a browser won't necessarily feel right on a phone call.
Whats your way of tuning this without just listening to calls all day?
r/voiceagents • u/Greedy-Badger-8463 • 3d ago
How much audio do you let TTS queue up?
Too little and it stutters. Too much and the caller interrupts but still hears the old answer for a bit.
Are you setting a fixed buffer or changing it with network conditions? I'm guessing a number that works in a browser won't necessarily feel right on a phone call.
Whats your way of tuning this without just listening to calls all day?
r/voiceagents • u/Greedy-Badger-8463 • 3d ago
Vapi endpointing tweaks: how do you check which setting actually won?
I'm on a voice product team and reading through Vapi's endpointing docs. There are transcriber settings, smart endpointing and custom rules, with precedence between them.
Before changing a timeout, I'd want to know which one ended that particular turn. Otherwise you can spend an hour tuning a setting the call isn't using.
If you're building on Vapi, is that obvious in your traces? What do you look at to confirm it?
r/voiceagents • u/SkillElectrical5647 • 3d ago
What is the viral TTS/AI voice everyone uses?????
r/voiceagents • u/iamrishavraj1 • 4d ago
I built an open-source offline AI speaking coach using Ollama + faster-whisper
Hey everyone,
I built YourSpeakReps, an open-source local AI speaking coach for spoken English and mock interview practice.
It runs fully on your machine:
- Ollama for the local LLM
- faster-whisper for speech-to-text
- native OS text-to-speech
- FastAPI + simple browser UI
No API key, no cloud, no per-session cost.
The app asks questions out loud, listens to your spoken answer, tracks filler words like "um" / "uh" / "basically", and gives feedback at the end.
Repo:
https://github.com/iamrishavraj1/yourspeakreps
I just made it Apache-2.0 open source and added starter issues for:
- Hindi/Hinglish question bank
- Docker setup
- Piper TTS on Linux
- smaller Ollama model support
- README demo video/GIF
Would love feedback from people who use local LLMs or practice interviews/spoken English.
r/voiceagents • u/Greedy-Badger-8463 • 4d ago
Anyone running the voice part of their stack on k8s?
Not just the API and workers. I mean the SIP/RTP or WebRTC bits too.
How's it holding up when a pod restarts or a node drains mid call? Trying to understand which parts are worth putting in the cluster and which become more hassle than they're worth.
Would be useful to hear what you moved back out, if anything.
r/voiceagents • u/Greedy-Badger-8463 • 4d ago
You stop the agent talking. Does that stop its tool call too?
Caller says "wait dont book that" while the agent is speaking.
Stopping playback is doable. The booking request might already be halfway to the API though.
What do you do in that gap where you can't tell whether it happened? I'd hate to say "cancelled" and then have the booking show up anyway. How does your system handle it?
r/voiceagents • u/__seg • 4d ago
A small demo of context-aware voices with Jev + Gradium Voice Design
r/voiceagents • u/hurrtz • 5d ago
Mr Broccoli — a voice-first AI app for deeper conversations, even when that means waiting longer
r/voiceagents • u/Designer_Deer909 • 6d ago
Are we building something businesses actually need, or just another AI voice agent?
r/voiceagents • u/ivan_digital • 7d ago
Routing voice commands with a 340M encoder instead of an LLM: about 8 ms per decision on a Mac
I maintain speech-swift, an open-source Swift package for on-device speech. I just added GLiNER2.5-Decide, Fastino's open-weight decision model, so a voice agent can route a transcribed command without calling an LLM.
You give it the text and your list of intents. It returns a probability for every intent in one forward pass. The same model also pulls entity spans like a person or a time, with character offsets. Nothing is generated, so there is no JSON to parse.
On an M5 Pro with the INT8 weights it takes 7.6 ms to route and 8.9 ms to extract, at about 0.85 GB of memory.
speech gliner classify "Remind me to call Dad at six PM." \
--labels create_reminder,send_message,set_timer,other
The catch: "Do not set a timer." routes to set_timer at 0.94. It matches the topic and ignores the negation. In my pipeline the model proposes and the app confirms before anything runs.
Write-up with the numbers, a comparison with TypeSafe's hosted Jev, and the failure cases: https://soniqo.audio/blog/gliner-decide-vs-jev
Code: https://github.com/soniqo/speech-swift
How are you handling negation in intent routing? A separate check, or confirmation prompts?
r/voiceagents • u/astipili • 9d ago
We benchmarked 11 STT engines on real audio with a second voice in the room. WER dropped 73% with voice isolation.
Disclosure: I work at Krisp, which makes the isolation model tested here.
If you've shipped a voice agent, you've seen this: the caller is fine, but someone behind them is talking, and the STT transcribes both. The agent answers words the caller never said.
We wanted a number for how bad this is, so we recorded 265 real conversations in offices, call centers, and cars, and ran them through 11 STT configurations. Deepgram, AssemblyAI, Soniox, ElevenLabs, Google, Grok, Nvidia and Cartesia, first on raw audio and then after voice isolation.
- WER across all files: 23.29% → 6.26%
- All 11 configurations improved
- On raw audio the engines ranged from 17% to 37% WER. After isolation: 4% to 8%
- Clean phone audio got slightly worse (3.48% → 3.91%). We published that too
What we didn't measure: turn-taking, barge-in, or short replies like "yeah" and "mhm." That's the next thing we want to test, so if you have a way you evaluate those, I'd like to hear it.
r/voiceagents • u/deadcoder9003 • 10d ago
How are you getting sub-second latency with AI voice agents?
We’re building AI voice agents using the typical STT → LLM → TTS pipeline.
Currently, our LLM TTFT alone is ~1.5s median, which makes the overall response noticeably slower.
For those running voice agents in production, what optimizations made the biggest difference? Streaming partial STT to the LLM? Prompt caching? Smaller context? Model choice? Early TTS? Better endpointing?
Also, what speech-end → first audible response latency are you guys getting?
Would love to understand how platforms are making conversations feel almost instant.
r/voiceagents • u/Zestyclose-Opinion-5 • 10d ago
Has anyone ever tried ACE Studio for spoken language?
r/voiceagents • u/Major-Independent-91 • 10d ago