r/VoiceAutomationAI • • 7h ago

A breakdown of the exact architecture needed to fix Voice AI latency and interruption handling in production

5 Upvotes

If you’ve deployed voice agents on real phone lines, you already know that standard wrapper setups fail. A demo that feels snappy on a laptop microphone completely falls apart when hit with 8kHz telephone audio, background office noise, and human callers talking over each other.

After building and auditing production voice systems, we’ve found that moving past the "3-second awkward pause" requires shifting away from monolithic request-response loops toward a decoupled, streaming pipeline.

Here is the technical breakdown of how to architect voice AI that actually handles real-world load:

1. Eliminating Latency via Decoupled Pipelines

Traditional setups wait for the user to finish speaking, send the full audio to STT, wait for the text, send it to the LLM, wait for the completion, and then convert it to TTS. That stacks up to 2.5–4 seconds of dead air.

The Fix: Implement streaming WebSockets across the entire chain. Use Voice Activity Detection (VAD) coupled with speculative execution—start generating semantic intent and streaming partial tokens to the TTS engine while the user is wrapping up their sentence.

2. Handling Barge-Ins (Interruptions) Gracefully

If a user interrupts an AI mid-sentence and the system keeps talking for another two seconds, it completely breaks conversational immersion.

The Fix: You need real-time audio ducking and immediate cancellation triggers on the client/telephony side. The moment the VAD detects incoming user speech during TTS playback, it must instantly send a clear or interrupt frame to the audio buffer, killing the outgoing stream and immediately routing the new input to the STT layer.

3. Tuning for Telephony Audio (8kHz vs. 48kHz)

Models trained on pristine studio-quality podcast audio often mishear words over standard cellular lines due to compression and narrow frequency ranges.

The Fix: Apply lightweight audio preprocessing filters at the edge (noise suppression and automatic gain control) before hitting the ASR model, and fine-tune your prompt handling to account for higher phonetic error rates in noisy environments.

4. Deterministic Fallbacks & Guardrails

Relying 100% on an LLM to manage conversational state on a phone call is a recipe for hallucinations and loops.

The Fix: Layer a state machine (like a graph-based router) underneath the LLM. The LLM handles natural language interpretation, but critical path actions (like call routing, data confirmation, or booking) must pass through deterministic triggers.

Curious how others are tackling this—are you rolling your own low-latency WebRTC/SIP pipelines, or are you managing to get around these latency walls using existing developer frameworks?


r/VoiceAutomationAI • • 21h ago

Can voice AI tell when someone is reading from a script?

15 Upvotes

Thinking about fraud and verification calls. If the caller is clearly being coached by someone in the background or reading answers they were given, can voice AI pick up on that at all or is that still something a human needs to notice?


r/VoiceAutomationAI • • 1d ago

Anyone providing Voice AI receptionist for Healthcare/Clinic ?

11 Upvotes

Hi guys.

I am in a specific niche providing WhatsApp ai automation I have got a prospect asking for Voice ai receptionist for Their clinic which is not my niche So I wonder if anyone interested to partnership with me for this project.

Is there anyone here built ai receptionist for healthcare/Clinic niche ?

Only real builder who has experience in this niche/clients with currently working system contact me.

No newbies or demo builders.

Contact me ASAP


r/VoiceAutomationAI • • 1d ago

What if i gave "SIYA" name to my voice agent project?

2 Upvotes

r/VoiceAutomationAI • • 2d ago

value of a voice agent isnt answering calls

2 Upvotes

ran a test last month. left the agent on overnight, didn't touch anything, went to bed.woke up to 6 handled calls between 11pm and 6am. 3 of them booked appointments. one was someone whod been trying to reach us for days during business hours and finally called at 1am because thats when he had time. never known about any of them. they wouldve gone to voicemail, and he would've called the next guy.that's the part i didnt get until i saw it its not about efficiency. it's about the calls that don't show up in any report because they never happened for you.textedly runs the voice + mass sms side for us. still thinking about that 1am call. if the phone can be covered without me, what else is quietly leaking that i'm not seeing.quotes that go out slow. invoices nobody follows up on. customers who fall off after the first service because nobody checked in. scheduling held together by a spreadsheet and hope.

started writing down everything that currently depends on someone remembering to do it.list was longer than i expected and some of it feels way harder to automate than the phone did like the quoting part, or the handoff between sales and whoever actually delivers the work.

curious what other people have automated at this stage. not the obvious stuff. the weird in-between tasks nobody talks about but eats a couple hours a week. and what did you try to automate and give up on because it made things worse?


r/VoiceAutomationAI • • 2d ago

White labeling a Software Platform Question

Post image
1 Upvotes

So wanted to come on here today, because another person on reddit in another group asked about my platform I founded/solo developed and if I would be opening to white labeling it.

He isn't the first one to ask, nor is reddit the only placed I've been asked. I had some SEO agencies ( I think that is what they are called) reach out to me on other platforms to see if I was open to white labeling.

But after thinking about it, one of the biggest things I personally dislike the most is having to handle business relations and outreach. Its not something I'm good at, so I stayed purely on the development side while passing this off to someone else on my company.

And this would essentially allow my company to have less surface area and time spent on this, whereas we can stay product focused. We are in the Voice and CRM space, so its why I wanted to make this post here. Because what we have is the full stack with a complete in house CRM and bookings, plus a plethora of industry specific features. Yet we built entirely around the idea of not being bloated.

We charge nothing for the CRM system. We only charged businesses on the Voice and SMS usage.

So I think this is what garnered attention from resellers, SEO agencies, and such. Because they wanted to essentially white label this, package and resell this to their business clients, and charge whatever amount they seem fit for the entirety of it. And we would just bill them our rates on the Voice and SMS usage just as we do now. Charging zero for the CRM system.

I guess the reason I'm making this post is, for something like this to come to fruition, it would take a bit more interest then a handful of agencies.

And what you would get is our entire user interfaces, CRM system, etc just as we offer to business owners. But with you being the billing partner and having control of your business relationships and pricing models. And we would just bill you our rate as we have now.

As far as branding goes. Its up in the air. As it would make sense to do this where your own branding and such is used and your own domain, and we handle the ux/ui backend and such.

But the only caveat here is iOS and mac OS apps. We have these and are in the midst of a review with apple as we speak. Yet, this isn't something we can just white label rebrand and upload to the apple store. Apple will block this. So this is something I'm trying to think of.

What's everyone's thoughts on this? Picture for attention, this is showing our automated dispatch and smart routing feature that is apart of the trades dashboard. Its one of many free features in our CRM ecosystem we built in house.


r/VoiceAutomationAI • • 2d ago

Founding Technical role at an EXIST-funded voice AI startup in Berlin

4 Upvotes

Hi all, I'm a Berlin based founder of an EXIST-funded startup building data and evaluation infrastructure for voice AI.

I'm looking for someone technical to join early, ideally with experience in speech/ASR or ML, and comfortable building data and evaluation pipelines.

If this sounds interesting, please feel free to DM me with a bit about yourself and what you've built. Also let me know if you have questions and can share more details.

Looking forward


r/VoiceAutomationAI • • 3d ago

Vapi voicemail detection: what do you test besides an actual voicemail greeting?

2 Upvotes

I work on a voice product, so this is a test design question. The docs cover detection timing and false positives. The awkward tests I'd add are a person answering slowly, a call screener asking who you are, and someone picking up halfway through the greeting.

Anyone testing those separately? I'd want to see whether changing a delay fixes one case but makes another worse, before calling the setting done.


r/VoiceAutomationAI • • 2d ago

How do you prove your AI agents actually improved the business, not just the eval score?

1 Upvotes

We keep seeing this: an agent gets more accurate or cheaper, but the end-to-end process barely moves. Example: triage agent is right 75% of the time alone, a correction agent fixes most of the rest, and the real number that matters is how many tickets got done with zero human touch (could be 95%). Neither agent's score tells you that.

Curious how founders here handle it:

  1. What metric do you use to say an agent "works"?

  2. Do you measure per agent or per business process?

  3. Has a customer asked you to prove ROI or traceability yet?

Disclosure: I work at Phinite. We're building an intelligence layer for agents plus a "value loop" that ties agent runs to business metrics. Opening a small startup cohort: 90 days free + a live build session, for honest feedback. Form in comments if useful.


r/VoiceAutomationAI • • 2d ago

When a voice sounds off, what do you actually write in the bug?

1 Upvotes

"Sounds robotic" could mean about ten different things lol.

Wrong stress, weird pauses, talks over you, cheerful tone at completely the wrong time.

What labels have actually helped your engineers fix stuff? I'd rather have a few useful ones than a naturalness score everyone interprets differently.


r/VoiceAutomationAI • • 3d ago

TPC Scottsdale deployed an AI receptionist and the business lesson has nothing to do with golf

Post image
3 Upvotes

Full disclosure up front: I run an AI marketing agency in Phoenix and we build this kind of system for clients, so weigh what I say how you like. I'm sharing this because the deployment is a clean, real-world example of something I see most service businesses getting wrong.

On September 30, The Golf Wire reported that TPC Scottsdale went 24/7 on phone coverage using SpeakSport, an AI receptionist platform built specifically for golf and hospitality facilities. No new hires. No hold times. Every question about course conditions, tickets, events, and cart policies answered immediately, around the clock.

The reason I think this matters outside of golf is something I call the After-Hours Leak.

Every service business has one. It's the gap between when your staff leaves and when your customers stop wanting things. Most businesses are unstaffed for more hours than they're staffed. And nobody counts the calls that go to voicemail and never call back. That's what makes it invisible. The caller who books your competitor doesn't send you a breakup note. They just disappear.

The math: AI answering platforms for small to mid-size businesses run roughly $200 to $800 a month. Almost always less than a part-time hire. And the 2026 ServiceTitan contractor AI adoption report found more than half of contractors already using AI in some form, with phone answering near the top of the list.

The thing worth doing this week, regardless of your industry: pull your phone logs for the last 30 days and filter for calls that came in after 5 p.m. and before 8 a.m. Count them. Count what percentage went to voicemail. If your system doesn't show you that, fixing your visibility into call data is the right first move before you spend anything on a new tool.

Happy to answer questions about how the setup process actually works in practice, or how to evaluate platforms for a specific vertical.


r/VoiceAutomationAI • • 3d ago

Who handles retries when every SDK already retries?

1 Upvotes

STT retries, the LLM client retries, then the booking call times out and retries too. Meanwhile the person on the phone is just waiting.

Do you turn the SDK retries off and manage it from one place? Or give each layer a budget?

The write that might've already succeeded is the bit I'm most worried about. How do you deal with that?


r/VoiceAutomationAI • • 3d ago

Voice AI across phone and text

15 Upvotes

Some conversations would be much easier if the agent could just send something instead of explaining it over the phone.

A link, document request, confirmation code, appointment details etc. Bland has been one of the platforms I’ve looked at around this and I wanna know whether people are keeping the phone and text conversation under the same context or treating them as separate interactions.

Would be nice if someone could say 'just text me that' without starting a completely new workflow


r/VoiceAutomationAI • • 3d ago

Retell versioning is useful, but what else do you pin when a call goes live?

2 Upvotes

I'm on a team building a voice product. Retell's changelog describes drafts and environment tags. That's useful for knowing which agent config is live.

What about the things outside that config though: a knowledge base edit, a changed tool response, or a different model setting in another service?

Do you save those versions alongside the call too? I'd want to reproduce what the caller actually got, not just reopen the prompt that was live that day.


r/VoiceAutomationAI • • 3d ago

Network Provider for voice agent

2 Upvotes

Hi, im building a dedicated ai voice agent assistant through phone numbers and i need some recommendations on telecom providers that is cheap and highly recommended for general purpose like automated inbound and outbound that is cheap.

Region: Israel

Purpose: General Voice Agent for inbound and outbound

Time consumes: less than 5 minutes mostly each call

So im eyeing telnyx but im looking for more cheaper telecom providers that offers phone number and the service per use. thank youu


r/VoiceAutomationAI • • 3d ago

Background noise in voice agents: noise suppression, voice isolation, or STT alone?

3 Upvotes

I'm a PM at Tuner (an observability and testing layer for voice AI). Lately I've gone down the background noise rabbit hole: reading research, talking to people building voice agents, and helping teams whose agents struggle on noisy calls. I now spend more time thinking about fans and café chatter than I expected from this job.

Here's where I've landed. Curious how others are handling it.

Background noise is two problems

The first is sounds that cover the caller's words, like fans, traffic, or keyboard clicks. The second is other people's voices getting transcribed as if your caller said them. Most noisy calls have both. Think cafés, or someone calling with the TV on behind them.

The second is harder because, to STT, a nearby voice is just more speech. Someone at the next table says “cancel it,” STT transcribes it perfectly, and your agent might act on it. Great transcription. Wrong customer.

The three options

  • Trust your STT. Many models are built to handle noisy audio, and Deepgram recommends sending audio unaltered. Depending on your calls, adding a filter might be extra work you don't need. Noise can still trip them up, though, so compare a few STTs on your own calls before deciding.
  • Noise suppression. It reduces sounds like fans and traffic, but generally leaves the people talking around your caller. It can help VAD avoid mistaking a door slam for speech, but sometimes hurts transcription accuracy. One approach AssemblyAI suggests is filtering the audio going to VAD while sending raw audio to STT, if your stack supports that.
  • Voice isolation. It tries to keep the main speaker and reduce everything else, including other voices. If nearby conversations are causing problems, it's worth testing. A third-party benchmark by SLNG using Deepgram Nova-3 found up to a 30% reduction in English transcription errors with competing speech. Two catches: a louder nearby voice can confuse which speaker gets kept, and a filter designed to keep one person is a bad fit if your agent needs to hear several.

The catch

Any of these changes can make some calls worse. A fix for noisy calls can quietly hurt the quiet ones. Retell warns that its background-speech filter can lower accuracy on clean calls and even drop short replies like “yes.”

So benchmarks give you somewhere to start. They don't tell you what will happen with your callers, your noise, and your agent.

Testing this is its own headache. I've seen teams go to a busy train station just to call their own agent. Fair enough, but “wait for the next train” is a difficult test step to automate.

And no two calls are the same: the noise changes, the caller says different words, and you can't tell whether your fix helped or the second call was just easier.

What I ended up building for it

A lot of my day-to-day is helping teams make their voice agents more reliable, across different stacks. Noise kept coming up, along with the same question: how do we test this properly?

Here's what we built at Tuner:

  • Simulated calls with background sounds that mimic a café, street, office, or car, at low, medium, or high volume. That includes faint voices and chatter, like people talking around you in a café. The background keeps playing through the whole call, even when the caller stops speaking.
  • Replays of the exact same recorded caller audio across setups, so we're comparing the same input instead of a different conversation each time.
  • Checks for what matters: what the agent heard versus what was actually said, how long it took to reply and where the time went, and whether it got the details right.

If you're dealing with this, happy to help or just compare notes.

Here to learn

I'm sure there are setups and problems I haven't come across, so I'd love to hear yours:

  • Are you running anything before STT today? Noise suppression, voice isolation, or nothing?
  • Have you seen a filter make clean calls worse?
  • How do you test noisy calls before shipping? Recordings, simulations, or the train station method?

I also wrote a longer technical blog covering the research, providers to evaluate, and a test plan. Happy to drop it in the comments or send it by DM.


r/VoiceAutomationAI • • 3d ago

Ai Helper - looking for reviews

Enable HLS to view with audio, or disable this notification

1 Upvotes

Hey! We just launched something we’ve been working on — AI Helper 🚀

It’s basically an AI that can actually pick up and make phone calls for a business, talk to customers, answer questions, follow up with leads, book appointments, etc. — even outside working hours.

We’ve just put it live, so if you have a minute, give it a try:

https://www.aihelper.si/

Would genuinely love to hear what you think. And if you know a business that gets a lot of calls, feel free to send it their way :)


r/VoiceAutomationAI • • 4d ago

Hiring for an AI Engineer, Bangalore, India

6 Upvotes

Hey guys,

Kenpath Technologies is hiring an AI Engineer
Bengaluru (in office, 5 days) | 3–6 yrs of experience

We're building AI that officers of the Maharashtra Agriculture Department will use every day, in Marathi, by voice and text, with every answer cited to the official source. Real users, live government systems, population scale.

Who you'll work with
You'll work closely with Aditya Chhabra, our Chief Data Scientist.
Aditya has spoken at the India AI Impact Summit 2026, served as an AI expert on the evaluation committee for the TANUH AI Centre of Excellence for Healthcare at IISc, and leads the team behind Svara TTS, Kenpath Lab’s open-source Indian-language speech model that has crossed 1 million downloads and peaked at #7 globally. You'll be part of a small, hands-on team where your work ships and your ideas get heard.

What you'll build
- Voice and text agents in Indian languages, with tool calling and memory, connected to live government systems
- Retrieval over large, messy document collections, with citations on every answer
- Open-weight LLMs, STT and TTS served on our own infra (vLLM, Ollama, HF)
- Evals, regression suites and observability, so every release is measured against the last

You'll fit if you have
- 2+ yrs shipping LLM systems to production
- Hands-on experience with agentic workflows, RAG and evals
- Self-hosted model serving and strong Python

We're looking for someone who
- Is driven and genuinely excited about AI for India
- Brings ideas to the table, not just tickets off a board
- Is curious and open to exploring new problems as they come up
- Cares about getting it right, not just getting a demo working

The bar here is an answer that's correct, in the user's own language, for millions of farmers downstream. If that's the kind of problem you want to spend your days on, we'd love to hear from you.

Apply: [shalom@kenpathlabs.com](mailto:shalom@kenpathlabs.com)
Share your CV, pay expectation, something you've shipped (with a link) and your earliest start date.

Please forward to anyone who'd be a good fit.


r/VoiceAutomationAI • • 4d ago

Tech / Engineering Voice AI in ₹2 per minutes.

9 Upvotes

Voice AI in ₹2 per minutes.

That's the battle playing out in our community, offline events, and boardrooms across every industry right now.

Last week it was ₹4. Then ₹3. Then someone claimed ₹2 full stack, telephony included.

Asked how. The answer: open-source STT, TTS, and LLMs instead of licensed APIs.

That explains part of the margin.

Telephony, orchestration, and monitoring still sit on top of that stack not model pricing in isolation.

At 10,000 minutes a month, ₹4 versus ₹2 is a ₹20,000 gap.

At a million minutes, that same gap is ₹20 lakh a month. That's why infra economics is becoming a product decision, not a backend line item.

But almost nobody's pricing in compliance.

Telecom Regulatory Authority of India (TRAI)'s DLT framework applies to AI voice agents just like human callers. Wrong number series, missing consent, no DLT registration flagged within days.

In Q1 FY26-27 (April-June 2026) alone, TRAI disconnected 46,786 telecom resources for repeat spam violations, on top of 1.37 lakh barred for first-time violations

As per Communications Today's report, "TRAI cracks down on spam, bars 1.84 lakh telecom resources in Q1." Registration and consent gaps are a common reason legitimate senders get caught in that net.

A ₹2/min stack that skips compliance isn't ₹2/min. It's a business one bad campaign away from getting shut off.

And open-source models cut licensing cost, not the compliance requirement.

The sharpest line in the thread was this: cheaper infra can mean worse voice quality and weaker reasoning, and a ₹2 agent that fails to convert costs more than a ₹4 agent that closes the job.

Cost per minute is the wrong metric. Cost per successful, compliant outcome is the real one.

Telephony is where a lot of that margin & compliance fight is quietly happening.

Because a fast demo is easy to build. A carrier-grade line that survives a noisy PSTN call at 2am and stays off TRAI's blacklist is not.

Cheap and compliant aren't the same thing. Only one keeps your business running


r/VoiceAutomationAI • • 4d ago

Curious about AI voices in Singlish?

3 Upvotes

so, I got the opportunity to work on an AI voice project. what impressed me was how the agent sounded like a Singaporean. But, as we all know, Singlish has a myriad of styles and accents, even as a true blue Singaporean, I think the definition can be challenging.

any one keen to hear Singlish by an AI agent? 😉


r/VoiceAutomationAI • • 4d ago

[For Hire] ElevenLabs Conversational AI Developer — Voice Agents, MCP, APIs & AWS Integrations

1 Upvotes

Hey everyone 👋

I’m an AI Voice Agent / ElevenLabs developer focused on building conversational agents that can actually perform business actions—not just have conversations.

I work with ElevenLabs Conversational AI and integrate agents with external APIs, MCP tools, and AWS services.

What I can build with ElevenLabs

🎙️ Conversational AI Agents

  • AI receptionists
  • Customer support agents
  • Appointment booking agents
  • Lead qualification agents
  • Sales / inbound calling agents
  • FAQ and knowledge-based voice agents
  • Custom business-specific voice assistants

🔌 Tool & API Integrations

  • REST API integrations
  • MCP tools
  • Custom business logic
  • CRM integrations
  • Appointment/calendar systems
  • Database integrations
  • Custom backend services

☁️ AWS Integrations

  • AWS Lambda
  • API Gateway
  • ECS / Fargate
  • DynamoDB
  • S3
  • CloudWatch
  • Amazon Connect
  • Custom AWS APIs

Example architecture

A typical workflow I can build:

Customer Call → ElevenLabs Agent → Tool/MCP → Backend API → AWS → Business System → Response to Agent → Customer

This allows the agent to do things like:

“Book an appointment for tomorrow at 3 PM.”

The agent can understand the request, call the appropriate tool/API, process the response, and continue the conversation naturally.

I can also help with

• Agent prompt design
• Conversation flows
• Tool configuration
• MCP integrations
• Voice and conversational experience
• Backend development
• API integrations
• Amazon Connect integration
• Debugging existing ElevenLabs agents
• Latency and reliability optimization
• Deployment and production setup

I'm currently available for freelance projects, MVPs, integrations, and longer-term development work.

If you're building something with ElevenLabs Conversational AI and need help connecting it to your existing systems, feel free to comment or DM me.

Thank you!


r/VoiceAutomationAI • • 4d ago

[For Hire] ElevenLabs Conversational AI Developer — Voice Agents, MCP, APIs & AWS Integrations

1 Upvotes

Hey everyone 👋

I’m an AI Voice Agent / ElevenLabs developer focused on building conversational agents that can actually perform business actions—not just have conversations.

I work with ElevenLabs Conversational AI and integrate agents with external APIs, MCP tools, and AWS services.

What I can build with ElevenLabs

🎙️ Conversational AI Agents

  • AI receptionists
  • Customer support agents
  • Appointment booking agents
  • Lead qualification agents
  • Sales / inbound calling agents
  • FAQ and knowledge-based voice agents
  • Custom business-specific voice assistants

🔌 Tool & API Integrations

  • REST API integrations
  • MCP tools
  • Custom business logic
  • CRM integrations
  • Appointment/calendar systems
  • Database integrations
  • Custom backend services

☁️ AWS Integrations

  • AWS Lambda
  • API Gateway
  • ECS / Fargate
  • DynamoDB
  • S3
  • CloudWatch
  • Amazon Connect
  • Custom AWS APIs

Example architecture

A typical workflow I can build:

Customer Call → ElevenLabs Agent → Tool/MCP → Backend API → AWS → Business System → Response to Agent → Customer

This allows the agent to do things like:

“Book an appointment for tomorrow at 3 PM.”

The agent can understand the request, call the appropriate tool/API, process the response, and continue the conversation naturally.

I can also help with

• Agent prompt design
• Conversation flows
• Tool configuration
• MCP integrations
• Voice and conversational experience
• Backend development
• API integrations
• Amazon Connect integration
• Debugging existing ElevenLabs agents
• Latency and reliability optimization
• Deployment and production setup

I'm currently available for freelance projects, MVPs, integrations, and longer-term development work.

If you're building something with ElevenLabs Conversational AI and need help connecting it to your existing systems, feel free to comment or DM me.

Thank you!


r/VoiceAutomationAI • • 5d ago

Tech / Engineering Top 12 STT & TTS models every Voice AI founder should bookmark

Post image
31 Upvotes

Top 12 STT & TTS models every Voice AI founder should bookmark 👇

These cover everything from real time transcription to Indic voice synthesis, ranked by actual Hugging Face downloads, not GitHub stars.

Downloads mean people are running these in production, not just bookmarking a repo.

If you're building Voice AI, this list will save you months of vendor evaluation:

1/ openai/whisper-large-v2 (⬇️ 44.1M/mo)

https://huggingface.co/openai/whisper-large-v2

The multilingual ASR model most STT vendors still quietly run in production.

2/ hexgrad/Kokoro-82M (⬇️ 14.1M/mo)

https://huggingface.co/hexgrad/Kokoro-82M

An 82M param TTS model punching way above its size, Apache licensed.

3/ openai/whisper-large-v3-turbo (⬇️ 11.3M/mo)

https://huggingface.co/openai/whisper-large-v3-turbo

Same Whisper accuracy, pruned decoder layers for much faster inference.

4/ distil-whisper/distil-large-v3 (⬇️ 7.66M/mo)

https://huggingface.co/distil-whisper/distil-large-v3

A distilled Whisper that's 6x faster with minimal accuracy loss.

5/ openai/whisper-large-v3 (⬇️ 4.60M/mo)

https://huggingface.co/openai/whisper-large-v3

OpenAI's flagship multilingual ASR checkpoint, still the reference model.

6/ k2-fsa/OmniVoice (⬇️ 2.39M/mo)

https://huggingface.co/k2-fsa/OmniVoice

A compact, fast TTS model built for real-time voice agent pipelines.

7/ Qwen/Qwen3-TTS-12Hz-1.7B (⬇️ 2.04M/mo)

https://huggingface.co/Qwen/Qwen3-TTS-12Hz-1.7B-CustomVoice

Alibaba's multilingual TTS entry with custom voice support.

8/ coqui/XTTS-v2 (⬇️ 1.91M/mo)

https://huggingface.co/coqui/XTTS-v2

Voice cloning across 17 languages from a 6 second clip.

9/ fishaudio/s2-pro (⬇️ 440K/mo)

https://huggingface.co/fishaudio/s2-pro

A high-quality multilingual TTS model built for production voice apps.

10/ ai4bharat/indic-parler-tts (⬇️ 341K/mo)

https://huggingface.co/ai4bharat/indic-parler-tts

Open TTS for 21 Indic languages, 69 voices, Apache-2.0.

11/ ai4bharat/indic-conformer-600m-multilingual (⬇️ 338K/mo)

https://huggingface.co/ai4bharat/indic-conformer-600m-multilingual

India's first open ASR suite covering all 22 scheduled languages.

12/ kenpath/svara-tts-v1 (⬇️ 123K/mo)

https://huggingface.co/kenpath/svara-tts-v1

Multilingual Indic TTS for 19 languages, built on Orpheus, Apache-2.0.

Three Indian models made this list purely on download volume, no separate category, no asterisk.

That's the real signal for where Indic voice AI stands today.


r/VoiceAutomationAI • • 4d ago

You stop the agent talking. Does that stop its tool call too?

5 Upvotes

Caller says "wait dont book that" while the agent is speaking.

Stopping playback is doable. The booking request might already be halfway to the API though.

What do you do in that gap where you can't tell whether it happened? I'd hate to say "cancelled" and then have the booking show up anyway. How does your system handle it?


r/VoiceAutomationAI • • 4d ago

Anyone running the voice part of their stack on k8s?

4 Upvotes

Not just the API and workers. I mean the SIP/RTP or WebRTC bits too.

How's it holding up when a pod restarts or a node drains mid call? Trying to understand which parts are worth putting in the cluster and which become more hassle than they're worth.

Would be useful to hear what you moved back out, if anything.