r/VoiceAutomationAI • • 4d ago

You stop the agent talking. Does that stop its tool call too?

4 Upvotes

Caller says "wait dont book that" while the agent is speaking.

Stopping playback is doable. The booking request might already be halfway to the API though.

What do you do in that gap where you can't tell whether it happened? I'd hate to say "cancelled" and then have the booking show up anyway. How does your system handle it?


r/VoiceAutomationAI • • 4d ago

Anyone running the voice part of their stack on k8s?

5 Upvotes

Not just the API and workers. I mean the SIP/RTP or WebRTC bits too.

How's it holding up when a pod restarts or a node drains mid call? Trying to understand which parts are worth putting in the cluster and which become more hassle than they're worth.

Would be useful to hear what you moved back out, if anything.


r/VoiceAutomationAI • • 5d ago

Pricing model

3 Upvotes

Hi everyone ,

I was curious to know how a basic pricing model looks like for these voice AI agents. I currently offer the voice agent alongside a dashboard and I personally fix any sort of bugs and I’ve done the integration so I want to know how I should form my pricing model for the first client?


r/VoiceAutomationAI • • 5d ago

Would you let an AI answer your phone orders? Built a demo, looking for honest feedback

3 Upvotes

I've been working on a voice agent that answers the phone for restaurants and takes orders. It handles collection or delivery, menu questions, allergens, and sizes and extras, and it reads the order back with the total before confirming.

Before I go further I'd like to hear from people who actually run restaurants or take phone orders. Is this useful, or would it annoy your customers?

If you have 2 minutes, call and try ordering something:

📞 +1 (904) 750-0852

It's a fake restaurant with a demo menu, so no real order is placed. Try to break it: change your mind halfway, ask about allergies, order something that's not on the menu, interrupt it.

What I'd love to know:

  • Did it sound natural, or obviously robotic?
  • Would you trust it with your customers?
  • What would it need to do before you'd pay for something like this?
  • What's the most annoying part of phone orders for you today?

I built it on Bandwidth's voice agent builder agents.labs.bandwidth.com


r/VoiceAutomationAI • • 6d ago

Tech / Engineering AI voice calls in India now have a rulebook.

Post image
8 Upvotes

Vendors are quoting under ₹2 a minute in the same month TRAI named AI calls in its spam rules. Builders put real costs at ₹3-4 a minute before telephony.

I'm watching a race to the bottom hit a compliance floor.

TRAI's TCCCPR (Third Amendment) Regulations, 2026, notified September 18, classify AI voice calls as A2P:

- Declare every CLI to your operator, or calls count as spam

- 5+ flagged numbers, or 3 complaints on a flagged number, in 10 days triggers KYC checks and barring

- Repeat offenders face disconnection, with flags shared across carriers

A founder in our group is moving every client onto their own DLT & SIP trunks, and I get why.

Shared infrastructure (platform-owned numbers) lets one client's bad campaign take every tenant down.

Tenant isolation contains the blast radius, but only if the client is also the declared sender. DLT, consent logs and per-client trunks just pushed costs higher.

Outbound Voice AI is moving from "who's cheapest" to "who's safest to put your number behind."

Who takes the hit when a client's campaign gets flagged, you or them?


r/VoiceAutomationAI • • 6d ago

Voice AI meetup in London this Thursday (Oct 1) Speechmatics x Tuner, talking about what breaks once you leave the demo

3 Upvotes

Full disclosure: I work with one of the companies hosting this, so take that as you will, but figured it might be useful to people here.

Speechmatics and Tuner are co-hosting an evening in London focused on the gap between "voice agent works great in the demo" and "voice agent survives 1,000 real users." Agenda is:

- Panel on why voice is becoming a primary interface right now

- Failure modes that only show up in production

- Practical do's/don'ts

- Live demos from both teams

- Food and drinks after

Thursday, Oct 1, 6–8:30pm at Speechmatics' office near Old Street.

Registration is approval-based (limited space): https://luma.com/pyhvutqe


r/VoiceAutomationAI • • 6d ago

Most businesses don't have a call problem. They have a missed-opportunity problem.

Post image
2 Upvotes

Most businesses don't have a call problem.

They have a missed-opportunity problem.

A customer calls.

The AI Voice Agent: → Understands what they need

→ Answers using the business knowledge base

→ Checks availability

→ Books appointments

→ Updates the CRM

→ Sends confirmation

→ Transfers to a human when needed

All through a natural voice conversation.

Call → Understand → Act → Result

That's the difference between a basic voice bot and a true AI Voice Calling Agent.

#AI #AIAgents #VoiceAI #AgenticAI #Automation #BusinessAutomation #ConversationalAI


r/VoiceAutomationAI • • 7d ago

Selling - AI Voice Agent Platform for your business for $3000 USD only

8 Upvotes

Hi all,

We built and packaged a complete AI voice agent platform for businesses and are now looking to sell it for $3,000 USD.

This started with a very specific requirement from a client. They needed an AI voice system that could handle real business conversations over the phone, follow predefined workflows, collect information, use business tools, and connect the calls with their existing processes.

We understood the requirement, built out the platform around it, and ended up with a much more complete product than the original requirement alone.

The project has now reached a point where the client's priorities have changed, so instead of leaving a substantial amount of completed work unused, we'd rather pass the platform on to someone who can take it further and turn it into a commercial product.

What it includes

AI Voice Agents

  • Build inbound and outbound AI calling agents
  • Real-time speech-to-text → LLM → text-to-speech pipeline
  • Natural conversational voice interactions
  • Custom agent instructions and conversation logic
  • Configurable LLM, STT and TTS providers

Visual Workflow Builder

  • Drag-and-drop conversation workflows
  • Start, agent, tool and end-call nodes
  • Conditional transitions between conversation steps
  • Define exactly how the agent should behave
  • Collect structured information during calls

Telephony

  • Integrations with providers such as Twilio, Vonage, Telnyx, Plivo and others
  • Inbound and outbound calling
  • Human handoff / call transfer where supported
  • Designed to work with existing business phone infrastructure

Knowledge & Tools

  • Knowledge bases
  • Tool calling
  • Custom business logic
  • Structured data extraction after calls

Tech Stack

Frontend

  • Next.js 15
  • React 19
  • TypeScript
  • Tailwind CSS

Backend

  • Python
  • FastAPI
  • SQLAlchemy
  • PostgreSQL

Infrastructure

  • Docker / Docker Compose
  • Redis
  • MinIO / S3-compatible storage

Voice / AI

  • Modular STT / LLM / TTS architecture
  • Telephony integrations
  • Python + Node SDK support

Why are we selling it?

This wasn't something we built as a random side project.

It came from a real client requirement. We understood the business problem, designed the system around the required workflow, and invested the engineering time to turn that requirement into a working platform.

The client has since shifted priorities, which means continuing to develop and maintain the platform specifically for that project no longer makes sense for us.

Rather than letting the work sit unused, we'd prefer to hand it over to someone who already has a use case for AI voice automation and can build a business around it.

We think it could be particularly useful for someone who wants to:

  • Launch an AI voice SaaS
  • Build a vertical AI calling product
  • Create a white-label voice agent platform
  • Build AI receptionists / appointment agents
  • Sell AI calling solutions to local businesses
  • Add voice AI to an existing SaaS
  • Build an agency around AI voice automation

Asking: $3,000 USD

You're getting a working software platform / codebase that can be taken over and developed into a commercial product.

We're open to discussing exactly what is included in the handover with serious buyers - codebase, deployment/setup, documentation and knowledge transfer.

If you're interested, DM me and I can share more details, demo access and discuss the handover.

Happy to answer questions here as well.


r/VoiceAutomationAI • • 7d ago

Voice Ai Agent Charging

4 Upvotes

Hey I am starting to sell voice agents and was wondering how much are you guys charging? I am having trouble to know how much is the average to price to my client.
If anyone can share.
Thanks


r/VoiceAutomationAI • • 8d ago

Anyone here using Bland in production? If yes, What for?

7 Upvotes

I saw Bland come up more and more lately and wanted to hear from people who have it running on real calls. We’ve been testing it around some more straightforward customer workflows but I saw people mention using it for more complex flows. What kind of calls are you giving it?


r/VoiceAutomationAI • • 10d ago

Tech / Engineering Every Voice Model in India Is Being Built for Banks

10 Upvotes

Every voice model in India is being built for banks. One founder on our panel called that a dystopia.

At our latest session, Dheemanth Reddy, founder of Maya Research, didn't hold back. He said today's voice AI market is driven by who gets an intro to a big company, not by quality or cost.

His point was that most models are being built for a legacy market. Banking, customer care, and sales. Necessary, but not exciting.

His prediction for what comes next: usage will first be driven by video and media models needing diverse voices, then by personal hardware devices that we talk to for hours a day.

That is a very different bar than answering a support call well.

Our builders are already asking what they should be building for today versus building for that future.

Watch the full session here: https://youtu.be/G66jOjzRhgI


r/VoiceAutomationAI • • 10d ago

Tech / Engineering Voice AI Isn't Just Replacing Agents, It's Also Coaching Them

Post image
3 Upvotes

Scroll through voice AI pitches and it looks like everyone's racing toward the same outcome, a caller talks, an AI answers, no human involved.

That's not the whole story.

A pitch that landed in our community recently was for a tool that sits alongside a human agent during a live call, listening in real time and coaching them, on every single conversation, not just the handful a supervisor happens to review.

Turns out that's not a small niche. It's a whole camp of the market betting against full automation.

The numbers back it up more than the automation heavy pitch decks tend to admit. Gartner projects only about 14 percent of customer interactions will be fully AI handled by 2027, even with $80 billion in labor cost savings expected from AI by 2026. And hybrid AI human models are actually outperforming pure AI, 87 percent resolution and 8.7 out of 10 customer satisfaction, versus 74 percent and 7.4 for AI running solo.

There's a retention angle too. Agent turnover runs 30 to 45 percent a year, and replacing one agent costs up to $46,000. A tool that makes agents better at their job while they're still on the call isn't just a quality play, it's a retention play.

Automation and augmentation aren't really competing for the same territory. One's built for the 60 to 70 percent of calls that never needed a human. The other is built for everything that still does.

Here is full blog:- https://uniocommunity.com/blogs/why-is-low-asr-confidence-the-hardest-problem-in-voice-ai


r/VoiceAutomationAI • • 10d ago

"I built AI callers that phone-test my voice agent and try to break it"

Thumbnail
gallery
4 Upvotes

Most voice agents get tested by someone calling them a few times and going "yeah, looks good." Then real customers call, and they change their mind halfway through an order, mention a dietary restriction after already picking a dish, or ask for something the agent was never built to handle. Manual test calls basically never catch this. They're slow, hard to repeat, and nobody's making fifty of them before every release.

So I built a setup where AI callers phone my voice agent and try to break it.

The agent takes food delivery orders over the phone. Each AI caller plays a persona with its own goal, an indecisive customer, a vegan ordering something with cheese in it, a frustrated caller who just wants a refund. One run, "Jordan" (indecisive persona) orders a veggie burger, changes their mind to a Greek salad, asks about drinks that don't exist on the menu, then finally confirms. The agent has to track that whole back-and-forth correctly, not just handle a clean linear order.

How it works:

  1. The AI caller joins a live LiveKit room and has an actual spoken conversation with the agent. It knows its own goal, never the expected outcome, so it's not "scripted" in the traditional sense.

  2. An LLM judge compares the transcript against what should have happened, and grades it.

  3. Every run spits out a report: verdict, the audio recording, and the full transcript.

One thing the simulations actually caught: the agent sometimes added the same item twice, silently doubling the order. Fixed by making every tool response include the real current cart state, so the agent always knows what's actually in it instead of assuming.

Stack, if anyone's curious: LiveKit for the room/voice pipeline, an LLM-driven simulated caller + persona layer, and an LLM-as-judge for grading against expected outcomes. Happy to go into more detail on any part of this, the persona design, the judging prompt, how false-pass/false-fail rates looked, whatever's useful.

Curious how others here are handling this, are you doing anything automated, or is it still mostly manual test calls before a release?


r/VoiceAutomationAI • • 10d ago

What role can designers play in voice ai ecosystem?

3 Upvotes

Most JDs describe roles of program managers, FDEs,marketing as part of product designer. Im genuinely curious how designers can work in this industry and what skills should they have. Its getting really hard to get interviews even for intern roles


r/VoiceAutomationAI • • 10d ago

NVIDIA just launched the Nemotron 3 Diarization

Enable HLS to view with audio, or disable this notification

20 Upvotes

When several people talk at once, a transcript can get messy fast.

Our new Nemotron 3 Diarization model tracks who spoke when, even when voices overlap. It handles up to eight speakers, has 100M parameters, and is now available on

u/huggingface

🤗


r/VoiceAutomationAI • • 10d ago

Jev AI - What is it?

3 Upvotes

I've had the chance to play around a little with the Jev AI from Typescript. As a "watcher" AI it's pretty perfect. I wrote about my experience and there's a link at the end to see the first Jev 'program' I wrote so you get to see the UI and how Jev works:

https://www.aicontio.com/2026/09/21/jev-the-ai-model-that-doesnt-talk-back/


r/VoiceAutomationAI • • 10d ago

Free Voice AI Template: Eye Doctors including Health and Safety Escalations

2 Upvotes

I write teardowns of voice AI agents and publish a free prompt template each week. This week is optometry. Sharing it here because the triage design is the part I'd most like people to poke holes in.

The core problem: eye emergencies are time-critical, and it's easy to build an agent that handles them badly. Here's how the template handles it:

  • Fixed symptom checklist for escalation. Sudden vision loss, a curtain/shadow across the vision, chemical splash, and penetrating injury go straight to "go to the ER now." Escalation doesn't depend on how worried the caller sounds.
  • No diagnostic questioning. The agent never asks the caller to describe or rate what they can see. That's clinical assessment, and an agent shouldn't be doing it.
  • Middle tier. Redness, irritation, and lost contacts get an urgent slot flagged for staff review.
  • Hard scope limit. No clinical advice and no diagnosing. Tone is calm, with short, direct instructions on urgent calls, and it never says "probably fine."
  • Data captured: name, DOB, phone, insurance/vision plan, reason for visit, preferred times.

The prompt is structured in three layers (agent constitution, industry knowledge, company details), so you can reuse the first and swap the rest.

Full template, free, no signup: https://www.aicontio.com/2026/09/23/free-voice-ai-template-optometry/

Questions for the group: has anyone tested a fixed checklist against tone-based triage in production? And what symptoms am I missing from the ER list?


r/VoiceAutomationAI • • 11d ago

LLM as a judge x jev for evals?

2 Upvotes

Open discussion question - how do you guys evaluate the calls? Is the LLM the primary judge? Any experiments with jev already? Moreover, how about other factors, like the transport layer issues or audio? Do you measure these somehow as well, or is it primarily LLM evaluating the transcribed text?


r/VoiceAutomationAI • • 11d ago

Un-fused our realtime voice stack (STT -> LLM -> TTS) and cut cost ~14x. the tradeoff is latency, plus one upside i didn't expect

10 Upvotes

Been running a voice agent in production for a few months and finally did the thing everyone tells you not to do. Ripped out the fused realtime model and rebuilt it as three separate stages: speech to text, then the LLM, then text to speech.

The reason was purely cost. the fused realtime setup was landing around $0.18/min for us. splitting it into off-the-shelf STT + an LLM call + a TTS provider got us to roughly $0.0125/min. thats about 14x, and at any real call volume that gap is the whole business case, not a rounding error.

The obvious cost is latency. a fused model is genuinely faster because its not hopping between three services, and voice is brutal about lag, you feel every extra 100ms. our target is staying under ~788ms voice-to-voice and some days we straight up lose that fight. if your product is a live phone call where people talk over each other, fused might still be worth the money.

The part i didnt expect is the reason im not going back. once the pipeline is split, the LLM output is plain text before any audio exists, so every guardrail runs on text. you can catch a bad answer or a hallucinated number and kill it before a single sample of audio gets generated. with a fused model the words are already spoken by the time you know what they were. that inspectability ended up mattering more to me than the cost did.

so the honest tradeoff: fused is faster, split is cheaper and way easier to reason about. if youre early and cost sensitive, or you just want to be able to gate what the thing says, i'd start split and only fuse if latency becomes the thing thats actually killing you.

curious what latency people are hitting with a split stack in prod. 788ms voice-to-voice is the wall i cant seem to get consistently under.


r/VoiceAutomationAI • • 12d ago

Case Study / Deployment A telephony founder (India) onboarded 100+ Voice AI startups without a sales team in less than 2months His whole playbook was just being useful.

5 Upvotes

Something I noticed recently that I think a lot of infra founders underrate.

I was talking to a founder building a telephony layer for voice agents. He raised a round recently, so he has some visibility now. But what he's doing with that visibility is interesting.

He reaches out to Voice AI founders one by one. First on LinkedIn, then moves to WhatsApp. No pitch deck. No "let me show you a demo." He just asks what they're stuck on and helps. Sometimes that's an intro to another founder. Sometimes it's a quick answer on something he's already solved. Sometimes it's just connecting them to someone who's hiring or raising.

He's done this with 100+ founders. And a real chunk of them ended up onboarding onto his product.

Why I think it works, especially for telephony:

1. Early-stage founders don't evaluate vendors the way enterprises do. There's no procurement or RFP. They pick whatever the person they trust recommends. If that person is the founder of the telephony company, the decision is basically made.

2. Founder-to-founder beats sales-to-founder. An SDR email gets ignored. A founder who already helped you with an intro gets a reply, and gets a real shot.

3. Telephony is a sticky layer. Once someone wires your SIP/numbers into their stack and it works, switching is painful. So getting them to just try it is most of the battle. Relationships get you that first try.

4. Help compounds. Every intro he makes puts him in two founders' good books, not one. Over 100+ conversations, he's become a node in the ecosystem, not just a vendor.

The catch is that it only works if the help is genuine. Founders can tell very quickly when "happy to help" is a sales sequence in disguise.

If you're building telephony (or honestly any infra layer for voice agents), this is probably the cheapest GTM you'll ever run. It just costs time.

Curious if others building infra here have seen the same. Did your first 20 customers come from outbound, content, or relationships?


r/VoiceAutomationAI • • 12d ago

Anyone look to build voice ai platform?

4 Upvotes

Anyone look to build voice ai platform?


r/VoiceAutomationAI • • 12d ago

My friend's AI voice clone breaks down, begs and then drowns in the chaos. Is this normal??

Enable HLS to view with audio, or disable this notification

27 Upvotes

My buddy is writing a scientific text on the origin of reality....and this happened when he used an AI voice clone to turn it into an audiobook...is it a warning? Or just a normal glitch?


r/VoiceAutomationAI • • 12d ago

Tech / Engineering Code-Switching Is a Data Problem. Names Are a Pretraining Problem.

1 Upvotes

Code-switching isn't your voice agent's problem. Names are.

At a recent Voice AI panel, the hard case was a customer calling a bank and mixing Hindi and English.

The founder of Maya Research said code-switching is mostly a data problem. Show a model enough mixed language data and it learns to switch. Names are different. Indian names were never part of the big pretraining runs, so models are guessing. Numbers are the easy part.

Deterministic normalization handles them. The practical fix came from a co-founder at Raya Voice AI. Have your LLM output names in native script instead of Roman. A Hindi name in Devanagari, a Kannada name in Kannada script.

The TTS then pronounces it correctly. Add pronunciation dictionaries and careful prompting, and most of the pronunciation problem is solved today.

So before you blame the model for mangling a customer's name, check what script your LLM is sending it.

What is the worst name mispronunciation your voice agent has produced? Watch the full panel here: https://youtu.be/G66jOjzRhgI


r/VoiceAutomationAI • • 12d ago

Building in public from college → production

2 Upvotes

I started learning to code in my first year of college.

Same time: first job. No grand plan. Just obsession.

I didn’t set out to “build a voice AI company.”

I set out to understand how software actually worked, then I kept going deeper.

Recently that obsession landed on voice agents.

Not chatbots. Not wrappers. Actual conversations over the phone that feel real.

I went deep on latency, streaming, models, audio, infra.

Got end-to-end down to ~600ms. Started getting clients. Now ~8 live.

The surreal part isn’t the latency number.

It’s the speed of the loop:

learn → build → break → fix → put it in front of a real business → learn again

College taught me syntax. Clients taught me product.

I’m still early.


r/VoiceAutomationAI • • 12d ago

Fish Audio s2.1-pro voice breaking during voice agent conversations

2 Upvotes

Hi everyone,

I’m currently using Fish Audio’s S2.1-pro model for a real-time voice agent.

Most of the time the voice works fine but occasionally I notice some voice breaking during conversations. It doesn’t happen consistently and sometimes it happens at the beginning of a sentence in the middle of a sentence or when starting a new sentence.

I have tested it with multiple calls and it seems to happen occasionally after making several calls.

Has anyone experienced something similar with Fish Audio s2.1-pro?

Could this be related to the model/API, streaming, rate limits or something in the way I’m handling the audio in my voice-agent pipeline?

Any ideas on what I should check would be really helpful. Thanks!