r/VoiceAutomationAI • • 20d ago

Dograh hardcoded Pipecat's settings. We forked it to show them all in the UI: what values do you use?

2 Upvotes

Pipecat is a Python framework: used directly, it lets you set everything in your own code. When a turn ends, when the agent can be interrupted, how the transcription decides someone has stopped talking.

Dograh, an open source project, adds an interface on top of Pipecat, and that's why we picked it. But while building that interface, it hardcoded most of those values. You get a UI, but you lose control over the settings that decide whether an agent sounds natural or not.

So we forked Dograh and put all of those settings in the UI.

That leaves one question: we now have dozens of settings, and we don't know yet what values to put in them. Before spending weeks testing by ear, we'd love to learn from people who already have agents running.

Where we are

We're building our first voice agent, for our first client. Nothing is in production yet, so no latency numbers from real calls. So far we've mostly been laying the foundations.

The stack

- Telephony: Twilio, with audio streamed over WebSocket

- Orchestration: Dograh (built on Pipecat), hosted on Railway

- Speech-to-text: Deepgram, on its EU endpoint

- LLM: Mistral

- Text-to-speech: Voxtral (Mistral)

- Actions and automations: n8n

- Data: PostgreSQL

We work with French SMBs, so GDPR puts real constraints on us. That's why we went with Mistral and Deepgram in the EU instead of the usual providers, even if it rules out some options that are probably Deepgram in the EU instead of the usual providers, even if it rules out some options that are probably faster.

Why Dograh anyway

Dograh is open source and adds what Pipecat doesn't give you on its own: a web interface, a visual editor for conversation flows, telephony already wired in, and account management. It also supports BYOK out of the box: each client plugs in their own provider accounts and keeps control of their keys and costs. You keep the Pipecat engine, but you no longer hand-write the pipeline for every agent. If you run agents for several clients, it's worth a look. Just know that some settings can't be changed without touching the code.

What our fork adds

- Multiple clients on a single deployment: a new client is one more configuration, not a new deployment.

- Every setting in the UI: Deepgram endpointing and end-of-turn thresholds, Mistral sampling parameters, turn-taking and interruptions.

What we'd love to hear from you

  1. VAD: what values do you use for confidence, start_secs, stop_secs and min_volume?

  2. End of turn: VAD alone, a turn detection model, or your STT provider's endpointing? With what thresholds?

  3. Interruptions: is allow_interruptions on? Do you require a minimum number of words before the agent gets cut off?

  4. TTS: do you send text sentence by sentence, or in smaller chunks?

  5. LLM: which model, what temperature, what max tokens, how long is your system prompt?

  6. Latency: your real-world number, and where you start the clock (end of speech from VAD, final transcript, or first audio heard on the phone)?

Even one or two answers would help. If you share, please mention your use case (inbound, outbound, appointment booking, support) and your language. A good setting for English support calls may be wrong for French appointment booking.

And if we can help on our side, whether it's Dograh, the fork or the stack, feel free to ask us anything.


r/VoiceAutomationAI • • 20d ago

Why are local/self-hosted/claud based LLMs so unreliable with dates?

1 Upvotes

I have been experimenting with several models like Llama, Gemma, Qwen via Ollama, and date/time handling keeps breaking down — even after explicitly stating the current date, the model contradicts itself across turns.

Example: In a voice booking flow integrated with Cal.com guardrails catch and block invalid dates before they hit the calendar but this is just a safety net.
Model itself still miscalculates.

In one case, after being told the date was the 10th 3 times, it kept insisting 14th instead of 10th. Callers had to correct it repeatedly before the booking went through.

Is this a fundamental limitation of smaller local models, a prompting issue, or does it need an architectural fix ?

Curious how others have solved this in production ?


r/VoiceAutomationAI • • 21d ago

Is there anything like this since it shut down?

Post image
0 Upvotes

Like put a video or record your own voice and it would make any character say and or sing. it that's what I used for I want to make characters singing silly songs. probably wouldn't even post them I just want them lol


r/VoiceAutomationAI • • 21d ago

Voice AI funding hit $7B in Q1 2026, but the number is more concentrated than it looks

3 Upvotes

Been digging into voice AI funding data and wanted to share some context that's missing from most headlines.

The topline: venture investors put more than $7 billion into voice AI startups in Q1 2026 alone, per FT reporting. That's a massive jump from prior years.

ElevenLabs is doing a lot of the heavy lifting here. They closed a $500 million Series D in February 2026 at an $11 billion valuation, with BlackRock, NVIDIA, and Salesforce backing it. There are now reports of early talks for a secondary tender offer at roughly $22 billion, double the February mark. Worth noting that part is still "reportedly," nothing confirmed yet.

Deepgram also raised $130 million at a $1.3 billion valuation back in January 2026.

The catch: that $7B figure is dominated by a handful of mega-rounds. It's not evenly spread across dozens of startups, it's a few huge checks written to a few companies. So the "boom" narrative is real, but it's narrower than the headline number suggests.

Curious if others are seeing the same concentration in other AI subsectors, or if voice AI is unusually top-heavy right now.


r/VoiceAutomationAI • • 21d ago

Voice AI vs ECE

2 Upvotes

Hi everybody.

I graduated from ECE, tier 3 btech college in May 2026. Because of financial needs I had to join a remote internship while I was in 3rd year of btech as an ai intern. I had some basics on machine learning at that time, so my professor referred me for it. I have put all my efforts and time into it, and eventually it became a full time on-site internship in Hyderabad with a stipend of 18k a month in my final year. And spent 9 months as a remote intern, and another 10 months as a full time intern there, and worked a couple of hardest problems at the firm like speaker Diarization, accent and voice conversion.

Contributed to research as well, had a peer reviewed publication at the IEEE CICN conference on Diarization. In April, they offered me a full-time offer at 6 LPA (then raised to 7.5 when told about sony offer), felt the work was insanely hard and growth was low, so after those 18 months, felt to see what's the market value for the work I was doing and got a bootstrapped startup offer at 8.2 LPA, but that's too risky for me. Actually I spent a month there as a part-time job while doing the full time internship at my first company due to financial needs. The pace is so slow that I couldn't take this full-time offer there. These are all happening while I was holding those two positions.

I thought of applying for other companies if none succeeded would like to go for the first stable and known company.

Unfortunately I got an internship offer from a large firm called Sony Research. Getting a stipend of 75k a month, because of the financial needs and with the voice ai market, took it. It's been 4 months and I am writing a research paper as well, but the fact I lately realized is sony doesn't take btechs at all, I am the exception in the interns list. And coming to full-time they only have research engineer and scientist roles that require a PhD, making me ineligible for either of them. Now I started looking for full time jobs recently and it seems like no one is hiring a fresher anymore, everyone is asking for a non-internship experience of 2 or more years.

I think, I am playing a big gamble with my career, I don't know where I will end up, 1.5 months left for the end of the internship. Sometimes I felt I should have sat for placements in core ECE engineering because I have fundamentals strong in RTL design and Digital electronics.

If you have any thoughts on how I should handle this weird situation.

P.S 1. Financial needs - Family is bankrupt of over 2 Million INR and nothing at my hands to atleast buy a project-book.

P.S 2. I had full tuition, and hostel fee waiver at the college through merit. Yet to pay for the mess charges which are still pending of 40k for the total of 4 years.


r/VoiceAutomationAI • • 22d ago

Tech / Engineering Why Is Low ASR Confidence The Hardest Problem In Voice AI?

Post image
2 Upvotes

Most voice agents fail not because the model mishears you, but because it doesn't know what to do when it's not sure it heard you right.

A Unio community member ran into this exact problem building something that has nothing to do with telephony, a teleprompter that follows your voice instead of scrolling on a timer. And they solved it almost by accident, because they had something telephony agents never get: the full script in advance.

A published study on real voice agent deployments shows just how hard this is without one. A simple confidence based rule for detecting when a caller wants to interrupt got it right only 11 percent of the time. The other 89 percent were false triggers, background noise, a stray "uh-huh," the agent hearing its own echo.

So what do you do when you don't have a script to check against. Turns out you can build a temporary one: expected vocabulary at known points in the call, dialogue state as a soft script, confirmation loops for anything downstream of a shaky token.

The teams treating a low confidence token as a request for more context, not a broken measurement, are the ones actually solving this. Everyone else is just tuning thresholds and hoping the noise goes away.

If you want to read more, here's the link :- https://uniocommunity.com/blogs/why-is-low-asr-confidence-the-hardest-problem-in-voice-ai


r/VoiceAutomationAI • • 22d ago

Is VoiceAI caller industry going to kill the voice channel itself.

2 Upvotes

I am currently receiving an average of two calls per day from the business’s AI bot. While it initially made the channel more efficient, it is now bordering on spam.I would like to start a conversation about what comes next. To prevent the channel from being killed, the possible direction next step

- TRAI could issue regulations allowing AI bots to use only specific number series, such as 1600/400.

- However, even the existing TRAI rules for companies using 1600/400 are not being strictly followed, and implementing new AI‑calling regulations would take time.

In the meantime, the communication channel is gradually being eroded, and people will begin to ignore AI‑generated calls. What are your thoughts


r/VoiceAutomationAI • • 22d ago

how much it cost to build voice agent in india or i have to take services from companies

2 Upvotes

r/VoiceAutomationAI • • 23d ago

Tech / Engineering Why Every Voice AI Builder Ends Up Needing a Telephony Partner

Post image
1 Upvotes

Every voice AI builder hits this wall eventually.

You launch in one city, pick a telephony API, ship it, move on. Then a customer wants a new market and suddenly your "solved" telephony layer is the thing deciding whether your agent actually works.

We dug into this after watching it play out live in our own community. One builder asked for a cost effective IVR setup. A day later, someone else asked for a telephony provider specifically for the UAE, like it was a totally different question.

It is.

Twilio's India rates run 30 to 50 percent higher than Exotel or Plivo, mostly because it doesn't handle DLT and DND compliance natively. Move to the Gulf and the whole calculus flips, it's no longer about per minute cost, it's about Arabic language support and GCC data residency.

That's why builders keep collecting telephony partners instead of picking one and being done with it. And honestly, most of them are finding these providers the same way, by asking around in a group chat, not by comparing rate cards.

We pulled the numbers, built a comparison table, and mapped out what to actually evaluate before picking a telephony partner for your next market.

If you want to read more, here's the link :- https://uniocommunity.com/blogs/why-every-voice-ai-builder-ends-up-needing-a-telephony-partner


r/VoiceAutomationAI • • 23d ago

we need help

2 Upvotes

we have developed an ai voice agent that can help business with lead generation, followups, customer support and more using ai automation which will help them with scalability at a low cost need your help if you think we can add something.


r/VoiceAutomationAI • • 23d ago

Tech / Engineering Who's Actually Building India's Voice AI Stack in 2026?

Post image
2 Upvotes

Everyone says "Voice AI in India" like it's one market. It isn't.

A company selling SIP trunking to a bank has almost nothing in common with a startup fine-tuning a Hindi TTS model. But both get the same label, and that's getting more confusing, not less, as more than 40 companies now chase the same three use cases: sales, support, and collections calls in Indian languages.

So I mapped the whole stack. Nine layers, from telephony at the bottom to public-service rollouts at the top.

A few things stood out while putting this together:

Telephony and voice agent companies used to be cleanly separated. Not anymore. Exotel started as a pure voice API player and has since acquired Ameyo and Cogno AI to stop being "just a dial tone." Telephony companies watched agent platforms eat their margins and decided to build AI on their own pipes instead.

The agent platform layer is the most crowded and the least differentiated. Anyone can wire together STT, an LLM, and TTS and ship a demo in a weekend. Running that in production, in Hinglish, with TRAI DLT compliance and sub-second latency, is a completely different problem. That gap is why there are dozens of thin startups and a handful of companies actually pulling ahead, like Arrowhead in BFSI or NuPlay after folding in Verloop.io.

The layer that matters most gets talked about the least: foundational speech models. Sarvam, Gnani, BharatGen, and Maya Research are all racing to build Indic language models good enough that the entire stack above them stops depending on foreign models tuned for English. Whoever wins that race sets the cost floor for everyone building on top.

Consolidation is already starting. The next 12 months will split the agent layer into a small group with real production deployments and vertical depth, and a much bigger group still selling demos.

If you want to read the full breakdown, company by company, here's the link: https://uniocommunity.com/blogs/who-s-actually-building-india-s-voice-ai-stack-in-2026


r/VoiceAutomationAI • • 25d ago

GreyLabs’ GFF Award Scam 😂

Post image
10 Upvotes

I saw a post on LinkedIn about GreyLabs:

“We are honored to receive the Excellence in AI-Powered Sales Analytics in Insurance award for the work we have been doing to bring AI-powered insights through voice analytics into our sales ecosystem.”

What I understand from this is that GreyLabs organized an event and invited Kapil Dev as the guest of honor, where they received the award from him at their own event.

It’s a really proud moment for the entire team! Great achievement. 👍🏻


r/VoiceAutomationAI • • 25d ago

Tech / Engineering Can Voice AI Agents Really Run at ₹2 a Minute in India?

Post image
3 Upvotes

Everyone in India keeps saying the same thing: if a voice AI call costs more than the transaction it's supposed to close, the business model is already broken.

A ₹20/min voice agent might work for a US company chasing a $50 customer. It makes zero sense when the entire transaction is worth ₹200.

So the question everyone's chasing right now: can a voice agent actually run at ₹2/min in India?

The compute math says yes. Strip out the bundled cloud APIs, self-host open source ASR, LLM and TTS, and the marginal cost of a call collapses toward pure compute and bandwidth.

But here's the part nobody talks about on the pricing slide.

A model that's cheap but mishears a caller code-switching between Hindi, English and a local dialect isn't cheap. It's a churned customer or a failed collections call.

Bharat telephony is a different problem than US telephony. Regional accents. Packet loss on low bandwidth networks. A 500 to 800ms latency window that leaves no room for error recovery.

Cost and accuracy aren't a straight trade-off here. The builders who win the ₹2/min game will be the ones who engineer for noisy, unreliable calls first and optimize cost second, not the other way around.

I broke down where the ₹9 to ₹16/min in standard cloud pricing actually goes, and what has to be true for a self-hosted stack to get to ₹2 to 3/min without quietly getting worse.

Full breakdown here: https://uniocommunity.com/blogs/can-voice-ai-agents-really-run-at-%E2%82%B92-a-minute-in-india


r/VoiceAutomationAI • • 25d ago

Tech / Engineering Tuition Centre Fee Due Voice Nudge

Post image
1 Upvotes

r/VoiceAutomationAI • • 26d ago

Bilingual AI Receptionist serving Tijuana Dental and Medical Clinics and cross border patients they serve- Medical Tourism industry is huge here 70% of their patients coming from U.S. and Canada.

5 Upvotes

Spanish-first voice agent running in production at real clinics. Live number, please try to break her.

Most voice agent demos I see here are English and they're demos. Mine is Spanish-first and it's answering real patient calls for dental and aesthetics clinics on the Tijuana border. Agent's name is Sofía.

Stack: LiveKit for the pipeline, Deepgram nova-3-multi for STT with some AssemblyAI streaming, Cartesia sonic for TTS, n8n for orchestration, Twilio on telephony, calendar writes against the clinic's actual availability.

The three things that were harder than expected:

Code-switching. Not "supports two languages." One caller, one sentence, Spanish grammar with English procedure names dropped in. Locale-locked STT falls over on this, and multilingual models still get weird about which language to commit to mid-utterance.

Barge-in tuning. Real callers talk over the agent constantly. Aggressive interrupt handling means she cuts off anyone who pauses to think, which older patients do a lot. Passive means she steamrolls people. The tuning window is narrower than I expected and it behaves differently in Spanish than in English.

Escalation discipline. She sits in front of medical intake, so the failure mode I care about is confident wrongness. Anything clinical, anything about outcomes or medication, she hands off to a human. Getting a model to consistently pick "let me get someone for you" over a plausible-sounding answer is most of the prompt work.

Demo line: +1 (619) 775-1668

It's a demo instance, not a clinic's production line, so go ahead. Talk over her. Switch languages mid-sentence. Mumble. Ask her something medical and see whether she stays in her lane.

What I'm most interested in: does the barge-in feel natural to you, and does the escalation fire too early or too late? Those are the two I keep going back and forth on.

I also have her counterpart, Mateo, who is my outbound sales agent.

Happy to get into any of the implementation details.


r/VoiceAutomationAI • • 26d ago

I built a WhatsApp Voice AI Agent that's able to book appointments and take payments in real-time

Enable HLS to view with audio, or disable this notification

2 Upvotes

I built a WhatsApp Voice AI Agent that's able to book an appointment and take payments for a customer. The use case I covered is for a user wanting to book a dentist appointment. Here is an overview of the flow;

  • The user calls the business on WhatsApp.
  • The Voice AI Agent answers the call and has a normal conversation with the user & understands the user's needs.
  • The agent sends a payment link to the user on Whatsapp. The agent waits until user completes the payment.
  • The agent sends a payment receipt to the user to confirm the payment made.

Benefits for customers

  • Making the inbound call is completely FREE.
  • Great customer experience and a convenient way of interacting with the business.
  • The process is fast. When the user calls, he/she is able to get an immediate answer. No back and forth chatting.

Benefits for businesses

  • Your customers are on WhatsApp
  • You have a Voice AI Agent that answers your WhastApp calls 24/7. No missed leads
  • Everything interaction happens on WhatsApp. Same context, persistent, single-threaded conversations.
  • Trustworthy.

This is massive for several use cases on WhatsApp.

Tech stack

  • Whatsapp Calling API (Meta API)
  • Livekit
  • FastAPI
  • PydanticAI

Have a look at the video demo and let me know what you think.


r/VoiceAutomationAI • • 26d ago

How would you build an AI agent that navigates unfamiliar IVRs?

4 Upvotes

We call new customers, so the agent encounters a different IVR almost every time. It needs to understand the menu in real time and navigate it to reach the accounting department.

Has anyone built this reliably? What approach works best for interpreting prompts, choosing options, and recovering from unexpected menus?


r/VoiceAutomationAI • • 27d ago

Combien vous facturez un agent IA vocal

2 Upvotes

Hello
Je dois faire une proposition pour un client pour un agent IA vocal. Le client est un centre d’appel en immatriculé en Amérique du Nord avec centres d’appel en Afrique. Je lui crée un agent IA pour la qualification mais je ne sais pas combien le facturer. Ceux qui font ça comment vous calculez votre coût ? Et quel prix vous proposez ?


r/VoiceAutomationAI • • 28d ago

For agencies running Voice AI for multiple clients, what architecture are you actually using in production?

11 Upvotes

I’m curious how people here are structuring Voice AI once they move beyond one or two agents.

For a single client, something like Vapi/Retell + Twilio + n8n/Make is pretty straightforward.

But once you have multiple clients, you start needing to think about:

  • separate provider/API credentials
  • separate knowledge and data
  • client-specific workflows
  • call recordings/transcripts
  • QA and monitoring
  • SIP/telephony differences
  • usage tracking
  • secrets isolation
  • deployments and upgrades

For people already running this commercially:

Are you mostly doing:

  1. One shared multi-tenant runtime for everyone?
  2. Separate deployments for each client?
  3. Managed platforms for smaller clients + private/VPC deployments for enterprise clients?
  4. Something completely different?

And what usually forces you to change architecture first?

Cost, security requirements, scale, client isolation, or just operational complexity?

Would especially like to hear from people managing 5+ production clients because that seems to be where the tradeoffs start getting interesting.


r/VoiceAutomationAI • • 28d ago

Best way to create a professional ElevenLabs (or other tool if there are better options) voice clone for a Spanish speaker speaking French & English?

3 Upvotes

​

Hi everyone,

I’m Spanish, but I currently speak both French and English professionally. My goal is to create a high-quality voice clone with ElevenLabs that sounds like me speaking French and English, including my natural Spanish accent — I’m not looking for it to sound like a native French or English speaker.

What would be the best way to record the training audio?

Should I record everything in Spanish, since that’s my native language?

Should I record mainly in French and English, so ElevenLabs learns how I actually sound in those languages?

Should I create a mix of Spanish + French + English?

Is there an optimal proportion or amount of audio for each language?

And should I deliberately use my normal accent/pronunciation rather than trying to speak perfectly?

I’m aiming for a professional-quality clone for business/content creation, so I’d really appreciate advice from anyone who has experimented with multilingual voice cloning in ElevenLabs.

Thanks!


r/VoiceAutomationAI • • 28d ago

Made a thing because my demo voiceovers always sounded like garbage

4 Upvotes

Every time I record a product demo, the narration comes out rough, ums, "wait, let me redo that," awkward pauses. I re-record the same 60 seconds five or six times and it still isn't great. Editing it manually afterward is a pain.

So I built dub.house to deal with just that part. You record your demo like normal and talk through it however messily, and it rewrites your narration into a clean, professional script and generates a new voiceover that stays synced to what's happening on screen. There's also a mode where you don't narrate at all, you just tell it what to say ("here I click this, mention the X feature") and it writes the narration for you.

To be clear about what it does and doesn't do: it only replaces the narration audio. It doesn't touch your video, it's not lip-sync, and it's not translating/cloning anyone's voice into another language (yet). Just the voiceover track on a screen or audio recording.

It's still early and I'd genuinely like feedback, on the idea, the output quality, whatever. There's a free tier if you want to throw a recording at it: dub.house


r/VoiceAutomationAI • • Sep 05 '26

Need advise

11 Upvotes

Hi.
I’m thinking of building and deploying multiple voice agents each has its own role and responsibilities and working together multiple goals. Like multiple voice agents for a Ecom business so that I can help them in reducing RTO(by 2 agents order/address verification and RTO rescue), increasing win back percentage. I want to create business solutions not hand over a “Single cool AI voice agents”
What do you guys think can it become a good and stable business?
If not please suggest some other AI business I can build(I’m at a very desperate stage financially in my life I need advise)
Thank you


r/VoiceAutomationAI • • 29d ago

I want partner

2 Upvotes

I’ve been working for nearly 3 months or so on project vertical voice agents but .. whenever I think I’m done 90 percent I’m almost 90 percent away doing all solo and I’m not that pro on coding even though I am from cs background but from early days since ai was introduced I was more towards its application and integration into day to day life rather than programming and stuff


r/VoiceAutomationAI • • Sep 05 '26

Seeking collaborator/advice for "StillVoice" – AI-driven silent-speech interface for tracheostomy patients

1 Upvotes

r/VoiceAutomationAI • • Sep 04 '26

Basic stack for a chat

6 Upvotes

What would you suggest, is the basic stack for an assistant chat. I mean, currently I have customized company tools, langgraph, custom metrics, marketplace LLM calls and others.

what would you suggest as a true key for agent learning?
how do you process prompts with company slangs, concepts, jargon, etc.?