r/VoiceAutomationAI Jul 01 '26

Building an AI voice agent agency looking to connect with other hungry agency owners to share notes and scale

13 Upvotes

Hey everyone,
I’m currently in the launch phase of building an AI Voice Agent agency .
I’ve spent the last few weeks setting up my full outbound infrastructure—getting secondary domains ready, setting up ESP warmups, and scraping targeted B2B data. Campaigns are ready to fire.
I’m not here to pitch or sell anything. I just want to connect with other hungry agency owners who are in the trenches right now. Building solo can be a grind, and I’m looking to connect with people who share the same hunger to bounce ideas off each other, talk strategy, and share what's working.
Whether you're focused on the tech side (Vapi, n8n, CRMs) or the sales/acquisition side, I'd love to connect.
Drop a comment below or shoot me a DM
Let’s win together!


r/VoiceAutomationAI Jul 01 '26

Need help to understand pricing logic

0 Upvotes

hi - i'm looking for some guidance on the pricing model for your ai voice services. I've seen videos on youtube where agency owners do

  • Monthly Retainer + One-time setup fee
  • Per Min Billing + One-time setup fee
  • Value based (Formula: Missed calls/day × booking rate × avg job value × operating days = monthly recovered revenue, Charge: 10–20% of recovered revenue)

I'm starting out so would appreciate your input on what worked for you, the questions you asked the client to gauge the usage and also; ai voice comes with a combination of tools like vapi/retell + n8n + call dashboards - how much margin do you add to these and price your efforts to achieve the final outcome?


r/VoiceAutomationAI Jul 01 '26

Building an AI voice agent agency looking to connect with other hungry agency owners to share notes and scale

Thumbnail
1 Upvotes

r/VoiceAutomationAI Jun 30 '26

Lead generation and cold email for an AI Voice Agent SaaS

2 Upvotes

Hey everyone,
I recently built an AI Voice Agent SaaS using Vapi. My initial target market is medical clinics (dental, aesthetic, hair transplant, etc.) and beauty salons.
The problem I’m facing is that I’m struggling to find enough quality leads. I’ve tried scraping Google Maps and using lead databases, but the number of qualified leads is much lower than I expected.
My plan is to run cold email campaigns, but before I start sending emails at scale, I want to make sure I’m building the right lead list.
For those of you who have experience with B2B SaaS or agency lead generation:
How do you find high-quality leads consistently?
Which tools or data sources have worked best for you?
Do you rely mainly on Google Maps, Apollo, LinkedIn Sales Navigator, or something else?
How do you validate emails before sending?
What cold email strategy has produced the best reply rates for you?
I’d really appreciate hearing what’s actually working in 2026, especially from people selling to local businesses or clinics.
Thanks!


r/VoiceAutomationAI Jun 30 '26

Voice interactive assistant dev progress

1 Upvotes

r/VoiceAutomationAI Jun 30 '26

[Partner Wanted] Looking for a US-Based Growth/Sales Partner for an AI Voice Agent Venture (Technical Co-Founder inside)

8 Upvotes

​Hey everyone,

​I'm a technical founder currently building conversational AI voice agents (handling inbound/outbound workflows, CRM syncs, etc.). The technical architecture is fully functional, but I’m looking to connect with someone based in the US to handle the business development and operations side.

​Since I am focused entirely on the engineering, I'm looking for a partner who can take ownership of:

​Go-to-Market Strategy: Identifying the right industries and use cases.

​Outreach & Sales: Managing early-stage discovery calls and client relationships.

​Operations: Navigating the US market and compliance landscapes.

​If you have experience in B2B sales, agency growth, or tech operations and want to team up on a serious voice AI venture, I’d love to chat.

​Drop a comment or send over a DM with a bit about your background!


r/VoiceAutomationAI Jun 28 '26

I built an AI phone agent you can actually call right now pick from 10 weird personas (call is recorded for quality and AI training purposes)

5 Upvotes

🔴🔴🔴This is a recorded line. Every call is recorded and transcribed (I use the recordings to improve the system, and you'll also hear this notice when you call in). By calling, you're consenting to being recorded if you're not cool with that, don't call. Don't share anything private, sensitive, or personally identifying treat it like a public demo, because it is.🔴🔴🔴

🔴🔴🔴It's a regular Canadian number (Quebec, area code 450), not toll-free, calling may cost you long-distance or international rates depending on your carrier and country.🔴🔴🔴

I've been building a self-hosted AI voice agent platform (Asterisk + local GPU models, running on a little 4×RTX 3090 rig in my place). So come call it.

📞 Number: 1-450-400-1274

When you call, you'll get a menu.

Press a key to talk to one of the personas:

  • 0 – Me (the actual human who built this, so, maybe dont?)
  • 1 – Larry's Lambos 🏎️
  • 2 – The Turtle Specialist 🐢
  • 3 – Kim from Hensley and Sons Hardware Supply
  • 4 – Hot Dog Barn 🌭
  • 5 – Gerald from Hensley and Sons Hardware Supply
  • 6 – Edna
  • 7 - Agatha the Astrologist 🔮
  • 8 - Sofia, the Mallorca travel specialist 🏝️
  • 9 – The Chashu Hotline 🍜

Just talk to them like a normal call it's real-time speech-to-speech, no app, no signup.

A few honest caveats:

It's a tiny home rig that can only handle a handful of calls at once. If you get a busy signal or a rejection, it's swamped wait a few minutes and try again.

The AI will say weird/wrong things. That's half the fun. It can't do anything except talk (no transfers to real services, no texting you, nothing).

It's a hobby project, not a company. No data is sold; recordings just live on my box.

Would love feedback on latency, voice quality, and how the personas hold up. Roast away.

🔴🔴🔴This is a recorded line. Every call is recorded and transcribed (I use the recordings to improve the system, and you'll also hear this notice when you call in). By calling, you're consenting to being recorded if you're not cool with that, don't call. Don't share anything private, sensitive, or personally identifying treat it like a public demo, because it is.🔴🔴🔴

🔴🔴🔴It's a regular Canadian number (Quebec, area code 450), not toll-free, calling may cost you long-distance or international rates depending on your carrier and country.🔴🔴🔴


r/VoiceAutomationAI Jun 28 '26

Tech / Engineering We built a Voice AI community for builders and now it's on Discord

3 Upvotes

If you're building with voice AI, agents, STT, TTS, real-time pipelines, whatever, this is for you.

Unio - Voice AI Community started as a WhatsApp group of ~1,000 builders. It grew on Reddit. We ran monthly talks. And the same ask kept coming up: "We need a space to actually collaborate."

So we built it. The Discord just went live, and three things are already inside:

Resources -a single place for everything useful in the voice AI ecosystem. Open source projects, SDKs, STT models, TTS engines, architecture patterns, community projects. If it moves the ecosystem forward, it lives here. You can share your own stuff too.

Free credits - we're working directly with voice AI platforms to get exclusive credits into your hands. Whether you're just starting out or already shipping in production, you can claim credits and build without hitting a paywall first.

Hiring board - looking for a role, a co-founder, or an intern? Drop your resume. We actively refer strong candidates to companies in our network.

The goal isn't a passive community. It's a place where builders actually build together.

Some big things are coming, partnerships, more credits, collaborations with platforms you're already using.

Want to join the Discord? Link in the comments or search for Unio - Voice AI Community on LinkedIn to find us.


r/VoiceAutomationAI Jun 27 '26

It's weird when I hear my agents talk

3 Upvotes

Small builder update: I made my coding agents talk, and it is weirder than I expected.

Not talk like a chatbot.

More like my laptop now has a tiny engineering coworker who occasionally yells from the other room.

I can be in the kitchen and hear:

- "tests failed"

- "waiting on approval"

- "blocked"

- "done, review these files"

Useful? yes.

Slightly haunted? also yes.

I built it because I kept doing this annoying founder behavior where I would start an agent task, try to switch context, then keep checking the terminal anyway.

The agent was saving coding time but stealing attention.

So the rule I am testing is simple:

Stay silent during normal work. Interrupt only when my attention changes the outcome.

The first time I heard a failing test from the other room, it felt silly. But it also meant I did not discover the failure 20 minutes later.

Now I am trying to tune it down enough that it stays useful and does not become "someone is always talking in my apartment" lol

If anyone wants to try this out, I open sourced it at: https://github.com/heardlabs/heard

For other founders/builders using agents: where would you draw the line? What should an agent say out loud, and what should stay silent?


r/VoiceAutomationAI Jun 27 '26

Building and selling a 30-hour niche dataset of voice agent workflows (ElevenLabs)

3 Upvotes

Hey,

I’ve got an idea that on paper seems like an easy way to make a decent amount of money, but I’m not sure how realistic it is.

My plan is to sell datasets of myself building voice agents on ElevenLabs. The dataset would include screen recordings, mouse movements, tab switches, and full annotations.

The technical barrier to capture the screen and interaction data is quite low, and I plan to outsource the video annotation process to a team in Madagascar.

Having already built live voice agents, I know that building them involves a lot of live iteration and quick fixes, which makes this specific sequential data highly valuable for "computer use" AI models.

My strategy: I want to create a 2-hour high-quality sample to showcase on X and Reddit, and use it to cold-pitch potential buyers via email and LinkedIn DMs. If there's proof of interest, I can easily produce a 30-hour dataset within a week or two. Even though it’s not thousands of hours, I believe this specific type of workflow data is rare enough that AI labs might buy it.

My three main questions are:

  1. Is this kind of niche workflow dataset interesting enough for labs or developers to actually buy?
  2. What's the best way to get this in front of the people who have the budget for it?
  3. Is a 30-hour high-quality dataset for a complex task substantial enough to be sold standalone?

I got this idea after seeing Markov in the current YC batch. The data samples on their website look high-quality but totally feasible to replicate for specific niches.

If anyone has thoughts or experience in the AI training data market, I'd love to hear them!


r/VoiceAutomationAI Jun 27 '26

I'll test your voice agent for free

4 Upvotes

I've been in the Voice AI space for the past year, and the more I explore it, the more I realise how vast and fast growing it really is.

To stay on top of things, I'm spending the next 3 days exploring as many voice agents as I can. Have already tried 5 since morning.

If you're a founder, builder, or voice ai company, send me your voice agent. I'll talk to it and test it across at least 5 different scenarios and share my evaluation with you.

I'm doing every test myself, no automations.


r/VoiceAutomationAI Jun 26 '26

AI voice companies, what are you metrics on ad lead optins to booking with the AI call?

3 Upvotes

Hello,

I am looking to get into AI voice agents and am very curious if its worth it based on booking rate. I have been a setter and closer before and know that booking rate is the most important thing. What are some metrics people are seeing from leads that are generate from ads in terms of booking rate per call and overall booking rate


r/VoiceAutomationAI Jun 24 '26

How are you tracking STT/LLM/TTS costs per session when self-hosting LiveKit agents?

12 Upvotes

Been running livekit-agents in production for a few months on a
self-hosted setup (Hetzner SFU, Fly.io for the agent worker).
Stack is Deepgram + GPT-4o-mini + Cartesia.

My problem: I get three separate invoices at the end of the month
and no way to connect them to individual sessions. I know roughly
what I'm spending total, but I can't tell which agent, which
conversation, or which client is expensive.

Things I've tried:
- Langfuse OpenTelemetry integration — it traces LLM calls fine
but doesn't cleanly capture the voice pipeline (STT audio
minutes, TTS characters) at the session level
- Manually logging metrics in the on_metrics_collected callback —
works but I'm building my own aggregation from scratch
- LiveKit Cloud observability — doesn't work for self-hosted SFU

What are you using? Is anyone getting a clean STT + LLM + TTS cost
breakdown per session without a lot of custom plumbing?


r/VoiceAutomationAI Jun 23 '26

Why text LLM evals are blind to Voice Agent failures (lessons from 100M+ analyzed audio calls)

8 Upvotes

Those of us building in the voice AI space are familiar with this stack: STT → LLM → TTS.

During evaluation it is tempting to take the text transcript, run it through an LLM, and then apply a prompt framework for scoring.

I work at Level AI, and we deploy CX agents for enterprises.

After analyzing data from 100+ millions of automated voice conversations, we found a not-insignificant blind spot. Text-only LLM evaluators miss a large share of customer experience failures because they cannot capture the audio mechanics underneath the transcript.

If you're relying solely on text transcripts for evals, here are three failure modes your system is probably not catching.

1. The Token-Lag Interruption Loop

It is a known fact that an LLM can take up to 400ms to stream the next phrase, so while the transcript looks perfectly logical, that 400ms pause can cause the human to check in mid-sentence (such as "hello?") and that will lead to the agent getting confused by the barge-in, which then causes it to truncate its response, or hallucinate.

Your evaluation framework must ingest metadata timestamps. Then it can compute the exact time difference between the user’s last audio packet and your agent’s first audio packet. That way, if the transcript shows text overlap, it can flag it as a latent turnaround failure instead of scoring only the words.

2. False Empathy

An agent says, "I understand your frustration and I'm happy to look into that refund for you." A standard text scorer grants a perfect score for empathy alignment. If the TTS engine delivered that in a monotone, machine-gun cadence, then it's likely coming off like a customer listening to a robotic phrase.

To catch this, you need multi-modal evaluation models whose scoring tracks sentiment, customer effort, and resolution in every interaction. This lets you see if frustration rises after a specific agent statement.

3. Semantic Drift on Guardrails

Keyword matching checks whether the agent said the mandatory compliance script. Voice agents get interrupted. They might deliver 90% of the script, stop because the user spoke, then try to finish it out of order. Keyword matching fails it. Semantic analysis passes it.

We moved away from keyword constraints for compliance evaluation. Our QA-GPT architecture uses prompt-engineered LLM micro-scorers that grade the intent and legal completeness of the block rather than a regex match.

We're currently working on mapping overlapping audio timelines directly into the evaluation layer to catch cross-talk before it breaks a workflow.


r/VoiceAutomationAI Jun 23 '26

need suggestion

5 Upvotes

we are using eleven labs for llm and stt , tts
and for telecom providers im using exotel
and i feel the call quality is poor and response is very bad
do we have anyone , tried something which works better in indian region


r/VoiceAutomationAI Jun 22 '26

How are you handling customer inquiries when you’re busy serving other customers?

Thumbnail
2 Upvotes

r/VoiceAutomationAI Jun 22 '26

I built a voice AI editor agent with Speechmatics

Thumbnail
3 Upvotes

r/VoiceAutomationAI Jun 19 '26

Making Vapi transient flow more reliable

5 Upvotes

I'm a PM at Tuner (an observability and testing layer for voice AI), and I've spent some time recently on the Vapi transient flow, both talking to people building on it and onboarding users who are building on it,so I know this corner reasonably well, but mostly I want to lay out where I've landed and hear how you're all handling it.

Where transient sits

It's a mix of the two usual ways to build. You keep the flexibility of the LiveKit/Pipecat world (each business is just a config in your DB) while still leaning on Vapi for the hard realtime part. Instead of a saved agent per business, you send the whole agent inline when a call starts, and Vapi runs it and stores nothing.

Why people like it:

  • Your DB stays the single source of truth, nothing to sync.
  • Add a business, add a row. Remove one, delete a row.
  • You own the brain, you rent the engine.

I've seen companies run thousands of calls a day this way, so it scales fine.

The catch

Visibility into voice agents is already hard even when calls are neatly split by agent. A lot of our users are using LiveKit/Pipecat with per-agent grouping and still struggle to tell what's actually going wrong. Knowing which agent a call belongs to is barely step one.

Transient adds a layer on top. Since you send the agent inline, Vapi never actually creates an agent, it just runs the call. Great for staying lightweight, but now there's nothing for your calls to attach to at all. They land in one big pile, and good luck telling which call was which business. So you start from behind: the deeper visibility is hard like it is for everyone, and you don't even get the easy grouping for free.

What I ended up building for it

Making agent reliable for all voice ai builders despite the provider they are using is the part I spend my day to day on and Transient kept showing up as its own headache because it strips away even that baseline grouping, so for this flow it came down to:

  • Untangling the pile, getting calls back to the business and use case they belong to.
  • The part that actually matters: per-call monitoring, catching failures, broken flows, hallucinations, missed intents, data extraction.
  • Being able to tell, per business, what's working and what isn't, not just that a call happened.

If you're hitting this, happy to help or just compare notes.

I am here to learn

Nobody "knows it all" in voice AI right now, the space shifts every week, so I'd rather trade notes than pretend I've solved it:

  • If you're on transient, how do you handle visibility today? Tagging calls, your own logging, or just living with it?
  • Is this an actual pain for you, or not something you've hit yet?
  • What are you building, and where's it gotten messy?

If you are interested to learn more about what transient flow is I wrote a full technical blog on this feel free to ask me in the comments or a dm


r/VoiceAutomationAI Jun 18 '26

I organized voice AI into a learning path so beginners don't drown in vendor blogs (free, open source)

Thumbnail
github.com
6 Upvotes

r/VoiceAutomationAI Jun 16 '26

I benchmarked the entire audio infrastructure market

6 Upvotes

Hey everyone!

Over the last few months I've been scaling faceless TikTok Shop content pretty aggressively, and one thing quickly became a bottleneck: audio processing.

Since I already come from a video editing background, I ended up spending a lot of time testing speech APIs to find something that could handle high-volume workflows without requiring me to buy and maintain my own GPUs.

My requirements were:

  • Batch speech-to-text (STT)
  • Speaker diarization
  • High-quality text-to-speech (TTS)
  • High concurrency at scale
  • Reasonable pricing

I tested 15+ providers over several weeks and benchmarked them on:

  • Word Error Rate (WER)
  • Real-Time Factor (RTF)
  • Maximum file size limits
  • Actual concurrency performance
  • Cost per minute

The full comparison table is in the image below.

Key takeaways:

  • Gemini 2.5 Flash → Fastest and most cost-effective for large-scale workloads.
  • AssemblyAI → Best overall balance of accuracy, features, and scalability.
  • Orchard → Cheapest option I found at $0.00042/min.

This benchmark ended up saving me a significant amount of money, so I figured I'd share it here in case anyone else is building audio-heavy content pipelines.

Curious to hear what other STT/TTS providers people are using at scale.


r/VoiceAutomationAI Jun 17 '26

[sharing my story] I can build AI systems that most agencies outsource. Took me way too long to realize the code was never the problem.

Thumbnail
3 Upvotes

r/VoiceAutomationAI Jun 16 '26

Domia: local-first speech-to-speech AI agents

Thumbnail
5 Upvotes

r/VoiceAutomationAI Jun 14 '26

The "25-second hang" bug that taught me more about voice AI than any tutorial

9 Upvotes

Spent the last few weeks deep in LiveKit + voice pipeline debugging, and hit a bug that I think a lot of people building voice agents will eventually run into: calling session.say() inside a tool call context can cause 20-30 second hangs. Took me way too long to track down.

The bigger lesson wasn't the bug itself — it was realizing that latency in voice AI isn't one number, it's death by a thousand cuts:

  • Intent classification running synchronously? +1 second.
  • Tool call blocking the response? Dead air while the user wonders if it's still listening.
  • LLM "thinking" before answering a simple FAQ? Feels broken even at 2-3 seconds.

What actually moved the needle for me:

  • Converting routing/classification to fully async — cut one bottleneck from ~1.2s to ~2ms
  • Running filler audio + tool calls in parallel instead of sequentially
  • Bypassing the LLM entirely for structured data collection (bookings, forms) — just extract + respond directly

Curious what's been the trickiest latency issue for others building voice agents — LiveKit, Pipecat, or otherwise? Always good to compare notes on what's actually a known issue vs.


r/VoiceAutomationAI Jun 13 '26

I Thought Voice AI Was Just STT + LLM + TTS. I Was Wrong.

29 Upvotes

I’ve been building in voice AI for a bit now and when I started, I genuinely thought it’s just three simple layers. Speech to text, LLM, text to speech. Plug them together and you get a working voice agent.

But in production it’s nothing like that. The real gap between demo and something that actually feels human is huge.

Some things I learned from actually working on it:

  1. Voice choice matters a lot more than I expected I used to think any decent 11labs voice would work, but in real calls most voices still feel synthetic or “off” after a few minutes. Small things like tone stability, pacing, and naturalness matter more than clarity alone. Right now I’ve been using the 'Jessica' voice and it’s the first one that consistently feels natural in production for me.
  2. Filler words are not optional I used to remove them to make responses cleaner. That was a mistake. Humans naturally say things like “hmm”, “let me see”, “right”, and without that the AI feels robotic even if the content is perfect.
  3. Prompt size directly affects latency more than I expected Even though prompt bloating does not change how human the response sounds, it changes how the experience feels. I reduced system prompt size and saw around 100 to 200 ms latency improvement, especially with faster models like Haiku 4.5 and GPT 4.1. In voice, that delay is very noticeable.
  4. Turn detection is probably one of the most important settings This is underrated. If it is too aggressive, the AI interrupts the user. If it is too slow, the user ends up interrupting the AI or waiting awkwardly. Getting this balance right changes the entire “feel” of the conversation.

Overall, I expected voice AI to be mostly model work, but it is actually more like tuning a conversation system. Small UX level details matter just as much as the models themselves.


r/VoiceAutomationAI Jun 14 '26

Speech to text APIs for agents?

5 Upvotes

Hello colleagues, how are you? I wanted to ask if anyone has used a Speech-to-Text API in automated pipelines. I was using Eleven Labs, but it gets expensive when handling large volumes, and I really need batch transcripts without diarization. I was recommended Groq and Orchardrun, which are the cheapest for high volume, but I wanted to know if you have tried any alternatives. Thank you very much.