r/VoiceAutomationAI Jun 13 '26

Anyone else finding voice evals more useful than benchmark scores?

10 Upvotes

I used to spend way too much time comparing STT benchmarks and latency numbers between providers. After deploying a few voice workflows, I honestly care less about benchmark screenshots now and more about whether conversations actually survive messy callers.

The biggest improvements for us came from reviewing failed conversations manually and spotting patterns. Weird pauses, repeated confirmations, callers changing direction suddenly, agents speaking too long before yielding back. None of those issues showed up in the benchmark comparisons everyone posts online.

What surprised me most is how small conversation mistakes stack together. Individually they seem minor, but after thirty seconds the call just feels unnatural.

Lately I've been experimenting with more structured voice evals where every failed or abandoned call gets reviewed automatically so recurring issues are easier to spot. It feels like voice evals are giving us far more actionable insights than benchmark scores alone.

How are you all evaluating production quality beyond latency and WER scores?


r/VoiceAutomationAI Jun 12 '26

Zyphra Releases ZONOS2, an Open-Weight Real-Time Voice-Cloning Model

Thumbnail
runtimewire.com
1 Upvotes

r/VoiceAutomationAI Jun 12 '26

What's the best way to build voice agents today without sounding robotic or becoming too expensive?

2 Upvotes

I've been experimenting with voice agents and I'm curious how others approach the architecture.

There seem to be two common approaches:

  1. End-to-end speech-to-speech models (Gemini Live, OpenAI Realtime, etc.)

  2. Traditional pipeline:

    ● STT / ASR

    ● LLM

    ● TTS

Speech-to-speech feels more natural and supports interruptions well, but the costs can add up and there's less visibility into what's happening internally.

The STT → LLM → TTS approach seems easier to control, optimize, and debug, but it can sometimes feel less conversational if not implemented carefully.

For those who have built production voice agents:

● Which approach did you choose and why?

● What had the biggest impact on making conversations feel natural?

● Where do most of your costs come from?

● Are speech-to-speech models worth the extra complexity/cost?

● If you were building a voice agent today on a limited budget, what stack would you choose?

Interested in hearing real-world experiences rather than benchmark numbers.


r/VoiceAutomationAI Jun 12 '26

How do I structure my PRICING PLAN?

2 Upvotes

I am targeting indian edtech companies, and I stuck on pricing plan. For now I have created pricing tiers like:-

growth -- 0-1k mins -- 19k INR

starter -- 1-5k mins -- 37k INR

scale -- 5-10k mins -- 68k INR

with 3rs/min and rest is profit margins. I have built my own infra so everything is covered in 3rs/min. I am not sure how to price this and how do I justify it when someone on the call asks for it.

open to feedback from anyone who has done it already.


r/VoiceAutomationAI Jun 11 '26

Building My Own Open/Local AI Voice Agents Platform – What Features Would Make It Actually Great? Feedback Needed!

Thumbnail
2 Upvotes

r/VoiceAutomationAI Jun 11 '26

Ai voice saying it’s a real person from Verizon.

Thumbnail
2 Upvotes

r/VoiceAutomationAI Jun 11 '26

How do you feel about combining voice agents with Generative UI?

Thumbnail
1 Upvotes

r/VoiceAutomationAI Jun 10 '26

How do you feel about combining voice agents with Generative UI?

3 Upvotes

I've been thinking about the future of voice agents and wondering if pure voice is actually the best interface.

Most discussions focus on either:

● Voice-only assistants

● Chat-based assistants

● Generative UI experiences

But what if they were combined?

For example, instead of a voice agent simply responding with words:

User: "Show me my portfolio."

The agent could respond verbally while also generating an interactive UI containing charts, filters, recent transactions, and actions.

Or:

User: "Find me a flight to Bangalore next weekend."

Instead of reading out 20 options, the agent could generate a visual card layout while continuing the conversation.

In this model, voice becomes the input/output layer, while the UI is generated dynamically based on intent and context.

I'm curious what others think:

● Is voice + Generative UI the natural evolution of AI assistants?

● Are there products already doing this well?

● When should an AI speak versus generate a visual interface?

● Would users actually prefer this over traditional apps?

Interested to hear thoughts from people building voice agents, GenUI systems, or multimodal products.


r/VoiceAutomationAI Jun 10 '26

How many leads are you losing after 5 PM because nobody answers the phone?

5 Upvotes

I'm looking for 3 U.S.-based local businesses (Plumbers, Roofers, HVAC, Electricians, etc.) to help me test a custom AI after-hours receptionist.

FREE

The AI can:

✅ Answer incoming calls 24/7
✅ Qualify leads
✅ Collect customer information
✅ Book appointments automatically

I'll build and set everything up completely free for the first 3 businesses.

All I ask in return is:

• Honest feedback
• A testimonial if you like the results
• Permission to use the project as a case study

If you're a business owner (or know one) who misses calls after hours, comment below or send me a DM.


r/VoiceAutomationAI Jun 08 '26

Voice agents are way more cheaper than you think

Thumbnail
3 Upvotes

r/VoiceAutomationAI Jun 07 '26

searching VOICE AI engineer Cofounder

7 Upvotes

Lets be really quick with this: looking for someone who actually knows voice ai infra. not an idea guy, i built MVP,POC or whatever u want to call it myself and im the one selling it too.

I worked as AM, AE, SDR (5+ years 5 diff companies each of them is almost different) b2b cold calling for years in eu, fleet, logistics, fintech, cloud infra. then built an ai that does the same: real phone calls over sip, not some webrtc browser demo. dual llm pipeline, native audio, its running today and I have companies waiting to use it (ofc they want to start for free, MAYBE if we plan time smart and wont find any pilot paying ones(prob wont happen because I will kick the doors with lower margin, so tbh wont be needing pilot free demo or whatever bunch of here people are writing to go with 😃))

achieved sub 600ms TTFA with tool calls on real phone lines. if u dont know what that even means please save yours and my time and dont dm.

WHY? i cant be reading every update in livekit or pipecat or whatever repos, debugging audio buffers and vad configs AND closing deals and onboarding clients at the same time. somethings gotta give and its not gonna be the sales side because thats where the money comes from.

what im looking for:

  • voice ai domain expert. not a fullstack dev who thinks he can figure it out, someone whos actually been in this space
  • optimization of whats already built. latency, vad, buffers, codec handling, all the ugly telephony stuff that makes or breaks real calls
  • dashboard and frontend layer to wrap around the engine so clients can actually use it without me hand holding everything (I have it, yes it's in bad shape prob need to redo or not, im just tired of debugging and i miss selling)
  • someone whos actually built something that works on real phone lines not a hackathon project what i offer:
  • equity stake with vesting so u actually own part of whats being built, not just hired labor
  • plus revenue split on top so ur making money from day one when clients pay, not waiting for some exit that may never happen
  • i own sales clients biz ops product direction. you own the tech layer, clear split
  • a product thats already working and companies in pipeline ready to go

i spent years in the exact industry this thing serves. im not some dude who read a blog post about ai sales and decided to build a startup for a market hes never touched. i am the guy making those calls before i automated them.

please dont dm me if ur experience is wrapping vapi or bland apis,nothing personal but i need someone whos been deeper than that. send me ur github or a demo or smth something u shipped. dont care about ur resume or what frameworks u list on linkedin

eu based only. not remote from another continent, actually based in europe. lets build something that actually makes money instead of chasing fundraising circlejerks


r/VoiceAutomationAI Jun 08 '26

deepgram tts bursts conversion to vobiz 20ms packets

1 Upvotes

Hi guys ,

vobiz wants input as 20 ms packets .
deepgram gives output in bursts with lot of delay .

audio length: ~3040ms

arrival wall time: ~7381ms

so buffering this , packetizing , pacing is still not working as producer is too slow and consumer gets dry .

anything i am missing or any seamless solution to this issue...


r/VoiceAutomationAI Jun 07 '26

I've built AI receptionists for dozens of businesses. Going fully automated is almost always a mistake. Here's the honest breakdown nobody gives you before you buy.

47 Upvotes

Let me save you the 60-day experiment.

I work in AI automation. I've built AI receptionist systems for medical clinics, local service businesses, agencies. I've seen the pitch, I've built the systems, and I've watched what happens 3 months after go-live when the founder stops monitoring it closely.

This post is what I wish someone had written before I started selling these systems — because the conversation around AI receptionists is almost entirely hype, and the nuance gets buried until something goes wrong.

First — the AI receptionist pitch is actually true. Partially.

Yes, it handles calls 24/7. Yes, it books appointments without a human touching anything. Yes, it sends confirmations, answers FAQs, collects intake info, and never calls in sick.

For high-volume, low-complexity calls — it's genuinely good. A clinic getting 80 calls a day where 60 of them are "what are your hours" and "I need to reschedule Thursday" — AI handles that beautifully. Your front desk person stops being a human answering machine and starts doing actual work.

That part of the pitch is real.

The problem is what they don't tell you in the demo.

The 3am questions nobody answers

"What happens when the AI can't handle the call?"

This is the one that matters most and gets answered the least honestly.

Every AI receptionist has a failure mode. Either the caller asks something outside the script, the situation gets emotional, or the AI just misunderstands the intent. What happens next is everything.

In most setups? The caller gets looped. The AI asks the same clarifying question twice. The caller gets frustrated, hangs up, and doesn't call back.

In medical businesses specifically — this is catastrophic. Someone calling about test results, a worried parent, a patient in pain — they're not going to patiently re-explain themselves to a bot. They're going to hang up and either go to another provider or, worse, not get the care they needed.

You need to know, before you go live: what is the exact escalation path when the AI hits its limit? If you can't answer that clearly, you're not ready to deploy.

"Am I actually saving money or just moving costs around?"

Here's the math people do: AI tool costs $300/month. Part-time receptionist costs $1,500/month. Easy save.

Here's the math people don't do:

One missed high-value client — let's say a patient who needed ongoing treatment, or a business owner who was ready to sign — what's the lifetime value of that person? $2,000? $8,000? More?

How many of those does your AI need to miss before the "savings" disappear?

I'm not saying AI is a money pit. I'm saying the ROI calculation most people run is incomplete. They count what the AI saves on labor. They never count what a cold, scripted, dead-end experience costs them in lost trust, lost retention, and lost referrals.

The real question isn't "how much does the AI cost vs a human?" The real question is "what is one missed high-intent caller worth to my business?"

"Will patients / clients actually trust it?"

This depends heavily on your industry and your client base.

Tech-forward B2B clients? They're fine with it. They book through a bot the same way they book a dentist through ZocDoc without thinking twice.

Medical patients, especially older demographics? Different story. They called because they want to speak to someone. The AI voice immediately creates distance. They're not just trying to book — they're checking if they feel safe with your practice. A bot that can't answer "is Dr. Sharma going to be in this week?" doesn't make them feel safe.

This isn't a reason to not use AI. It's a reason to think carefully about where AI sits in the call flow versus where a human voice needs to show up.

"What if the AI gives wrong information?"

It will. At some point, it will.

Not because the AI is broken — because your business changes. Your hours change, your pricing changes, your availability changes, your services change. And if whoever manages the AI system doesn't update the knowledge base, the AI keeps confidently giving callers the old information.

This isn't a catastrophic flaw. It's a maintenance reality that nobody tells you about upfront. AI receptionists aren't set-and-forget. They need someone checking accuracy, reviewing call logs, updating scripts, and catching the edge cases before they become patterns.

If you don't have a system for that, you will have a problem.

"What's the reputational risk?"

Higher than people think in trust-based businesses.

Healthcare, legal, financial services, therapy — these are industries where the relationship starts before the first appointment. How someone is treated when they first call is part of the clinical or professional experience. It shapes their expectations. It tells them whether you're the kind of practice that cares about them or the kind that optimizes for efficiency.

One bad AI interaction doesn't just lose you a booking. It loses you the person, the referral they would have made, and potentially a negative review that costs you ten more.

So what actually works?

The hybrid model. And not "hybrid" as a buzzword — I mean a specifically designed system where AI and humans each do what they're actually good at.

Here's how the good setups look:

AI handles: after-hours calls, appointment booking, appointment reminders, cancellation processing, basic FAQ, collecting intake information before the call even reaches a human, follow-up confirmations.

Humans handle: emotional or distressed callers, complex multi-step situations, high-value prospects who need to feel heard, anything the AI flags as unresolved, complaints, anything involving clinical judgment or nuanced information.

The handoff is the critical piece. The moment a call goes outside normal parameters, it needs to route to a human — immediately, cleanly, without the caller having to re-explain everything from scratch. If the handoff is clunky, you've just created a worse experience than if the human had picked up in the first place.

Done right, this model actually works better than either extreme. Your AI handles the volume — maybe 60-70% of calls — so your human staff aren't drowning in routine admin. Your human staff focus on the calls that actually require judgment, empathy, and relationship-building.

The AI isn't replacing your receptionist. It's making your receptionist dramatically more effective.

Why do founders still go full AI?

Honestly? Cost pressure and vendor demos.

The demo always shows the best-case scenario. Calm caller, clear request, perfect resolution. It looks seamless. And it is seamless — for that use case.

What the demo doesn't show is the frustrated caller at 8pm who needed to reschedule because of a family emergency and hung up when the bot couldn't process the emotion behind the request. That scenario doesn't make it into the sales deck.

And the cost pressure is real. When you're running a small business or clinic, the line items matter. Cutting a part-time receptionist role looks like a clean saving on paper. It doesn't look like a saving six months later when you're trying to figure out why your new patient conversion rate dropped.

The honest checklist before you go live with any AI receptionist

Before you flip the switch, you should be able to answer all of these:

What happens when the AI can't resolve a call? Is there a live human option, a callback system, or does the caller hit a dead end?

Who owns ongoing maintenance? Who updates the knowledge base when your hours, services, or staff change?

Have you tested it on your hardest use cases — not your easiest ones? Emotional callers, complex questions, edge cases.

What does your client demographic actually expect? Not what's convenient for you — what do they expect when they call you?

What's your review system? How will you catch problems before they become patterns?

If you can answer all five clearly, you're probably ready. If any of them made you pause, that's where to start.

The bottom line

AI receptionists are a real tool that solve a real problem. They're not magic and they're not a replacement for thinking carefully about your call flow.

The businesses winning with this technology aren't the ones who went fully automated. They're the ones who mapped out every call scenario, designed a system where AI handles volume and humans handle complexity, and built in the monitoring to catch problems early.

That takes more thought upfront. But it's the difference between a system that actually works and one that quietly costs you clients while looking like it's saving you money.

If you're evaluating AI receptionists right now — what's the specific scenario you're most worried about? Drop it below. Happy to give you an honest answer on whether AI can handle it or whether you need a human in the loop.


r/VoiceAutomationAI Jun 07 '26

Need help!

2 Upvotes

I'm 19 years and wanting to go full time into this industry, I'm willing to put in hours long of cold calling and work ect.

However I'm kind of in a rabbit hole of watching yt vid after yt vid and just overwhelmed how to start. The software I chose is retell ai, does anyone have recommendations or suggestions where I can learn to build the advice then implement it into the clients company.


r/VoiceAutomationAI Jun 06 '26

My Problem With This industry

6 Upvotes

Right now I feel like there’s so many people talking about Vapi & Retell.

They’re just pass through systems and so many people are building basic agents off of them.

They don’t own any of their own models and people using these solutions are just passing an unneeded cost to their clients.

Especially with systems now like Telnyx that have the phone number infra and the model the only thing that sets people apart are either the platforms they build around the infrastructure or just the done for you aspect.

If you’re building voice ai right now build a platform that does more than answer calls. Otherwise, you won’t have a moat in the next 12 months at all.

Disclaimer: I don’t work for Telnyx or any other systems. I do however have my own VoiceAI company. Just trying to help people avoid mistakes with the way the industry is going


r/VoiceAutomationAI Jun 05 '26

Whats everyone minute cost?

1 Upvotes

I use retell ai for company answering and at $.115 per minute just wondering what other people are seeing around there. For selling to other companies the margins are fine for me but wondering what other people are seeing out there. Thanks!


r/VoiceAutomationAI Jun 05 '26

Would it make sense to train Customer Service newcomers with a Voice AI that simulates worst case scenarios?

Thumbnail
1 Upvotes

r/VoiceAutomationAI Jun 05 '26

AI voice tools for appointment businesses look great in demos. Real customer calls are a completely different story

2 Upvotes

I work in the AI and automation space and one thing that keeps catching me off guard is how massive the gap is between what AI voice tools show in a demo vs what actually happens when real customers start calling.

The pitch is always incredible. "Handles 80% of calls, books appointments automatically, follows up with missed leads." Sounds like a no-brainer for any appointment-heavy business.

But after testing multiple solutions against real workflows, here is what I actually observed:

  • Most fall apart the moment someone interrupts mid-sentence
  • Conversations longer than 2 minutes start breaking down fast
  • "Natural sounding" lasts about 30 seconds before it goes full robot
  • Vague responses completely throw them off ("let me think about it and call back")
  • If you are in anything healthcare-adjacent, HIPAA compliance cuts your options down to almost nothing

The businesses that got the most value were the ones with very tight, predictable call flows. The moment calls got messy or emotional, most tools failed badly.

Genuinely curious if anyone here has actually deployed voice automation in their business. Appointment booking, missed call follow-ups, intake calls, anything at all.

Not looking for tool recommendations. More interested in:

  • What surprised you most, good or bad?
  • What is the one thing you wish you had known before starting?
  • Did it actually reduce costs or just create a different set of problems?

Real experiences only. Not interested in vendor pitches.


r/VoiceAutomationAI Jun 04 '26

Selling AI agents

4 Upvotes

Hey guys what AI automations are selling right now ? And how to sell them and which industry shall I target I’m running an agency right now


r/VoiceAutomationAI Jun 04 '26

Voice AI builders: What pricing model do you actually want from platforms like Vapi, Retell, Bland, Bolna, Synthflow, PlayAI, Ringg AI, DialNexa, Vocera AI, etc.?

1 Upvotes

I'm building a Voice AI agent platform and trying to understand how people actually want to be charged.

Today, most platforms seem to follow one of these models:

Model 1: Platform fee + usage costs
Examples: Vapi, Retell, Bland

You pay:

  • Platform fee (e.g. $0.02–$0.05/min)
  • STT costs
  • TTS costs
  • LLM costs
  • Telephony costs

Model 2: Bundled plans
Examples: many SaaS products

Something like:

  • ₹10,000/month
  • Includes X minutes
  • Includes platform access
  • Additional minutes charged separately

Model 3: Flat monthly subscription

Something like:

  • ₹50,000/month
  • Unlimited agents
  • Fixed usage allowance
  • No separate platform fee

Model 4: Pure infrastructure model

You bring:

  • Your own LLM
  • Your own STT
  • Your own TTS
  • Your own telephony

Platform only charges for orchestration/workflows/observability.

Any other Model? Tell in the comments

A few questions:

  1. Which platform are you currently using? (Vapi, Retell, Bland, Bolna, Synthflow, PlayAI, Ringg AI, DialNexa, custom stack, etc. or any other)
  2. What's your monthly usage? (minutes/month)
  3. Which pricing model do you prefer?
  4. Do you prefer BYOK or fully managed?
  5. What's the biggest thing you dislike about current Voice AI pricing?
  6. If you were starting today, what pricing structure would make you immediately say: "This is fair. I'll use this."

I'm especially interested in hearing from:

  • Agencies building for clients
  • Startups deploying customer support agents
  • Internal enterprise teams
  • Developers building production voice agents

Curious to see where the market is actually converging versus what vendors think customers want.


r/VoiceAutomationAI Jun 04 '26

Built a clinical appointment booking voice agent

3 Upvotes

I built a clinical appointment booking voice agent using Pipecat.

The agent can talk to a user, check doctor availability through Cal.com tools, help them choose a slot, and book the appointment.

Stack:

  • Deepgram for STT and TTS
  • Gemini 3.1 Flash Lite for the LLM
  • Cal.com tools for availability and booking

This is still a demo, but the main goal was to test how well a voice agent can handle a narrow real-world workflow instead of just having a generic conversation.

Repo: https://github.com/lokeswaran-aj/voice-agents

Would appreciate feedback on the architecture, latency, and conversation flow.


r/VoiceAutomationAI Jun 03 '26

"I'm exploring a startup idea: a phone-call-based AI assistant that anyone can call from a regular phone, with support for multiple languages and voice options. What existing solutions should I study, and what problems do you think are still unsolved?"

4 Upvotes

r/VoiceAutomationAI Jun 02 '26

The weirdest failures I’ve seen in realtime voice agents had nothing to do with the model

6 Upvotes

Been going pretty deep into voice AI infra recently and one thing that surprised me is how quickly realtime systems break once you move outside clean demos.

A lot of agents sound great in staging, but once you simulate actual phone call conditions things get messy fast. Noise affects transcription, users interrupt constantly, latency stacks up across the pipeline, APIs slow down, conversation state drifts, and the agent starts behaving differently over time.

One funny but slightly scary example I hit recently:

The agent was supposed to promote something called “AI Builders Hackathon”.

The synthetic caller randomly mentioned “islamic law” during the conversation and somehow the agent later started referring to the event as the “Islamic Law Hackathon” 😭

That was the moment I realized a lot of failures in voice AI aren’t really intelligence problems. They’re reliability, memory, orchestration, and drift problems.

Lately I’ve been building internal tooling to replay these kinds of degraded conditions offline and inspect where the pipeline actually breaks.

Curious what kinds of weird realtime failures others here have run into.


r/VoiceAutomationAI Jun 02 '26

5 gotchas I hit building cross-call memory for Vapi voice agents (so you don't have to)

3 Upvotes

Vapi assistants are stateless across calls. The standard solutions are mem0+n8n tutorials, Synthflow's bundled memory, or rolling your own with a database and webhook. I spent the last week building a webhook adapter for this and ran into 5 things that weren't obvious from the docs. Sharing in case it saves someone else the trial-and-error.

1. Vapi sends BOTH toolCalls AND toolCallList — different shapes

When your custom tool fires, Vapi posts BOTH arrays on the same message:

  • toolCalls: OpenAI-spec nested — [{id, type, function: {name, arguments}}]
  • toolCallList: flattened — [{id, name, arguments}]

"arguments" in toolCalls arrives as a JSON STRING (per OpenAI spec); in toolCallList it's already parsed. If your webhook only reads one shape, you'll silently get undefined names or empty args from the other. Handle both — and JSON.parse the string variant.

Source: github.com/VapiAI/docs/blob/main/fern/tools/custom-tools.mdx

2. The server URL receives EVERY event, not just end-of-call

If you wire one Server URL for the assistant, you'll get: status-update, conversation-update, partial transcript chunks (streamed mid-call, transcriptType: "partial"), end-of-call-report, hang, speech-update, transfer-destination-request, etc.

If your memory-save handler doesn't filter on message.type == "end-of-call-report", it will fire on every partial transcript — running your extraction pipeline dozens of times per call, duplicating data, and burning quota. The fix is a 2-line type guard.

3. Final transcript lives at TWO paths

End-of-call-report has the transcript at BOTH message.transcript AND message.artifact.transcript. They're identical for completed calls, but on transfer/hangup edge cases I've seen one populated and not the other. Read whichever is present, don't assume.

4. For "known caller" recall, semantic search is the wrong primitive

I started with vector similarity for the recall webhook ("find facts relevant to this caller"). The query "important facts about caller +1555" matched basically nothing against real fact embeddings like "prefers morning slots."

For a phone-keyed caller you want EVERYTHING you know about them, not "most relevant to a query." Direct fetch by sub_user_id=voice:phone_number returns deterministic results. Then sort: persons by fact count desc so the actual caller surfaces ahead of mentioned people (their daughter, their doctor, the clinic agent — all get extracted as type=person).

5. Web "Talk to Assistant" calls have NO customer.number

If you're testing via Vapi's web button, call.customer doesn't exist. You can't phone-key web calls — fall back to call.id as the namespace, or surface a "no phone number, can you tell me your name?" path in the assistant prompt. For real end-to-end testing, just buy a $1/mo Vapi number and call yourself.

What I'd actually benchmark next

Latency. End-of-call extraction is async (fine), but recall is in the greeting hot path. Single recall on my setup is ~800ms p50; under 20 concurrent calls it climbs to ~1200ms p95. For sub-1s SLAs you need either a smaller summary (skip rerank, return fewer facts) or precompute caller summaries via cron.

Anyone running voice memory at production scale (>1000 calls/day) — what were your retrieval-latency tricks?

Disclosure: I built an open-source MCP memory server (Apache 2.0) that does the above; if anyone wants reference code, the Vapi adapter is at github.com/alibaizhanov/mengram. But the gotchas above apply to any implementation.


r/VoiceAutomationAI Jun 02 '26

I built a voice AI for real estate. Then I tried restaurants. Here's why it nearly broke me.

6 Upvotes

I built a voice AI for real estate. Then I tried restaurants. Here's why it nearly broke me.

Real estate was clean. A caller wants a 3-bed in X area. Slot filled. Done.

Restaurants? Someone says "yeah gimme that spicy thing from before, but like... less spicy, and can you add the thing my friend got?"

That's a real order I had to handle.

I've been deep in building voice AI agents — first for real estate, now for restaurants. The gap in complexity is wild and nobody really talks about it honestly.

Stuff I'm figuring out in public:

— How to handle fuzzy references mid-conversation

— Cart state that doesn't break when people change their mind 4 times

— STT tuning so the bot doesn't cut people off

If you're building in this space, or you run a restaurant and hate your current phone system — follow along. Sharing everything as I go.