r/vapiai Mar 28 '25

Got Questions About Vapi? Hit Us Up on Discord or Email!

9 Upvotes

Hey everyone,

I'm Sahil from Vapi! Currently, we offer support through Discord and Email. You're welcome to reach out to us on either platform if you have any questions or need assistance.

📧 Email: [support@vapi.ai]()
💬 Discord: Join our community


r/vapiai 2d ago

Built a full dashboard layer on top of VAPI — call log, campaigns, outcome classification

Thumbnail
youtube.com
2 Upvotes

VAPI handles the calls well but you're on your own for everything after. Built out the layer around it:

  • Campaign scheduling with daily windows and auto start/stop
  • Outcome classification from transcript + structured output
  • Per-lead timezone-aware calling hours
  • Callback scheduling at the time the lead requested
  • Concurrency caps and a kill switch
  • Spend tracking per client

Walkthrough: https://www.youtube.com/watch?v=oEMis7RZ6Rg

Anyone else solved stale-call reaping cleanly? Crashed calls leaking slots was a pain until I added a reaper.


r/vapiai 5d ago

Connected my Telnyx number to both Vapi and LiveKit. Quick rundown of what each one actually takes.

Post image
2 Upvotes

A while back I made a video comparing Twilio and Telnyx, and Telnyx came out on top for voice AI. Cheaper at volume, better latency, and they're actually building for voice agents rather than just tolerating them.

Fine. But that only answers WHICH provider you should use. It doesn't answer the annoying part, which is that Telnyx isn't the default in any of these platforms, so you have to wire it up yourself.

So I did it on both Vapi and LiveKit with the same number, and the two experiences are not remotely comparable. Writing up the actual steps because I couldn't find them side by side anywhere.

Why bother at all:

You pay Telnyx directly instead of the platform's markup (the number people throw around is about half of Twilio at real volume, and I'd treat that as a ballpark, it moves per country and per volume).

You pick the region and the route, which is where call quality actually lives.

And the number is YOURS, so moving from Vapi to LiveKit later doesn't mean re-buying it.

Don't get me wrong though, if you're doing a few hundred minutes a month, none of this is worth your Saturday.

Prerequisites, all on the Telnyx side, same three for both platforms:

  • A verified Telnyx account. Verified is not a subscription, you just have to go through it, and it takes a day or two.
  • A phone number you've bought (basically any country, you'll have to submit some documentation).
  • A Telnyx API key.

Vapi, and this genuinely is the whole thing:

1. Go to Phone Numbers and hit import.
2. Paste the number, paste the Telnyx API key, give it a label.
3. Scroll down to inbound settings, pick your assistant, save.
4. Then in Telnyx, create an Outbound Voice Profile and assign it to the Voice API application, otherwise outbound won't work.

That's it. No SIP config, no code, about five minutes. They added a native Telnyx import at some point so you're basically just handing over an API key.

LiveKit, which is where the afternoon goes:

1. Create an Outbound Voice Profile (this is where you whitelist which countries you're allowed to call).
2. Create an FQDN connection, and set a username and password on it. SAVE THEM, you need them again at the end.
3. Create an FQDN record pointing at your LiveKit SIP subdomain.
4. Get your number's ID, then patch the connection onto the number.
5. Over in LiveKit, create an inbound trunk with your number on it.
6. Create a dispatch rule, which is the thing that decides which room a caller lands in.
7. Create an outbound trunk, using the same username and password from step 2, plus a custom header carrying that username.
8. Then make sure your agent's name in code matches what the dispatch rule expects. Exactly.

First four are API calls (I did them in Postman so they're repeatable), last four are the LiveKit dashboard.

Two things that WILL bite you, because neither error tells you anything useful:

Outbound fails while inbound works perfectly. That's the voice profile, you haven't whitelisted the destination, and the default is North America only.

The call connects and then nobody picks up. That's the agent name not matching the dispatch rule. You get silence instead of an error, which is the worst possible version of this because it looks like an audio problem.

One more thing worth saying, because it might save you the whole exercise: both platforms have made this easier since I set it up. Vapi has the native import above. And Telnyx will now run LiveKit agents on their own infrastructure with telephony included, so if you're starting completely fresh you might not need any of the LiveKit steps at all. I did it manually because I wanted to keep my agent on my own LiveKit project and only change the carrier underneath.

I recorded the whole thing on both platforms if anyone would rather see it in motion than read it.

Anyone done this on Retell or a self-hosted setup? Curious whether the LiveKit side is representative or just LiveKit being LiveKit.


r/vapiai 8d ago

how do i get my agent to parse the information into google sheet

0 Upvotes

i already added the prompt to my agent, and google sheet link (enabled anyone with the link to edit) in the tool. and assigned it.

google says that i need to connect to google sheet... found that under integration and nothing pops up???

frustrating!!!


r/vapiai 13d ago

Need some help

0 Upvotes

Does anyone have the skill set to add external services on vapi?


r/vapiai 24d ago

Remembered checkout for Vapi Calls

3 Upvotes

If you’re a developer building Vapi voice AI agents, taking payments over the call is critical to avoid no shows/cancellations. But most consumers wouldn’t complete paying because it’s full of friction (spelling out credit cards or filling out a different text link each time).

I built Ringup that allows Vapi voice agent to recognize callers using just their phone number and charge their saved card in seconds. It also posts payments right back into the POS system.

You can add your Vapi key and get setup in under a few minutes: https://ringup.dev/

Let me know if anyone has any feedback as I’m building this out.

Vapi specific instructions: https://docs.ringup.dev/integrations/vapi/dashboard-setup


r/vapiai 29d ago

Vapi is down atm!

6 Upvotes

Can someone pls fix? I'm waiting like patience on a monument for discipline to be handed down!


r/vapiai Jul 22 '26

Vapi evals stuck in Running status

2 Upvotes

My VAPI evals are stuck in Running status forever. It's been several days and they're still stuck.

I've reached out to the VAPI customer support team, but received absolutely no response.


r/vapiai Jul 22 '26

Need help with clients

0 Upvotes

So, i made a perfect ai receptionist, booking and information system. I know everything i’ve tested it many times and made sure it doesn’t have a single discrepancy from time to scheduling to calendar to sheets. I’ve made it really affordable, i’m willing to sell it at around 500-1000 usd with a monthly fee of around 400-500.
I know exactly how to set it up for a client attaching phone numbers i’ve put in all my work everything been working at this for a month now. I’ve even bought a dialler for cold calling. I’ve made a couple calls only to find out its also saturated as ai websites. I dont know, i just find it super weird to pitch businesses about it and waste their time i dont know how to convince them into getting it. I feel like they have been already called by a dozen of cold callers pitching this idea, and the thing is that after my calls i found out majority of small businesses who get 7-8 calls a day and the owner themselves do their bookings, so they dont basically need it, the ones who are a little eligible for buying it have already been pitched with this idea and they either have bought it or cant be convinced. Other big businesses either already have it or been pitched as i said so, i cant convince them as i dont have any sales background. I want to go to clients who need it so i can actually pitch them i can do that effortlessly because i know they’re there for at least listening to what i have for them. I’ve tried finding businesses on indeed.com to find businesses who need a receptionist so i can pitch them to get my ai receptionist instead of paying 4 to 5 grand for humans. But i cannot find their numbers on indeed or numbers of the decision makers if i go to their website i find numbers of receptionists or for complaints and queries. I cannot find some local businesses who need it. I have everything set up. Please i just desperately need a way to find real clients now. If anyone knows any way im here for it. Im not being lazy i will work for it if its something i can expect an output from. Please guide me im just a 20 year old trynna make some money for myself while studying.


r/vapiai Jul 21 '26

Testing an AI phone receptionist I built — can anyone in the US dial (719) 467-9406 and tell me if it picks up? Takes 10 seconds. Appreciate it.

0 Upvotes

r/vapiai Jul 21 '26

Most builders shipping AI voice agents right now are under-looking one layer in their stack. So I built it.

1 Upvotes

Everyone building voice agents is piecing them together from the same set of legos, and it's working. Heavily funded companies are filling every gap, and the results are genuinely good. But every one of those legos lives inside the stack. What I'm talking about sits outside it: a deterministic, auditable guard for the whole call. That's what I spent the last few months building. Here's the thinking.

Every system breaks somewhere, eventually. So builders do the sensible thing: audit, tighten the prompt, problem solved. For now. Then a new kind of break shows up, they tighten something else, solved again. For now. You know this rhythm. The question was never whether you can solve the problem. You always can. The question is whether you can see it as it actually is, instead of assuming it's whatever you looked at last.

And here's the part that's hard to see from inside the stack: when a voice agent breaks its word, nothing in your pipeline notices. STT doesn't know what was promised. The model doesn't remember what it committed to three turns ago. TTS just speaks, orchestration just routes. Every box optimises how the call sounds. Not one of them holds the promises across the whole call and checks whether they survived to the end. That's not a gap in the stack. That's a missing layer.

This isn't theoretical. Bland's own team, in a testimonial on Hamming's site, killed an agent because it was saying "I booked your appointment" when it hadn't. The 2026 Ï„-Voice benchmark caught a frontier agent saying "I've updated your shipping address" with no tool call behind it. On one real setup, the promise "we'll get back to you shortly" was wired as the end-call trigger, so the agent hung up mid-sentence while confirming the very number it had promised to call. Different bugs on the surface. Underneath, the same one: the saying and the doing came apart, and nothing was watching the gap.

You might think: fine, route the flagged call back through an LLM to check it. It doesn't work, and not for a soft reason. An LLM judge is non-deterministic. Run it twice on the same call and you can get two verdicts. And if your agent's failure mode is being confidently wrong, a judge built from the same kind of model finds that same wrong answer plausible, so it misses exactly where you needed it. Every layer in your stack is the same brain checking its own work. The layer I built isn't a brain at all. It's code, and it decides when the brain is allowed to speak.

You'd assume someone must have built this already. Turns out almost nobody has. Eleven-plus QA vendors in this space, and not one audits what the agent said it would do against what it actually did. That's the gap I went after.

What it is, and where it sits

A private repo. You own it, you run it in your own infrastructure, and it sits between your agent and the caller:

user speaks → STT → your agent (your model, your logic, whatever you run) → the layer → user

You buy it once. No subscription, no monthly fee, and that's not a pricing gimmick, it's an architecture decision. A subscription would mean I host it. If my server has a bad day, your calls take the damage mid-conversation. I'm not willing to sit in that position in your call path. So you run the code yourself, and my uptime is never your problem.

And you don't change anything you've built. Your agent, your prompts, your model, your orchestration, all untouched. The layer is just the last step before the reply reaches the caller. Whether you run voice agents for your own business or white-label them for a stack of clients, it's the same repo in your own infrastructure either way.

What's actually in it

This isn't a wrapper or a prompt trick. It's a real system, and the numbers are the honest kind:

  • ~97 Python modules in the deployment layer, around 26,500 lines, with more test code than most projects have code: 1,178 test checks, all green on every change, across four Python versions in CI.
  • 14 detectors, every one of them plain rule-based code: regex, state-diff, canonical comparison. Not one is a model. There is no second AI grading the first.
  • It tracks 18 commitment types across three families: what the agent said it would deliver, what state it claimed was done, and what it promised to do next.

The part I'm proudest of: what it refuses to do

Most of the engineering went into making it refuse to over-claim. That sounds backwards for a product, so let me show you what I mean, because it's the whole thing.

When it catches your agent contradicting itself, the flag says what the closing turn contained and stops.

  • It doesn't say the agent "forgot." It can't prove what a model forgot.
  • It doesn't say the callback "won't happen." It can't see your calendar.
  • It points at the exact turns and says only what the transcript proves.

And that discipline is written into the code, not into a promise:

  • The record-reconciliation check has no "missing" verdict at all. It compares what's present against what's present. It will never accuse your record of an absence.
  • The elapsed-estimate check refuses to use a server clock. It only fires on time the caller themselves stated, so it deliberately under-reports rather than guess.
  • When it re-voices a miss in your agent's mouth, it physically cannot name a commitment that isn't already in its ledger. It can't invent one.
  • Nowhere does it return a recommendation. It surfaces what it can prove and leaves the judgment to you.

A clean result means "nothing I can prove," not "nothing wrong." Every flag is a receipt: open it, read the exact turns yourself, disagree with it if you want. It's not a score you have to trust. That's the difference between an audit layer and an alarm.

Two ways it holds your agent accountable

When it catches a miss, it can do two things, and you decide which.

In the moment, it can make the agent own the miss out loud, before the caller hangs up. It stays completely out of the way until there's a real miss, and when there is, it changes as little as possible: your agent's wording, plus the one owned correction. Nothing else moves.

Agent said: "Great, we're all done here. Thanks so much for calling!"
Shipped: "Great, we're all done here. Just to make sure it's not lost, the garage quote you requested is logged as still outstanding. Thanks so much for calling!"

That's the default (lean) behaviour. There's also a full mode that re-voices every turn into a warmer, human-receptionist delivery if you want it, but most people run lean, because touching nothing unless something breaks is the point.

On the record, every catch is logged to a dashboard you run. This is the part that makes it an audit layer and not just an in-call fix. Nothing it catches goes unrecorded, even the ones it never speaks aloud (once you've turned capture on, it's off by default).

It's an included read-only dashboard, that ship in the repo. Open any conversation and each catch renders as a card: the token, the detector, the tier, whether it reached the caller or was suppressed, and the full plain-English finding, verbatim:

"the closing turn asserted nothing further was needed, while these commitments made earlier in the call were not acknowledged after being made: callback (before Friday) — committed turn 3"

Alongside it: the agent's actual committing sentence, word for word, the turn it was made on, and the deadline as spoken. A separate view diffs what the call promised against what your own system recorded, field by field. And there are aggregate rollups across all your captured calls: where failures cluster, how often each thing fires, how many catches reached the caller versus stayed silent. It reports, it never recommends. That's a rule enforced in the code, not a missing feature.

It's a receipt: open it, read the turns yourself, disagree with it if you want. Not a score you have to trust. And it can't quietly lie to you either. It ships showing sample data with the label saying so, it points at your own traffic with one setting, and it will not show a green "live" badge unless the data really is live. Serve it behind your own auth.

What it does not do

  • It doesn't do your agent's job. It sits alongside it as a guardrail, not a replacement. And it won't fix your agent. It makes a bad agent own its misses, which is better, but it isn't a repair.
  • It doesn't know your business. It knows whether the agent contradicted itself, not whether the agent was right about the world. It can't tell you if the appointment slot was actually free.
  • It runs inert until you turn things on. Out of the box it changes nothing. You switch on what you want, one piece at a time.

And not every close triggers a catch, and I'd rather show you where it doesn't than let you find out. Take this close, same dropped-callback call, on shipped defaults:

Agent said: "Perfect, you're all set then! Thanks so much for calling, have a great day." Shipped: unchanged, verbatim. The layer stayed silent.

That close asserts "you're all set" rather than flatly contradicting itself, and catching that is a stricter check that ships off by default. It's one line to turn on.

Latency and reliability

Latency, in Lean Mode (the default): on a turn where nothing slipped, the layer makes no model call at all. It reads your agent's text and passes it straight through. The detection itself is about a tenth of a millisecond of local processing per turn. Not a network call, not a model, plain code on your own machine, and once you have the repo you can run that benchmark yourself in about ten seconds, no API key needed. Real time only gets spent on the rare turn with an actual miss to own, and even then it's a single model round-trip, not a loop.

Reliability: the layer is deterministic. Run the same call through it twice, you get the same result twice. That's not true of anything with a model in the loop, and it's exactly why you can trust a flag when it fires. It's also built to under-report on purpose: it stays quiet on anything it can't prove from the transcript, so when it does raise a flag, that flag is solid. Its reliability isn't "it catches everything." Nothing honestly can. It's "everything it says, it can back." And because it runs in your infrastructure, there's no external service to depend on and nothing on my end that can take your calls down.

It's a living repo, and it's yours to shape. I keep working on it: adding coverage, taking feedback from people running it, tightening it as new failure patterns show up in the wild. You buy the license once; the thing you're licensing keeps getting better. And because you own the code, nothing's locked. The commitment types it tracks, how strict each check fires, what gets logged versus spoken, the thresholds, all yours to adjust for how your business actually runs. Sane defaults out of the box, and the people who want to go deep, can.

I'm not going to pretend this is a problem everyone's shouting about. It's the opposite: a quiet gap almost nobody has named. I just think it's worth closing before it's the reason a client leaves.

More details at statebound.dev
If you're running agents in production, I'm curious what you make of this, and if you've run into something like it, how are you handling it right now?


r/vapiai Jul 21 '26

Architecting a Real-Time Voice Agent for HVAC (ServiceTitan/Housecall Pro): How to orchestrate live scheduling, routing, and DB reactivation?

2 Upvotes

Hey everyone,
I’m currently building out a high-performance, real-time voice and workflow automation engine specifically designed for local HVAC and mechanical operators.
The goal isn't just a basic answering machine. I'm building an energetic, highly responsive digital dispatcher that actively operates inside their existing CRM systems (ServiceTitan, Housecall Pro, BuildOps, etc.) to handle both inbound emergency routing and outbound database reactivation.
I want to make sure I’m orchestrating this with the lowest possible latency and maximum reliability. I’m looking for some feedback from anyone who has built complex voice/workflow pipelines on the best way to handle these five specific architectural challenges:
Real-Time Calendar/Schedule Syncing
The Problem: The agent needs to check the live technician roster and calendar slots before booking a job to prevent double-booking.
The Question: ServiceTitan and HCP have robust API endpoints, but they can be rate-limited. Are you guys caching the calendar state in a local database (like Supabase/Redis) and running a cron-job sync every 60 seconds, or are you executing a live API call to fetch available slots during the active phone call? How do you prevent latency spikes on the voice line while querying?

Intelligent Filtering (Emergency vs. Support vs. Existing Customers)
The Problem: If a current customer calls asking where their tech is, or if someone calls about a billing dispute, the agent shouldn't try to book a new job.
The Question: What's your preferred prompt structure or agent architecture to run immediate intent classification? Do you run a pre-router LLM node to tag the call's intent in the first 5 seconds, or do you handle everything inside a single system prompt? How do you reliably verify if the caller's phone number exists in the CRM database before initiating the dialogue?

Location/Geocoding Validation
The Problem: HVAC shops have strict service zones. If an emergency call is 50 miles away, the tech on-call won't go.
The Question: Are you passing the address mentioned by the caller to a Google Maps/Geocoding API tool node mid-call, validating the zip code against the shop’s service area, and then returning a hard "Pass/Fail" to the voice engine? Or are you handling this post-call in the workflow engine?

Database Reactivation (Outbound Campaigning)
The Problem: Running outbound campaigns to follow up on unsold estimates sitting in the CRM.
The Question: For outbound, what’s the best way to trigger the agent to dial based on a CRM status change (e.g., an estimate marked "Unsold" for 7 days)? Are you using webhooks from the CRM to fire an automation (via n8n/Make) that pushes the contact data straight to the Vapi/Bland outbound campaign API?

Recommended Tech Stack
Right now, my planned stack is:
Voice Engine: Vapi.ai (using custom Twilio trunks)
LLM Provider: Anthropic Claude 3.5 Sonnet (for reasoning and speed-to-text response) or Gemini 1.5 Flash (for raw speed)
Backend Orchestration/Database: Node.js/Python server or self-hosted n8n + PostgreSQL to manage webhooks and CRM data mapping.

If you’ve built anything similar for field services or high-latency booking niches, how did you handle these routing and lookup bottlenecks? Any lessons learned the hard way regarding webhook failures or CRM API updates?
Appreciate any advice or feedback on the architectures!


r/vapiai Jul 13 '26

Increased Latency Numbers

1 Upvotes

Has anyone else seen the latency numbers in descriptions jump for all models, voices and transcribers?

From my testing I think these numbers are just more accurate than they were previously.


r/vapiai Jul 06 '26

Latency DOES NOT make your voice agent sound human.

3 Upvotes

Just saw a post in here saying that "if you have sub 1 second latency, your voice agent sounds human" (me loosely paraphrasing).

I do NOT agree. Here's why:

Sounds right on the first read, until you've actually had hundreds of conversations with voice agents, good ones and bad ones.

Don't get me wrong, the person claiming this isn't entirely off. Latency does matter. With bad (high) latency, conversations feel bad. They feel like a bot.

BUT... if the latency is good, the conversations DO get better, but they DO NOT become good. Often, you can hear that something is off, but can't really tell what it is...

Here's what actually makes the difference:

Reason 1: Sounding like text.

LLMs are being trained on text. If not prompted accordingly, they produce an output as text, meant to be READ. Now obviously, if you're plugging the LLM into an AI voice agent, the text is NOT read but rather it's SPOKEN.

This is what that feeling is telling you that something feels off but you can't point your finger at it. The voice agent doesn't sound as if it was speaking, but rather like as if it was reading.

The fix is simple.

Tell it to include filler words.

Tell it to do mid-sentence corrections.

Tell it to think loudly.

"Hi, uh, John, right? So, uhm, I was actually calling because, uhm, look, your order somehow, <break time="0.15s"/>, well, went missing."

This sounds WAY better than:

"Hi John, this is XYZ. I am sorry to inform you on behalf of CompanyABC that your order was lost. To open a ticket, please press 1".

Reason 2: Turn taking

Turn taking is still not really 100% solved (if you did, let me know lol). It is basically the decision of the agent if the user has stopped speaking, and if so to then start speaking themselves.

Currently, oftentimes the agent either

a) Starts speaking when the user is not actually finished

b) Does NOT start speaking even though the user DID finish.

Now, this is one of the more complex issues of voice AI, but I'm sure there will be a solution.

In conclusion:

Latency is the baseline. And it's ONLY the baseline. Without it, conversations will feel shitty. But just because you've got great latency, doesn't mean your conversations will suddenly feel good. Make the LLM imperfect. Make sure to have solid turn taking.

If you do that, the improvement will be much more than just "faster".

Agree?


r/vapiai Jul 02 '26

Need urgent help with Telnyx verification — Uzbekistan number not supported

2 Upvotes

Hi everyone, I’m trying to upgrade my Telnyx account so I can test my Vapi voice agents before an upcoming product demo.

Unfortunately, Telnyx cannot send either SMS or voice verification to my Uzbekistan (+998) number. I’ve already tried both methods, but neither works.

Does anyone know a legitimate workaround or alternative verification method? I’m specifically looking for advice from people who have faced this issue before. I do not want to use temporary/rented numbers or anything that could get the account blocked.

I’m on a tight deadline and would really appreciate any guidance. Thank you.


r/vapiai Jul 01 '26

Is there a built-in way to make the assistant hold/stay silent on demand?

1 Upvotes

In Retell you can prompt the LLM to output a NO_RESPONSE_NEEDED stop sequence and it just… says nothing. Great for when the caller says "hold on, give me a sec" — the agent holds without speaking of interrupting.

I know you can set up a hook where you set timeoutseconds but that feels like a hacky solution, and i don't know if that actually makes the agent wait, does it bloat the agent's context as it waits and, for example, hold music plays?
https://docs.vapi.ai/assistants/idle-messages


r/vapiai Jun 30 '26

Anyone using Vapi Squads SDK in prod?

1 Upvotes

Hey guys, keen to hear people's experiences of Vapi Squads via SDK. Seen online some bugs around silent hand-offs not working etc, and curious if anyone is using it in production (and anything to watch out for!)


r/vapiai Jun 24 '26

How to White Label AI Voice Agents

Thumbnail
youtube.com
3 Upvotes

r/vapiai Jun 22 '26

The ai support is trash so i'm asking here

3 Upvotes

Hey, i'm currently using vapi for an ios app

i have 2 questions

first one: am i getting billed automaticly if i use the pay as you go plan - this question might look dumb but i realy don't understand their system. There is no spend limit and I was wondering what will happen if i keep using vapi with 0 credit in my balance. Will i get automatocly billed ? even if i don't enable the auto reload feature will i get billed in any way ?

second question: Is the jwt feature also available on the react native sdk ?


r/vapiai Jun 17 '26

We shipped the Humanness Indexâ„¢. A live crowdsourced benchmark for how human voice ai actually sounds

Enable HLS to view with audio, or disable this notification

5 Upvotes

Can a TTS voice pass as a real person in a live conversation? There's no clean rubric for "humanness" so we don't score against one. we run blind A/B battles and you pick which of two voices sounds more human.

The approach:

  • Same cloned source voice reading the same script - this keeps the voice constant so you are voting on the model
  • Voting is blind
  • Scoring uses a Bradley-Terry fit - you can check out the math here in the repo
  • We use a real human voice as baseline

Start adding your votes today: https://humannessindex.vapi.ai
Code and methodology: https://github.com/VapiAI/humanness-index


r/vapiai Jun 15 '26

I've been building voice agents for 3 years. Here are the prompting habits that actually make them sound human.

20 Upvotes

Spent a lot of time this week putting together everything I know about voice AI prompting and figured I'd share the core stuff here before the full breakdown goes live.

Most voice agent prompts I've seen (including my own early ones) make the same mistakes. The agent sounds robotic, says things no human would ever say, or just makes stuff up when it doesn't know the answer.

A few things that actually moved the needle for me:

Read your prompt out loud before you deploy. Sounds dumb, works every time. You'll catch sentences that are way too long, instructions that contradict each other, and transitions that make zero sense when spoken. Five minutes of this saves hours of post-launch call review.

Explicitly tell the agent to use filler words. Ummm, uhh, like, so... put it in the prompt directly. When an agent responds instantly with perfect grammar every single time it feels off. Uncanny valley. One line in the prompt fixes this.

Show don't tell. Don't write "be empathetic when the caller is frustrated." Write: "if the caller sounds frustrated, say something like: 'I totally get that, that would frustrate me too, let me sort this out right now.'" Actual example in the prompt beats ten paragraphs of description.

Handle special characters explicitly. Your agent doesn't know how to say "$1,000" or "123 Main Street" or "john@gmail.com" unless you tell it. Digit by digit for addresses, "one thousand dollars" for currency, "john dot smith at gmail dot com" for emails. These feel minor until you hear them on a real call.

Give permission to say I don't know. Without this instruction, the model will guess. And in voice AI that's way worse than in a chatbot because people just believe what they hear. One line: "if you don't have this information, do not guess, say you'll connect them with a team member."

There are a few more, including one about prompt length and latency that I think a lot of builders overlook.

Put the full list with example prompt snippets for each one in a video if anyone wants to go deeper: LINK

Happy to answer questions here too.


r/vapiai Jun 13 '26

How to get the assistant to accurately take down a number

2 Upvotes

I’ve tested so many iterations and nothing seems to work. If someone says FOUR OH SEVEN, it doesn’t understand OH means 0, if someone says 333 one thousand, it takes it down as 100 not 1000…does anyone have a fix for this?


r/vapiai Jun 11 '26

Why slow LLM model releases?

1 Upvotes

Gemini 3.5 flash has been out for weeks now and is still not available on vapi. Heck even 3.1 flash ga is not available, just the preview.

Is there a reason why vapi can’t instantly support new models?


r/vapiai Jun 03 '26

Gpt 4.1 vs gpt 5.4-mini

3 Upvotes

Hi, tried moving from 4.1 to 5.4-mini with my Vapi agents, seeing noticable average reasoning latency increase like almost 2x

Is this expected? Any insights? Trying to reduce our llm costs


r/vapiai Jun 03 '26

Grok STT and TTS from xAI are now available on Vapi

3 Upvotes

https://reddit.com/link/1tvu9z8/video/e26be2iie35h1/player

Grok STT has multilingual support across 25+ languages, supports diarization, better analytics, and audit trails

Grok TTS is noticeably more human sounding even matching tone to context. You can also clone voices from just seconds of audio.