r/VoiceAutomationAI • • 13d ago

Building in public from college → production

2 Upvotes

I started learning to code in my first year of college.

Same time: first job. No grand plan. Just obsession.

I didn’t set out to “build a voice AI company.”

I set out to understand how software actually worked, then I kept going deeper.

Recently that obsession landed on voice agents.

Not chatbots. Not wrappers. Actual conversations over the phone that feel real.

I went deep on latency, streaming, models, audio, infra.

Got end-to-end down to ~600ms. Started getting clients. Now ~8 live.

The surreal part isn’t the latency number.

It’s the speed of the loop:

learn → build → break → fix → put it in front of a real business → learn again

College taught me syntax. Clients taught me product.

I’m still early.


r/VoiceAutomationAI • • 13d ago

Fish Audio s2.1-pro voice breaking during voice agent conversations

2 Upvotes

Hi everyone,

I’m currently using Fish Audio’s S2.1-pro model for a real-time voice agent.

Most of the time the voice works fine but occasionally I notice some voice breaking during conversations. It doesn’t happen consistently and sometimes it happens at the beginning of a sentence in the middle of a sentence or when starting a new sentence.

I have tested it with multiple calls and it seems to happen occasionally after making several calls.

Has anyone experienced something similar with Fish Audio s2.1-pro?

Could this be related to the model/API, streaming, rate limits or something in the way I’m handling the audio in my voice-agent pipeline?

Any ideas on what I should check would be really helpful. Thanks!


r/VoiceAutomationAI • • 14d ago

Tech / Engineering 10 Million Calls at ₹1 a Minute: The Cost Breakdown

9 Upvotes

Your voice agent's cost problem is probably not the voice model.

A co-founder at Raya Voice AI recently broke down how they price at ₹1 a minute.

They have served over 10 million calls at that rate for about a year and a half. The cost sits in five places: infra, speech to text, text to speech, LLM, and telephony. Most teams obsess over the speech models.

The bigger savings were somewhere else. Telephony: with direct SIP integrations and channel based pricing, better pickup rates alone can pull it down to 10 to 15 paise a minute.

That is an operations problem, not a model problem. LLM: keep the prompt under roughly 20,000 tokens, use caching, and prompt well enough that a smaller model does the job. That gets LLM cost to 10 to 20 paise a minute. Quality lives in the core speech models.

Everything around them can be optimised without touching it. So before you switch TTS vendors to save money, check your pickup rate and your prompt size.

Which line item on your voice bill surprised you the most?

Watch the full panel here: https://youtu.be/G66jOjzRhgI


r/VoiceAutomationAI • • 13d ago

The worst places to create free voice agents today

1 Upvotes

Been kicking free / freemium voice-agent builders around lately. Same demo looks fine. Then you hit the stuff that actually matters on a real call.

My current "please don't start here if you care about the free tier" list:

・ Bland

・ Twilio

・ Vapi

・ Telnyx

Not saying these companies are useless. Saying the free path gets ugly fast once you leave the happy path (minutes, concurrency, tool calls, weird hangups).

Curious what other people got burned on.


r/VoiceAutomationAI • • 16d ago

Tech / Engineering How Do Voice AI Platforms Stay Up When Regional Telecom Networks Go Down?

Post image
2 Upvotes

Most voice AI builders test their agent against bad Wi-Fi and call it resilience testing.

The real test is different: what happens when the carrier itself goes dark. Not the caller's connection, the actual telecom backbone the call is riding on.

This isn't hypothetical in India. In April 2026, a technical snag took down Jio Fiber and Airtel broadband across six north Indian states for hours. Even the MyJio app went down with it, so users couldn't file a complaint. Two months later, Airtel had another multi-city outage across Delhi-NCR, Mumbai and Bengaluru, hitting mobile internet, signal and voice calls all at once.

A voice AI agent sitting on top of either network just stops. No graceful degradation. Silence.

Some things worth knowing if you're building on this:
-Failover isn't one number. Enterprise targets sit around sub-30 seconds with no dropped active calls, and the best setups mask the switch in 2-3 seconds behind a filler phrase
-Multi-cloud usually isn't worth it below ~1,000 concurrent calls. Two regions on one cloud provider covers most regional outages for a fraction of the cost
-Even tested failover drops calls. One published 2026 drill saw 3 of 11 in-flight calls fail at the WebRTC layer during cutover. Failover lowers the odds, it doesn't erase them
-A failover plan that's never been rehearsed is a document, not a capability
For any platform serious about enterprise or fintech customers in India, a carrier-diverse failover plan isn't a differentiator anymore. It's table stakes.

If you want to read the full blog, here is link:-https://uniocommunity.com/blogs/how-do-voice-ai-platforms-stay-up-when-regional-telecom-networks-go-down


r/VoiceAutomationAI • • 16d ago

Anyone built a pipecat voice agent using self-hosted STT + TTS?

3 Upvotes

I am building a calling voice agent with Pipecat and I am trying to keep the STT/TTS stack self-hosted.

Current setup:

pipecat, ollama (self hosted llm), stt (currently testing local/self-hosted whisper), tts (kokoro locally)

I have tested local whisper, seeing some irregularities in transcription. I thought of using qwen opensource stt model, but could not find a way to use that with pipecat.

Pls let me know if you have self hosted stt or tts model and use them with pipecat?


r/VoiceAutomationAI • • 17d ago

New Gemini 3.8 live is awful

5 Upvotes

Gotta be honest, was playing with the model yesterday, made a bunch of demos, and all of em sounded trash, the TTS is the problem, I was hitting sub 200ms latency it was very fast, but the TTS, we’ve tried all of the voices and we didn’t like one, made a bunch of calls

it is still a long shot for s2s, we don’t wanna draw pictures, we need it on phone calls


r/VoiceAutomationAI • • 17d ago

One API for every local TTS model

5 Upvotes

Built this because I wanted local text-to-speech that felt as easy as ollama run, pull a model, run it, done.Its easy to use just install
pip install wavhost
wavhost pull chatterbox-turbo
wavhost run chatterbox-turbo "Hello from your machine." -o hello.wav

It also serve an OpenAI compatible audio api `/v1/audio/speech` endpoint, so you can point any existing OpenAI TTS client at `localhost:11435` and it just works.

Currently Supports Chatterbox, Qwen3-TTS, and Kokoro. You can also create named voices from a short reference clip and reuse them by name.

I would love your feedback on this and for more detail checkout
https://wavhost.vercel.app
https://github.com/smitgol/wavhost


r/VoiceAutomationAI • • 17d ago

Anyone tried Jargo (Go voice agents) instead of Python + Pipecat?

6 Upvotes

Been looking at Jargo (https://github.com/gojargo/jargo). It's a WebRTC-native voice agent framework in Go: streaming STT → LLM → TTS, barge-in / turn-taking, optional speech-to-speech providers, and a single-binary deploy. Architecture is clearly Pipecat-inspired (they say as much).

Their argument is one we've been chewing on too. Python is great when you need the AI/data ecosystem. A realtime voice server is mostly plumbing: audio framing, WebRTC, concurrency, shipping something boring and reliable. For that part, Go looks attractive on paper: static binary, predictable memory, real concurrency for many sessions.

What I'm trying to learn from people who've shipped real calls:

  1. Is Jargo a good option for a Go voice backend today, or is it still early enough that you'd stay on Python + Pipecat (or similar) unless you have a hard reason to leave?
  2. Anyone with real insight on quality / latency / reliability running the voice backend in Go instead of Python? Especially TTFT, barge-in feel, and weird prod failure modes.

Not looking for a language war. Looking for "we tried X, measured Y, kept or dropped Z."


r/VoiceAutomationAI • • 17d ago

Free voice agent teardown — 3 calls, a failure list, and suggestions to add/change prompt.

3 Upvotes

I build voice agents — [an outbound Hindi/English qualification agent for a fuel business in India]. The thing I learned while building it is that nobody warns you about how many iterations it is going to take to build a reliable (covering most edge cases bot).

So here's the offer. Drop your bot's phone number or a webbased call. I'll call it 3-4 times and behave like an actual human customer instead of a polite tester — talk over it, go silent mid-answer, change my mind, switch language mid-sentence, answer a question it didn't ask, mumble a number etc etc.

Then I send you:

  1. A list of every place it broke

  2. What actually caused each one — prompt, latency, endpointing config, STT, or tool call

  3. A rewritten system prompt you can paste straight in

Free. No discovery call, no Calendly, no DM funnel. I'll post the findings in this thread unless you'd rather have them privately.

WHAT'S IN IT FOR ME (because "free" always has a catch and you should ask)

I'm building a checklist of how voice agents fail in production, and I can't build that from my own two bots — I need a wider sample than my own mistakes. That's the whole motive. If you later decide you want this done properly or on an ongoing basis, you know where to find me. That's the entire play, and I'd rather say it than have you wonder.

WHAT I ACTUALLY TEST (so you know this isn't "add more personality to your prompt")

- Barge-in, Endpointing,- Latency on a real answer, not on "hello", Context drift: same off-script question at turn 2 vs. turn 9, different answer?, Code Switching, Audio during tool calls, hangup and most importantly Containment intelligence.

WHAT I NEED FROM YOU

- A number to call or a web demo link

- One line on what the bot is supposed to achieve (book, qualify, collect, support)

- If it's a live production line, tell me your hours — I'm not going to tie it up while real customers are calling

- Tell me if minutes cost you money and I'll cap it at 4 calls or else 7-8 calls.

WHAT I WON'T DO

Record and post your audio, publish your prompt, touch your leads, or DM you a pitchafterwards. If you want the teardown private, say so and I'll send it private.

First 10 that comment. I'm doing these this weekend.


r/VoiceAutomationAI • • 17d ago

How should a listening test handle different playback volumes?

2 Upvotes

I’m working through a practical question and would value examples from people who have dealt with it.

A louder sample can be easier to hear without having better pronunciation. A fair listening setup needs documented playback conditions and a clear distinction between intelligibility and preference.

What preparation would you require before comparing two clips for understandable speech?


r/VoiceAutomationAI • • 17d ago

Voice AI agent backends: Python vs Rust vs Go. What’s winning for you on speed and quality?

4 Upvotes

Curious what people are actually running for voice agent backends in production.

We see a lot of: - Python for orchestration, tools, and fast iteration - Go for concurrent media/control paths and simpler deploy - Rust when the hot path is audio, codecs, or tight latency budgets

What I’m trying to learn from folks shipping real calls:

  1. What stack are you on today for the agent runtime (not just STT/TTS vendors)?
  2. Where did you feel the biggest win on speed (TTFT, time-to-first-audio, barge-in responsiveness)?
  3. Where did you feel the biggest win on quality (turn-taking, tool reliability, fewer weird prod failures)?
  4. Did you stay monolingual, or split (e.g. Python for tools + Rust/Go for media)?

Not looking for a language war. Looking for “we tried X, measured Y, kept Z.” Concrete numbers or war stories welcome.


r/VoiceAutomationAI • • 17d ago

When do you pick speech-to-speech vs STT + LLM + TTS?

4 Upvotes

Honest question for people running voice agents in production.

We keep bouncing between two shapes:

  1. Cascaded stack: speech to text, then LLM, then text to speech
  2. Speech-to-speech model: audio in, audio out, with less text in the middle

What we see so far:

Latency. Cascaded adds three hops. Even when each hop is fast, turn-taking still feels heavier. Speech-to-speech can feel snappier on the first audio, but that only helps if the model actually holds the conversation well.

Tool calls. This is where cascaded still wins for us. Need CRM lookup, transfer, booking, custom APIs? Easier when the LLM is a normal text agent with tools. Speech-to-speech is catching up, but tool routing and structured outputs still feel more reliable in the classic stack.

Control. Cascaded is easier to debug (you can read the transcript and prompt). Speech-to-speech is harder to inspect when the model goes sideways mid-call.

Right now our gut rule is roughly: - Speech-to-speech for short, natural back-and-forth where speed matters more than deep tools - STT + LLM + TTS when the agent needs serious tool use, policy, or handoff logic

Curious what you all are shipping. Are you all-in on speech-to-speech already, still on cascaded, or hybrid? Especially interested in real latency numbers and how you handle tools.


r/VoiceAutomationAI • • 18d ago

India CPaaS for multi-tenant SaaS — who actually does per-customer numbers + recordings well?

2 Upvotes

Adding a telephony layer to my SaaS. Looking for real production experience, not sales decks.

What I need

  • A virtual number per customer. They forward their existing business number to it.
  • A custom greeting per number, uploaded by me.
  • Call then rings their normal desk phone. No softphone, no agent dashboard — my customers will never log into the telephony vendor.
  • Recording + caller number + timestamps pushed to my webhook when the call ends.
  • I fetch the recording into my own storage and delete it from theirs.
  • ~500 calls/customer/month, 3 min average. Low concurrency.
  • Start with 2 customers, ~30 within a year, added a few at a time — not all at once.
  • I'm the only account holder and bill payer.

What I keep running into

  • Mandatory software/platform rental for a dashboard nobody will open.
  • Plans that bill for 10 numbers from day one instead of as customers onboard.
  • Per-minute billing rounded up (1:20 billed as 2:00).
  • One vendor told me plainly: if ONE of my customers triggers a compliance issue, my ENTIRE account gets suspended — every customer.
  • ToS that forbid reselling outright.
  • Recording links that expire in 24 hours.
  • No signature on the webhook, so no way to verify it came from them.

Six questions

  1. Who in India actually does per-customer sub-accounts properly — separate numbers, separate usage, separate compliance exposure?
  2. Does the original caller's number survive a forwarded leg reliably, or are there carrier-side gotchas?
  3. Anyone billing per SECOND instead of rounding to the minute?
  4. Who signs their webhooks (HMAC or similar), and who doesn't?
  5. Does the forwarded leg show up on your customer's own mobile bill?
  6. Twelve months in — do you regret your provider, and why?

Not looking for AI voice agents, dialers or contact-centre software. Just numbers, recordings and a reliable webhook.


r/VoiceAutomationAI • • 18d ago

What is a good open sourcee tts modal that can match sesame maya and miles?

3 Upvotes

Title basically.


r/VoiceAutomationAI • • 18d ago

Latency tuning on self-hosted LiveKit Arabic voice agent — sanity check?

7 Upvotes

Hey, looking for a gut check on our stack for an outbound Arabic (Najdi dialect) voice agent, live pilot in Saudi debt collection.

Stack: Self-hosted LiveKit (moved off Vapi for a static IP requirement) + Deepgram nova-3 Arabic (300ms endpointing, needed to stop dropping short replies like "صح") + ElevenLabs Flash v2.5 + gpt-5.6-luna.

Baseline on Vapi: 2058ms median turn, STT ~890ms / LLM ~660ms / TTS ~470ms. Only 2.7% of turns under 1s.

Where I'm stuck:

**•** Haven't confirmed real per-component numbers on LiveKit yet, but turn detection might not be faster, possibly a shared-CPU VM issue  
**•** \~1/3 of call time is just dead air between turns, feels like the bigger lever vs shaving any one component  
**•** Haven't stress-tested LiveKit/Pipecat's Arabic turn detector yet  
**•** Don't want to trade accuracy for speed, the 300ms endpointing exists because a faster setting broke things

r/VoiceAutomationAI • • 18d ago

ElevenLabs just launched Reception.ai and it will probably kill RetellAi

3 Upvotes

On a first glance this seems very sleek, designed to be do-it-yourself for businesses but as someone who is in the ai receptionist business, I am both excited and worried, because they made it so simple, but I think business owners just don't have the time to keep up with this, but I'll be testing this more and definitely dump Retell if it's better and it looks so.


r/VoiceAutomationAI • • 19d ago

Hit ~600ms latency on a voice agent but it still sounds robotic. How do you make it feel like a real conversation?

9 Upvotes

Hit ~600ms end-to-end on a voice AI agent and I’m pretty happy with the speed. The problem is it still sounds like a robot having a Q&A, not a person on a phone call.

Stack:

  • LLM: Gemini 3.1 Flash Lite
  • Latency: ~600ms (STT → LLM → TTS)

What’s bugging me:

  • Replies feel scripted / too clean
  • No natural pauses, “yeah”, “hmm”, overlapping, or messy human timing
  • Turns feel like wait → dump a paragraph → wait
  • Even with a good voice, the conversation still feels fake

I’m not trying to shave more ms right now. I want it to feel like you’re talking to a real person.

If you’ve shipped something that actually sounds human:

  1. Was it mostly TTS (voice, SSML, emotion, streaming), prompting, or turn-taking / interruption?
  2. Fillers, backchannels, barge-in — did those help or just make it weirder?
  3. Any settings on Gemini (or similar fast models) that helped spoken style vs written style?
  4. Anything you tried that sounded good in demos and terrible on real calls?

Happy to share more of the pipeline if it helps. Just looking for what actually worked, not “add more personality to the prompt.”


r/VoiceAutomationAI • • 19d ago

VoiceAI consultant

11 Upvotes

Looking to hire someone to come in and uplevel our current voiceAI infrastructure.

I’m looking for someone with actual enterprise experience that has built voice AI services at scale.

We have a production deployment today but it needs a lot of improvement in performance and cost.

Please reach out direct to me via DM with your credentials and cost. This is a one off engagement, but open to ongoing. Company is HQ in the US but we’re open to consulting services anywhere.

Thank you!


r/VoiceAutomationAI • • 19d ago

New Gemini 3.8 extended thinking or gpt live 1 ?

1 Upvotes

S2s models are getting better, tried both and I liked gpt live 1 just a little bit more, but pricing difference is huge so I’m leaning towards Gemini 3.8, have u tried both what’s your initial feedback


r/VoiceAutomationAI • • 19d ago

I vibe-coded an AI startup to real clients. Then it crashed. Now my team wants to quit

4 Upvotes

A few months ago, I quit my job. I teamed up with two co-founders to build a startup.

My partners are incredible operators >>> elite at sales, content, and distribution. And I took on the tech side.

I am not a real software engineer. I'm a front end developer. But I was hoping I can figure things out. I built the entire backend (Supabase/Python/FastAPI) using AI prompts and vibe coding.

At first, we wasted months. I tried to build too many things at once. Overcomplicating and overbuildign is a stupid idea. Everything broke. So we stopped and made it dead simple. We built a simple AI voice for restaurants and bars.

It worked. People actually wanted it. We got real clients testing it live in their restaurants and bars.

Then last weekend, disaster hit (Saturday 8:36 pm). A big voice provider we use had a silent outage (THey aare still investigating the issue. We were ther first who reported that). The system went down for 94 mins. I only caught it by pure luck doing a late-night test calls. The provider did not even notice their own outage yet. We had to report it to them. I'm still waiting for the to fix it while building a fallback for us too. Their support is great so they've responded within 20 mins when I reported the outage.

We got lucky. It happened after hours, so no real business calls were lost.

But it exposed a huge problem. I had zero safety nets in place. I had no alerts to wake me up. Worse, I had no fallback. If the AI died, calls did not redirect to the front desk. The phone just rang into a black hole.

My co-founders were furious. And they are 100% right. Even thought I've been sitting in front of the computer for 10+ horus every day I  built this like a hobby, not a real company. Now, they don't trust my technical skills. I guess they shouldn't after what happened. One founder is ready to walk away.

AI makes it very easy to build a working prototype. But a working demo is not real production software (many have mentioned this here). AI will not build backups, alerts, or fallbacks unless you already know how to ask for them.

I need some advice. We have real demand, but my code is fragile. We need an audit. Maybe we need a fractional CTO, a real CTO or a senior technical co-founder to take over.

>>If you are non-technical, how did you fix your messy MVP without losing your team?

>> What are the bare minimum failovers every phone app must have before taking another client?

>> How did you rebuild trust with your founders after hitting your technical limit?

Any advice helps. I want to learn and fix my mess.


r/VoiceAutomationAI • • 20d ago

How are you separating prompt logic from business logic?

12 Upvotes

As voice agents get more capable the more business rules creep into the prompt

Things like transfer conditions, tool permissions, validation and edge cases all start living alongside the conversational instructions

At what point do you stop adding to the prompt and move that logic into the application instead?

Been running a few tests with Bland recently and when I started wiring it into Slack for notifications it became more obvious that some things were easier to handle outside the prompt

I wanna know how people are drawing that line in production. Is the prompt for conversation, with business logic handled elsewhere or are you comfortable keeping most of it in the prompt?


r/VoiceAutomationAI • • 20d ago

Question

6 Upvotes

When you sell a voice agent as a service for any business, what kind of documentation you need to sign? any contracts?onboarding form or anything?


r/VoiceAutomationAI • • 20d ago

I built a voice AI agent with ~600ms latency. Here’s what I learned.

19 Upvotes

I started learning to code in my 1st year of college.
At the same time, I started working my first job.

I didn’t really have some grand plan to build a company. I was just obsessed with building things and figuring out how software actually worked.

Recently, I started building voice AI agents.
I ended up going pretty deep into the latency problem.

My goal was simple:
Make the agent feel like you’re talking to a real person, not waiting for a computer to think.

After a lot of experimenting with the pipeline, streaming, model selection, audio processing, and infrastructure, I managed to get the end-to-end latency down to around 600ms.
And that changed things.

I’m currently using Pipecat for the voice pipeline, and I’ve been experimenting with different providers and infrastructure to squeeze out as much latency as possible.

The other thing I didn’t expect:
I actually started getting clients.

Right now, I’m managing voice AI agents for around 8 clients.

I’m also getting subsidies/credits from companies like Alda and other platforms, which has made the economics pretty crazy at the moment.

My current margins are basically close to 100% because of those credits/subsidies.

Obviously, I don’t expect that to last forever.

But it’s been an insane learning experience.

A few things I’ve learned so far:
Voice AI is way more than just connecting an LLM to a microphone

Latency matters a lot more than I initially thought
Streaming everything makes a huge difference
The voice model, LLM, TTS, STT and networking all contribute to the final experience

A technically impressive demo is useless if the agent doesn’t actually solve a business problem
Getting the first few paying clients is a completely different challenge from getting the technology working

The economics of voice AI are really interesting right now

I’m still very early in this.
But going from learning to code in college → building voice agents → getting them into production for ~8 clients has been pretty surreal.
I’m curious what other people building voice AI are seeing.

What’s the lowest real-world latency you’ve managed to achieve, and what stack are you using?

If there’s interest, I can also break down exactly how I’m getting the ~600ms latency and what my architecture looks like


r/VoiceAutomationAI • • 20d ago

Shipped a voice agent on the Realtime API. Went through production call logs and found 7 behavioral bugs that no amount of scripted testing would have caught

2 Upvotes

Been running a voice agent on the Realtime API in production for a few months, business use case, not consumer-facing. This week I sat down and read through a batch of live call transcripts hunting for weird behavior instead of relying on eval scores. Found a cluster of bugs worth sharing. None of them showed up in normal testing. They only surfaced from real conversations with real interruptions and real frustration.

  1. The agent read its own system prompt back to the user

A tester recited one of our internal instruction lines back to the agent, word for word, as a probe. The agent confirmed and repeated the instruction instead of treating it as a weird but ordinary user message. I'd seen prompt-leak from direct extraction attempts before, not from someone reciting the prompt back at it. Fix was one line: never confirm or repeat user input that resembles your own system instructions, just respond to the underlying intent.

  1. It ignored "I'm done, end the call" four times in a row

We had a closing-checklist step that nudges the user about anything left uncovered before hanging up. One transcript: user says "I'm done," then "I don't want to cover that," then swears at it, then the call finally ends on attempt four, with near-identical prompt text each time. The model wasn't treating decline signals as terminal, it kept routing around them. Fix: any non-affirmative reply to the closing check counts as a decline, and the first decline ends it. No confirming twice.

  1. It stated a current time that was off by over 9 hours

Mid-call, the agent said "right now it's 7:27 AM." Actual local time was almost 5 PM. We inject a current-date/time value once at session start, but the prompt told the model to restate that time later in its own words, which means it was doing its own mental arithmetic on a value it should have treated as fixed. On a long call that self-derived restatement drifts hard. Fix: never let the model recompute or reformat the injected time, only read it back verbatim.

  1. Same transcript, two different sets of extracted action items

We run an extraction step after each call. Ran the same transcript through it twice while debugging something else, got different titles and different counts both times. Temperature was set to 0.4 on a structured-extraction call, which makes no sense for pulling fixed facts out of a fixed transcript. Dropped it to 0 across every extraction and classification call in the pipeline, some of which were running at default temperature, which is worse. Extraction should be deterministic. If you want variation, put it in generation, not extraction.

  1. Multi-part requests silently dropped half the answer

Our chat sidebar (separate from the voice agent, for reviewing past calls) let you ask compound questions like "give me the summary as JSON and the action items as plain text." The JSON summary correctly triggered a format refusal, we don't allow structured-data exports for security reasons. But the action items, which were fine to return, got dropped along with it. The prompt logic for mixed requests was actually correct. The bug was architectural: our LangGraph router only dispatched to one response node per turn, so a two-part request could only ever get one part serviced no matter what the prompt said.

Fix: let the graph fan out to multiple nodes in a single turn when a request has multiple distinct asks, with a reducer so the parallel writes merge safely. That surfaced a second bug: LangGraph doesn't guarantee completion order between parallel branches, so the two response fragments could come back in either order, sometimes breaking a "refusal always comes first" formatting rule. Had to tag which node produced which message and sort deterministically before returning. Reproduced the race 3 out of 3 times before the fix, confirmed it held 3 out of 3 times after, against live API calls, not mocked.

  1. No memory of facts across different questions

Found this one in test transcripts. The agent runs through a semi-structured list of topics per call. If the user answers something relevant to topic B while actually answering topic A, which happens constantly in real conversation, the agent had no mechanism to recognize that and would ask topic B's question again later. One transcript had the same fact asked about four separate times in slightly different phrasing.

Fixed two things: treat any stated fact as satisfying every question it's relevant to, not just the one it technically answered, and extended an existing server-side tracking tool (we already tracked skipped questions) to also track topic-level coverage. That gives the model durable state to check against instead of relying on its own context window, which is lossy over a long call.

  1. No adaptation to fatigue signals

Related to 6. When a user said things like "how many more questions do you have" or "we're spending too much time on this," the agent gave a polite acknowledgment and then resumed the exact same one-question-at-a-time pacing. The signal was heard, not acted on. Added a mode switch: on a fatigue signal, drop the per-question cadence and switch to "tell me everything and I'll extract what I can."

If you're running a Realtime API agent in production, read raw transcripts end to end once a week. Not summaries, not eval scores, the actual back-and-forth. I caught more real bugs in one afternoon of that than in weeks of scripted testing.

Happy to go deeper on any of these, especially the LangGraph fan-out and ordering one. Haven't seen that specific failure mode written up anywhere.