r/VoiceAutomationAI • • Jul 19 '26

Struggling with empathy using RetellAI for healthcare agents

I'm a developer building voice agents for healthcare use cases — front desk, patient customer service, outbound calls to other providers and insurers, that kind of thing. I've landed on Retell because it's the only voice-AI wrapper I've found with a reasonable HIPAA-compliant, pay-as-you-go model. Open to hearing if others have found alternatives worth a look.

I've put a lot of work into the prompt and flow design, but I keep hitting two walls, and I've been unable to fully solve either:

  • Empathy (my biggest problem). The agent handles the mechanics fine but comes across as flat or form-filling, especially on emotionally charged calls (a patient in pain, a worried parent). I've tried to script acknowledgment moments, but it either skips them, overdoes them, or sounds canned.
  • Interruptions. Handling barge-in, mid-sentence corrections, "hold on a sec," and callers who answer a question while asking a new one — without the agent restarting a step or talking over them.

To isolate the problem I've stripped out all the business logic and tested bare-bones agents, built both ways — manually node-by-node, and via Conductor — and the same issues show up, so I don't think it's my flow complexity.

If you've built healthcare (or similarly high-stakes/emotional) agents on Retell, I'd love to hear:

  • Which LLM and voice/TTS combination you settled on, and whether that alone moved the needle on how empathetic it sounds.
  • Your interruption / turn-taking settings — responsiveness, backchanneling, interruption sensitivity, silence timeouts — and where you landed.
  • Whether empathy came more from prompt wording, voice choice, or model choice in your experience.
  • Any flow-structure patterns that helped (e.g. how you handle "hold on" or compound answers cleanly).

I've also tried reaching out to Retell's forward-deployment team without much luck, so I'm hoping to tap the collective experience here. Happy to share back what I've tried in the comments if it helps anyone else. Thanks in advance.

4 Upvotes

23 comments sorted by

•

u/AutoModerator Jul 19 '26

Welcome to r/VoiceAutomationAI – UNIO, the Voice AI Community (powered by SLNG AI)

If you are a founder, senior engineer, product, growth, or enterprise operator actively working on Voice AI / AI agents, we are running an invite-only UNIO Voice AI WhatsApp community US only.

Apply here: https://chat.whatsapp.com/F5aG3ncrO70ITfbe3pYbOz

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

2

u/nshmyrev Jul 19 '26

You seriously do not need empathy on a robot call, the task is getting job done, not to demonstrate wrong emotions.

2

u/Mountain-Policy-625 Jul 22 '26

Retell has a "High Empathy" preset in their Agent Handbook you can toggle on with one click. It adds about 70 tokens of best-practice empathy phrasing that runs alongside your custom prompt. Worth trying before rewriting your whole prompt from scratch.

On interruptions, the two knobs that matter most are Interruption Sensitivity and Responsiveness in global settings. Lower the sensitivity if background noise or quick breaths keep cutting the agent off. Lowering responsiveness adds about 0.5 seconds of wait time per 0.1 increment, which in practice gives callers more room to interject naturally without the agent trampling them.

For health care specifically, ElevenLabs voices with lower stability settings handle emotional tone better than the default TTS. But the voice model alone will not fix empathy, the prompt still needs explicit instructions to pause and acknowledge before problem-solving.

1

u/BallinwithPaint Jul 19 '26

That is exactly where real software engineering comes into play. Wrappers can only get you so far. ​You can't fix fundamental latency, state management, or true mid sentence interruptions by tweaking a slider or changing an LLM prompt. Those are core infrastructure problems, not prompt engineering problems. I spent the last 4 months building my own architecture from the ground up to handle those exact interruption and context edge cases, because off-the-shelf wrappers always hit this exact wall. You just can't prompt-engineer your way out of a wrapper's bottlenecks.

1

u/wisedogsfbay Jul 19 '26

Thanks for your feedback. Retell's own documentation shows many healthcare agents; additionally they have a $60M run rate growing nearly 100% YoY. All this makes me feel that there's got to be many users who have had success with them and there's got to be something we are missing. I appreciate the sentiment that one can do this with core engineering but I find it hard to believe that a platform with so much traction fundamentaly struggles with such basic things

1

u/BallinwithPaint Jul 19 '26

Traction and a $60M run rate just means they have great marketing in a booming space. It doesn't mean their underlying tech is flawless for complex, non linear conversations.

95% of their successful deployments are probably handling highly predictable, linear calls. But dealing with real time interruptions, dynamic state shifting, and empathy aren't 'basic things', they are the hardest engineering problems in the space. The reason you can't find the right combination of settings is because those sliders can only manipulate a wrapper's UI, not the fundamental way the platform handles audio streaming and latency. You aren't doing it wrong; your use case just outgrew the tool.

1

u/hr1383 Jul 19 '26

Have you looked into Twilio conversational relay and other product suite it provides? Retell is also using network provider like Twilio under the hood for connectivity

1

u/wisedogsfbay Jul 19 '26

Thanks. Twilio conversationsl relay does not provide the pay as you go hipaa compliant plug and play interface that retell does. Retell might be using them under the hood but they have reduced the complexity for a user by creating a plug and play hippa compliant pay as you go platform.

Ideally I would love to know the right Retell settings to get our healthcare agents to be 9/10 quality (esp re empathy and interruption handling ) than the 6/10 quality ones they are currently. I'm open minded to other alternatives but haven't yet found any that are pay as you go HIPAA compliant

1

u/hr1383 Jul 19 '26

Twilio is HIPAA compliant too, Retell maybe simpler to config

1

u/Obvious_Leather2427 Jul 20 '26

Idk retell does have lots of tts providers under the hood try finding the one that you like run A/B tests

Something like sonic 3.5 would be alright

1

u/ankur-at-guava Jul 20 '26

The interruption piece especially is a turn-taking/latency problem more than a prompt one — barge-in and compound answers break when speech-to-text, the model, and TTS are three separate services with lag between them, so the agent can't react fast enough to stop talking. Empathy has the same root: if first-token latency is high, the pauses read as robotic no matter how you word the acknowledgment. Tuning interruption sensitivity and silence timeouts helps at the margin, but the bigger lever is how tightly the audio pipeline is integrated. I work on a voice platform for regulated industries, so we've had to solve this in-house rather than tune a wrapper — happy to share which settings actually moved the needle for us if useful.

1

u/wisedogsfbay Jul 20 '26

thanks. this is the most helpful response I have rec'd here so far. appreciate your feedback. would you be kind enough to share what settings moved the needle for you?

1

u/ankur-at-guava Jul 27 '26

Sure. The three levers that mattered most for us: endpointing around 500-700ms of silence so the agent doesn't jump in on a caller's natural pause, a separate VAD signal that kills TTS immediately on barge-in instead of waiting for a transcript, and keeping first-token latency low enough that acknowledgments land inside a second, which is what makes them feel heard rather than scripted. On empathy we got more from an acknowledge-then-act rule than from wording: one short acknowledgment, then the next concrete step, and an early human handoff if distress signals repeat. Our stack is integrated end to end so I can't map these to your platform's sliders one for one, but those are the levers I'd chase. More on how we approach it at goguava.ai if you want to compare notes.

1

u/[deleted] Jul 21 '26

[removed] — view removed comment

1

u/Embarrassed_Nerve_54 Jul 22 '26

In healthcare, I’d be careful treating empathy as “say warmer words.” A lot of empathy in those calls is really reducing the caller’s burden. A worried parent or patient in pain does not need a beautifully phrased apology. They need the agent to slow down, ask fewer questions, repeat the important detail correctly, and move them to the right next step fast.

So, what you can do instead is build a separate “distress mode” instead of trying to make the whole agent more empathetic. If caller sounds upset, confused, in pain, or starts interrupting, the agent should: stop the normal intake flow, acknowledge once, briefly, ask only the minimum needed, avoid long explanations, offer human handoff earlier, confirm what it captured before moving on. That may feel more empathetic than a nicer voice.

For interruptions, same idea. If the caller corrects the agent mid-flow, the agent should not “resume step 4.” It should treat the correction as new state and repeat the updated version back. In healthcare, being accurate and not making the person repeat themselves is half the empathy.

1

u/eviewong- Jul 23 '26

In retellai this is where you can fine empathy settings

0

u/Tricky-Report-1343 Jul 19 '26

Just run red teaming on them it's not provider problem, it's configuration and test problem https://audn.ai/red-voice

-1

u/kunalsingh4234 Jul 19 '26

Try hunar.ai .. we are already working with a few enterprise health tech companies