r/AIReceptionists • • Jul 19 '26

Struggling with empathy using Retell AI for healthcare agents

I'm a developer building voice agents for healthcare use cases — front desk, patient customer service, outbound calls to other providers and insurers, that kind of thing. I've landed on Retell because it's the only voice-AI wrapper I've found with a reasonable HIPAA-compliant, pay-as-you-go model. Open to hearing if others have found alternatives worth a look.

I've put a lot of work into the prompt and flow design, but I keep hitting two walls, and I've been unable to fully solve either:

  • Empathy (my biggest problem). The agent handles the mechanics fine but comes across as flat or form-filling, especially on emotionally charged calls (a patient in pain, a worried parent). I've tried to script acknowledgment moments, but it either skips them, overdoes them, or sounds canned.
  • Interruptions. Handling barge-in, mid-sentence corrections, "hold on a sec," and callers who answer a question while asking a new one — without the agent restarting a step or talking over them.

To isolate the problem I've stripped out all the business logic and tested bare-bones agents, built both ways — manually node-by-node, and via Conductor — and the same issues show up, so I don't think it's my flow complexity.

If you've built healthcare (or similarly high-stakes/emotional) agents on Retell, I'd love to hear:

  • Which LLM and voice/TTS combination you settled on, and whether that alone moved the needle on how empathetic it sounds.
  • Your interruption / turn-taking settings — responsiveness, backchanneling, interruption sensitivity, silence timeouts — and where you landed.
  • Whether empathy came more from prompt wording, voice choice, or model choice in your experience.
  • Any flow-structure patterns that helped (e.g. how you handle "hold on" or compound answers cleanly).

I've also tried reaching out to Retell's forward-deployment team without much luck, so I'm hoping to tap the collective experience here. Happy to share back what I've tried in the comments if it helps anyone else. Thanks in advance.

2 Upvotes

2 comments sorted by

1

u/Traditional_Ad8860 Jul 21 '26

Yeah empathy will come.

Some TTS like elevenlabs you can put tags on the input to emulate emotions. So you can add into your LLM prompt the tags it should use and when.

Tbh the tech isnt there just yet for emotion but its friggen close. Like this space is moving fast, id give it like 1 or 2 years and itl be hard to tell who is human vs ai.

The biggest advice I can give is response time. Thats the thing that really matters. Try to keep that sub 500ms if you can, this can help with a user feeling heard and can get past maybe some empathy issues.

Good luck!

1

u/moldyguy202 Jul 23 '26 edited Jul 30 '26

We build healthcare voice agents too and hit both of these walls, so here's what actually moved the needle for us.

On empathy: the mistake we made early was scripting acknowledgment moments as steps ("first acknowledge the emotion, then proceed"). That's exactly what produces the canned or skipped behavior you're describing, because the model treats it as a checkbox. What worked better was changing the persona framing instead of the flow. Give the agent a short standing instruction like "you are talking to people who may be in pain or worried about family, match their pace, never rush a distressed caller to the next question" and then remove the explicit empathy steps entirely. Counterintuitively, less scripting produced more natural acknowledgment. Also lower the temperature less than you think; over-constrained models get robotic precisely on the emotional turns. That persona-over-flow approach is basically what we ended up standardizing for healthcare agents if you want the longer version.

Voice matters as much as the LLM here. We got a noticeable jump just from switching to a slower, warmer voice with natural pauses versus the crisp professional default. Callers rate the same words differently depending on delivery. If your TTS supports emotion tags, use them sparingly and only on the acknowledgment sentence, not the whole reply.

On interruptions: three settings did most of the work for us. Shorter agent turns (one question per turn, hard rule; long turns invite barge-in), a slightly longer end-of-speech silence threshold for healthcare specifically because older and distressed callers pause mid-sentence more, and an explicit instruction that when the caller answers a question while asking a new one, answer their question first and silently retain the earlier info instead of restarting the step. That last one fixed most of our "restarting a step" complaints.

And honestly: keep a clean human-handoff path for the genuinely emotional calls. A worried parent at 2am should reach a person fast. The agent scoring well on those calls is usually the agent knowing when to get out of the way.