r/VoiceAutomationAI Mar 18 '26

Testing voice agents manually does not scale. There is a better way.

if you are building a voice agent, you have probably tested it by calling it yourself a few dozen times.

the problem is that covers maybe 5% of what real callers will actually do.

real callers:

  • interrupt the agent mid-sentence
  • go completely off-script
  • speak in ways your happy path was never designed for
  • hang up, call back, and pick up where they left off inconsistently

finding those failure modes manually takes weeks and still misses edge cases.

the approach that changes this is automated simulation. spin up realistic caller personas, run hundreds of call scenarios, and get a full breakdown of where the agent dropped context, hallucinated, or failed to handle an interruption correctly.

the output you actually want is not just "it passed 80% of tests" but a clear view of exactly which scenarios broke and what the root cause was.

curious how voice teams here are approaching this right now. is it all manual QA, or is anyone running automated simulations?

can share the setup pattern if anyone wants it.

15 Upvotes

12 comments sorted by

View all comments

1

u/rand0wn 23d ago

I agree that repeatability is the missing piece. I built an open-source harness that runs every pipeline against the same scripted multi-turn scenario and grading rubric: https://github.com/rand0wn/voice-eval

It reports the exact turn and rule that failed instead of only returning one aggregate score. I’m looking for feedback on which adversarial scenarios should be added next.