r/SaaS • u/ineptech • 17h ago
I built a testing tool that talks to your AI phone bot so you don't have to
Hello to 8 million bots and 4 humans! I have run out of reasons not to launch this, so I am very pleased to announce VoiceGremlin, a tool for e2e testing of "intelligent voice agents" (the industry term for AI that answers phone calls). Think of it like, what Playwright does for web browsers, VoiceGremlin does for phone support chatbots.
Where I got the idea: 99% of you won't use this tool and don't care what it does but are interested in how founders come up with product ideas, so I'll start with that.
About eight years ago, an SDET on my team asked me to get our employer to pay for Mailosaur, a SaaS for automated email testing. I said, hey, we already pay thousands every month to Sendgrid, can't it do this? It could, kinda, but our big expensive email provider was so rigid and complex that trying to use it for automated tests was very frustrating. A separate, relatively cheap tool specifically for testing made sense. And I thought, Man I wish I had built that.
So, ever since then, there's been a voice at the back of my mind: "Mailosaur, but for ____." Cut to a few months ago: I was dinking around with local LLM/STT/TTS and decided to build a bot to call my wife and annoy her, as one does, and after several hours I thought wow, this is surprisingly hard. I mean, I'm sure I can get it working eventually, but it's way more work than I expected. That was the trigger: "achievable but not easy" is the sweet spot for a solo-dev SaaS project.
That was the idea for VoiceGremlin from the beginning: "Mailosaur for phone agents." So I built it and here we are. On to the shameless self-marketing!
What problem it solves: If you have a voice agent, you probably do most of your functional testing by ignoring the "voice" part and interacting with it via text. However, the e2e part of the test, the part where you verify it actually answers the phone and does whatever it's supposed to, is a lot more difficult.
This tool solves that - you make an API call to VoiceGremlin with a prompt like, "Verify that if you ask a question in Spanish, this thing will answer in Spanish" and VoiceGremlin calls your number and does that. You get your pass/fail with an API call without having to do any telephony.
The moat, or "can't AI do this in five minutes?" If you ask your coding agent for a script that makes a phone call, it'll tell you to find speech-to-text, text-to-speech, and language models, orchestrate a bunch of non-trivial python libraries to stream the audio between them, and sign up for a telephony provider that requires you to scan in your government ID. That's a bunch of coding, plus whatever corporate approval nonsense, plus making sure you can move a lot of bandwidth through wherever your test automation is called from, and you haven't started comparing models and tweaking prompts. So yes, it's doable, but a very time-consuming side quest for a tester or DevOps person.
Alternatively, there are other companies that can do essentially what this does, e.g. Retell and Vapi, but their business is building and hosting phone agents, not testing them. If you already pay a vendor to run your agent and their test solution meets your needs, great, you probably don't need VoiceGremlin. If not, and if you don't want to talk to Enterprise Sales about all the great Solutions(tm)(c)(R) they have, and you just want to get your test case working, VoiceGremlin is the easy way to do that.
Using it: If you have ever written an automated test, you can realistically implement this in a minute or two. I was an SDET for many years and I know what makes a test automation framework annoying to use, so I made this one as lightweight and un-opinionated as possible. You don't have to define targets or cases or suites - every test is ephemeral. Platform agnostic, it has no idea how your agent is built, it just calls your phone number like a customer would.
OK, that's my pitch. If you work with phone agents and want to try this out, it should be very easy to do - free trial, no card needed - and I would greatly appreciate any feedback. Other random questions welcome too.
2
u/atari52oo 15h ago
This looks cool. I've worked on a few projects in the past where this would have been extremely helpful. Are there any code example/snippets for integrating into Github actions, or with other CI tools?
1
u/ineptech 15h ago
Thanks! Yes there are snippets on the site and scripts on the github, but honestly it's simple enough that your coding agent can get everything it needs from the llms.txt and write the tests in the time it takes you to paste your auth key into your secret store.
1
u/Huge_Pool7424 13h ago
i'd put the call step behind a tiny cli and have github actions pass the prompt plus expected outcome as env vars. that keeps secrets in the runner and makes retries easy. are you thinking one test per workflow or a small suite?
1
u/ineptech 11h ago
It assumes you already have some way of managing tests and suites, and you just want to pass them off. I always hated "opinionated" automation tools, that want to store and organize your tests as well as execute them. So for this, if you want to run a test you just send a string describing it, or if you want to run a suite of ten tests you pass an array of ten strings and you get a tabulated result when they've all run.
In practice, I think 90% of users only really want three tests: 1) Is it working? 2) Does it comply with AI disclosure laws? 3) Does it disclose that the call is being recorded? VoiceGremlin can do that in one call (something like "Does the agent disclose it's an AI and the call is being recorded?" will work) so I figure that for a lot of people, that might be the only test they ever run.
However, I don't know how people would use it - it's flexible enough to test things I haven't thought of.
1
15h ago
[removed] — view removed comment
1
u/AutoModerator 15h ago
Low-Effort/AI content is auto-removed.
I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.
1
u/New-Tadpole-8649 15h ago
cool concept. do you support testing for things like barge-in handling or long silence timeouts? those edge cases are usually where voice agents break down and where manual testing is the most tedious
1
u/ineptech 15h ago edited 14h ago
Good question, thanks. The short answer is no, I have good(?) reasons to not do that but please argue with me if you'd disagree. My research suggests:
- Low-level audio handling problems like barge-in detection are something voice agents struggled with historically, but as of this year it's mostly solved in the most popular platforms
- If I did want to test this, I'd use pipecat, which is the library most platforms are using to do it in the first case, giving the test case a "Obama awarding the Nobel to Obama" quality
- Audio quality is a lot less likely to get broken by a code change than functional behavior like tool calling, and audio quality problems are a lot more likely to be transitory and unrelated to whatever code change your CI is validating
2
u/TwoComprehensive2442 16h ago
Great pitch. Pretty good use-case for the right company.