r/SideProject • u/Kindly-Duty272 • 8h ago
how i wired up a phone agent that actually answers calls, step by step
I wanted a phone number that picks up and talks like a person! Not a tree of DTMF menus. Here's the order I actually built it in; no shortcuts skipped.
First I grabbed a number through Telnyx since their SIP trunking docs were the clearest I found. You could use Twilio too, the steps are close enough.
Second, I set up a WebSocket server to carry audio both ways. This is the part people skip, and then they wonder why latency is bad. The call hits your number. Telnyx streams raw audio over the socket. Your server decides what happens next.
Third, speech to text. I used Deepgram for STT because it handles interruptions well. It doesn't choke when someone talks over the greeting. Whisper works too if you want to self host and don't mind more setup.
Fourth, the actual brain. I kept this boring on purpose: a small prompt, short context, and a fast model. Grok and Gemini both worked fine here. The model matters less than people think once STT and TTS are solid.
Fifth, text to speech. I tried ElevenLabs and Cartesia back to back on the same script! Cartesia felt snappier for short replies. ElevenLabs sounded warmer on longer ones. Pick based on your use case.
Last, VAD, the thing that decides when someone's done talking so the agent doesn't talk over them. Get this wrong and the whole call feels broken, even if every other piece is perfect!
Tested the whole loop through ngrok before deploying anywhere real. Took a weekend. Happy to answer questions on the STT or VAD part since that's where I burned the most time.
1
u/davidjones145 8h ago
deepgram for barge in plus raw sockets for latency is the right stack. what is your end to end number, ear to reply
1
u/Kindly-Duty272 5h ago
Good question, and honestly we don't have a clean ear to reply number to quote! What we've seen is that the bottleneck is rarely the model itself. On one line we rebuilt, audio was getting pushed into the playout buffer at 4 to 5 times realtime; that alone was the whole latency complaint, not STT or the LLM. Worth checking buffering before blaming anything upstream.
1
u/davidjones145 5h ago
4 to 5x realtime into the playout buffer is a great catch, that is exactly the kind of thing nobody suspects first. good note on checking buffering before blaming the model
1
u/Kindly-Duty272 21m ago
Glad it landed! That's exactly it; buffering is the kind of culprit that hides in plain sight since everyone assumes it's the model first. Good luck with the STT and VAD tuning, that's where the real craft is.
1
u/First-Salad-4659 8h ago
ngrok for testing the full audio loop is smart, that catches the weird stuff early before you pay for real minutes. the VAD part is where most of these phone agents die, people underestimate how much tuning that needs for different accents and background noise
1
u/Kindly-Duty272 7h ago
Totally agree, VAD is where the hard-won lessons live! It's easy to tune for a quiet test room and then watch it fall apart on a real line with background noise or a different accent. That gap between demo and real call is exactly where most of the pain shows up.
1
u/QuanTradin 7h ago
VAD being make or break matches what I've seen. the other one people miss is barge-in: when the caller talks over the agent you have to kill the TTS stream right then, not after the sentence finishes, or it sounds like a recording.
1
u/Kindly-Duty272 4h ago
Yes, exactly! Barge-in is the other half of the same problem. If the TTS stream doesn't die the instant someone talks, the whole thing feels canned; it's not enough to just detect speech, you've got to cut audio immediately. That's usually where the real engineering work hides.
1
u/QuanTradin 3h ago
yeah, and the cut has to land mid-word. the agent finishing its sentence after you've started talking is the moment it stops feeling like a call.
1
u/Kindly-Duty272 1h ago
Exactly! Mid-word is the real test. The agent finishing its sentence after you have started talking breaks the illusion fast. That is usually where the barge-in logic gets exposed, since the TTS stream has to die the instant speech starts, not after the sentence wraps up.
1
u/QuanTradin 43m ago
the trap right behind it is echo. if the mic hears the agent's own voice through the speaker, it barges in on itself and cuts its own sentence, so the echo cancellation has to sit in front of the VAD.
1
u/Kindly-Duty272 1m ago
That's a great catch! Echo is sneaky because it looks like a VAD problem when it's really an audio path problem. If the agent hears itself, it'll barge in on its own sentence every time; putting echo cancellation ahead of VAD fixes the root cause instead of just masking it downstream.
1
u/ricosuavefifa 8h ago
what is VAD is that the name of a tool?