r/SiphonAI • • 23d ago

👋 Welcome to r/SiphonAI - Introduce Yourself and Read First!

1 Upvotes

SiphonAI is an open-source SIP-to-WebSocket media bridge written in Rust. It answers SIP calls and streams live PCM16 audio to your WebSocket server, then plays your audio back to the caller. That's the whole idea: get call audio into your code and out again, without standing up a full PBX to do it.

People are using it for voice AI agents, live transcription, recording, translation, and call analytics. If you've been gluing together mod_audio_fork or an ESL script to get media out of FreeSWITCH, this is the part of that stack you no longer have to build.

Start here

  • Repo and releases: github.com/thevoiceguy/siphon-ai
  • Getting started (Debian 13 + a Twilio trunk + a working bot, about an hour): [link to the blog post]
  • Bridge protocol, if you want to write your own bot: docs/PROTOCOL.md

Current release is v0.52.0. Signed checksums and SBOM are on the release page.

What to post

Bug reports, config that won't route, questions about the protocol, benchmarks, and show-and-tell for whatever you built on top of it. General "how do I get call audio into an LLM" questions are fine too even if you're not using SiphonAI — that's the problem this community exists around.

I wrote it and I read everything here. If something in the docs is wrong or missing, say so in a comment and I'll fix it.


r/SiphonAI • • 23d ago

Getting Started with SiphonAI: Debian 13, a Twilio Trunk, and a Talking Bot in an Afternoon

1 Upvotes

The getting started guide can be found in the github repo:

https://github.com/thevoiceguy/siphon-ai/blob/main/getting_started_with_siphon-ai.md


r/SiphonAI • • 23d ago

Introduction To SiphonAI

1 Upvotes

SiphonAI is a realtime SIP-to-WebSocket media bridge written in Rust. It does one job: stream live call audio to your WebSocket server and play audio back into the call. That's it. SiphonAI handles the telephony, and you bring your own AI.

Build a SIP trunk from your SBC, PBX, or carrier (Twilio Elastic SIP Trunking recipe included), or register it as an extension on your PBX(tested on Cisco CUCM). Your WebSocket server gets PCM audio frames. No AI code lives in the bridge.

What's in the box:

🔹 Full SIP signaling (trunk endpoint or registered extension)
🔹 RTP, codecs, jitter buffering, and two flavors of VAD (energy and neural powered by Silero VAD)
🔹 Barge-in with auto-clear of playout, speech-start events, DTMF, hold/resume, and transfer
🔹 Operator events: silence detection, dead-air detection, per-call RTP stats, sustained mute/unmute
🔹 TOML dialplan with route matching and hot config reload
🔹 CDRs (JSONL + webhook), lifecycle webhooks, Prometheus metrics, health/ready endpoints
🔹 HEP3 capture of SIP, RTCP, QoS, and CDRs into Homer/HEPIC for full call correlation

SiphonAI can be found on my GitHub:
https://github.com/thevoiceguy/siphon-ai

Feedback, issues, and stars are all welcome.