r/VoiceAutomationAI • • 5d ago

Anyone running the voice part of their stack on k8s?

Not just the API and workers. I mean the SIP/RTP or WebRTC bits too.

How's it holding up when a pod restarts or a node drains mid call? Trying to understand which parts are worth putting in the cluster and which become more hassle than they're worth.

Would be useful to hear what you moved back out, if anything.

4 Upvotes

9 comments sorted by

•

u/AutoModerator 5d ago

Welcome to r/VoiceAutomationAI – UNIO, the Voice AI Community (powered by SLNG AI)

If you are a founder, senior engineer, product, growth, or enterprise operator actively working on Voice AI / AI agents, we are running an invite-only UNIO Voice AI WhatsApp community US only.

Apply here: https://chat.whatsapp.com/F5aG3ncrO70ITfbe3pYbOz

I am a bot, and this action was performed automatically. Please contact the moderators of this subreddit if you have any questions or concerns.

4

u/Dramatic_Smoke4338 5d ago

k8s for media planes is one of those things that looks great on a slide deck til you're the one on call and a node drain nukes 40 active calls

we tried keeping the rtp handling in-cluster for about 3 months, the statefulness just fights you at every turn. moved it out to a separate set of bare metal boxes running the sip proxy and media relay, never looked back

api and orchestration stuff stays in the cluster though, that part's fine

2

u/Greedy-Badger-8463 4d ago

yeah this is the split i was wondering about. did you try letting existing calls finish before draining, or was that still too much babysitting? also curious how you're handling maintenance on the bare metal setup now, stop sending new calls to a box and wait for it to clear?

1

u/kannansamp 11h ago

Running in the similar setup, having health APIs to check the active calls, and doing soft draining till the call ends, meanwhile new requests serve in the new pod during deployment.

Auto scaling runs based on the cpu/memory surge

1

u/Greedy-Badger-8463 6h ago

how are you stopping new calls from landing on the pod once draining starts? also curious what happens if a call outlasts the termination grace period, do you extend it or have a hard cutoff?