How are you running real-time Voice AI media workloads on Kubernetes in production?

0
0
Asked By MellowPine47 On

We're evaluating a Voice AI and agent-builder architecture involving real-time voice agents, SIP/PSTN traffic, WebRTC, STT → LLM → TTS pipelines, WebSocket streaming, workflow orchestration, concurrent calls, and GPU-based speech workloads. Kubernetes seems like a strong fit for the control plane, APIs, workers, scaling, and observability, but the real-time media layer is more difficult. UDP port ranges, NAT traversal, STUN/TURN, RTP routing, UDP load balancing, stable media paths during scaling or restarts, and avoiding unnecessary latency all create challenges that do not appear in a typical SaaS deployment. We're considering placing a dedicated STUN/TURN and media gateway in front of Kubernetes so it handles NAT traversal and real-time connectivity while Kubernetes manages orchestration and compute. For those running LiveKit-based systems, custom Voice AI platforms, or similar production stacks, what architecture has worked for you? Do you terminate SIP/RTP outside the cluster, run TURN inside it, use host networking or dedicated media nodes, or rely on MetalLB, direct server return, or an external UDP load balancer? What started breaking when you reached hundreds or thousands of concurrent calls?

4 Answers

Answered By NorthStarLime5 On

We had a rough experience keeping RTP inside Kubernetes. We eventually moved the media edge outside the cluster onto a small group of dedicated, high-capacity nodes running coturn and a thin SIP proxy. Those gateways terminate the external traffic and forward media internally, while Kubernetes remains responsible for orchestration. Once we passed roughly 300 concurrent calls, host networking became unreliable: packet loss increased and stale conntrack entries left pods in bad states. Moving media termination out of the cluster resolved most of those issues. It’s also important to test live calls during rolling deployments, because a functioning TURN path does not help if the process owning the session is killed.

BrightOtter31 -

What led you to build a custom SIP proxy instead of using an existing gateway?

QuietHarbor64 -

The key lesson was that external connectivity and session survival are separate problems. The media layer needs to stop accepting new calls during a drain and allow active calls enough termination time to finish cleanly.

Answered By CopperFable6 On

LiveKit-style draining behavior is a useful model even if you use a different media stack. During a deployment or node drain, stop assigning new rooms or calls to the instance while allowing existing sessions to finish, and provide a realistic termination grace period. Test this separately from TURN connectivity by running an active call through a rolling deployment. A reachable media path alone cannot preserve a call when its owning process is terminated.

Answered By CloudyMarble8 On

A practical approach is to treat this as two separate systems: Kubernetes for the control plane and elastic compute, and a deliberately simple media plane for SIP, RTP, and TURN. Terminate external traffic on dedicated gateways and keep each UDP flow tied to a stable owner for the duration of the call. Pod churn and mid-call rebalancing are usually more dangerous than raw CPU usage. If media workers remain in Kubernetes, isolate them on dedicated node pools, use topology-aware placement, add long termination grace periods and disruption budgets, and drain nodes by rejecting new sessions before affecting active ones. An external UDP load balancer or direct-return path is generally preferable to piling on kube-proxy hops and conntrack pressure. Track packet loss, jitter, round-trip time, retransmits, call setup time, and audio gaps—not just CPU and memory.

SilverCactus22 -

That separation makes sense: Kubernetes handles orchestration and scaling, while dedicated gateways keep SIP/RTP/TURN stable with one owner per call. Media quality metrics are a much better scoreboard than ordinary pod health metrics.

Answered By AmberKoala19 On

Don’t overlook clock and scheduling behavior. If FreeSWITCH or another process owns the media timers, use a monotonic clock and verify that node timekeeping and CPU scheduling remain stable. Containerized workloads can still experience timing drift when the node’s clock source or CPU throttling interferes with media processing. This is worth checking before assuming the network is the only problem.

BlueWillow73 -

Timing issues are easy to miss because the stack may look healthy while audio quality degrades. I’d test clock stability and CPU throttling alongside the network path.

Related Questions

LEAVE A REPLY

Please enter your comment!
Please enter your name here

This site uses Akismet to reduce spam. Learn how your comment data is processed.