OpenAI's engineering team has published details on how it handles real-time voice for ChatGPT and its Realtime API. The key innovation: splitting traditional WebRTC media routing and protocol termination into two separate layers.
A stateless relay handles UDP packet forwarding, while a stateful transceiver manages the full ICE (Interactive Connectivity Establishment) and DTLS (Datagram Transport Layer Security) handshakes, plus encryption. The relay never decrypts media or touches codec negotiation — it just parses the ICE ufrag (username fragment) from the STUN header to decide where to forward each packet.
This design solves the classic WebRTC-on-Kubernetes headache. In the traditional setup, each session burns a public UDP port, so high concurrency means exposing tens of thousands of ports — a nightmare for security auditing and autoscaling. With the split, the public UDP surface shrinks to a small fixed set of addresses and ports, the relay can scale horizontally, and if it restarts, the route is rebuilt as soon as the next STUN packet arrives.
OpenAI also ditched the Selective Forwarding Unit (SFU) architecture common in multi-party calls. Since the vast majority of its voice AI sessions are one-on-one, the transceiver model offers lower latency and means the backend doesn't have to act as a WebRTC peer.
The relay is written in Go, runs in userspace, and uses SO_REUSEPORT to let multiple worker threads share the same UDP port, with runtime.LockOSThread to pin threads for better cache locality. No kernel bypass framework is used. A global relay deployment, combined with Cloudflare geo-steering, routes signaling and media streams to the nearest OpenAI network edge, cutting first-hop latency.
The blog post also reveals that Justin Uberti — the original architect of the WebRTC protocol — and Sean DuBois, creator and maintainer of the open-source WebRTC library Pion, have both joined OpenAI to work on merging real-time AI with WebRTC.