Menu

Categories

Tags

OpenAI open-sources the network protocol that keeps its AI clusters running

May 7, 2026 | alex | OpenAI | 170 views 0 comments

OpenAI just open-sourced a networking protocol called MRC (Multipath Reliable Connection), developed alongside AMD, Broadcom, Intel, Microsoft, and Nvidia. The spec is now available through the Open Compute Project (OCP), the industry's largest open-source hardware standards body. AMD, Broadcom, Microsoft, and Nvidia all published companion blog posts.

Training large language models requires tens of thousands of GPUs to stay perfectly in sync. A single training step can involve millions of data transfers — if one packet arrives late, every GPU sits idle. The bigger the cluster, the more frequent the link flaps and failures.

Traditional networks have a problem: if a single link goes down, the entire training job can crash, forcing a rollback to the last checkpoint. Recouting paths through switches takes seconds or even tens of seconds. For OpenAI's Stargate project — its massive compute infrastructure initiative — the first bottleneck was the network.

MRC's innovation is packet spraying. Instead of sending an entire transfer down a single path, it shatters the data into pieces and sends them across hundreds of paths simultaneously. At the destination, they're reassembled by memory address.

Link failures are bypassed in microseconds — no need for switches to recalculate routing tables. OpenAI also ripped out BGP entirely, replacing it with SRv6 source routing: the sender specifies exactly which path each packet takes, turning switches into dumb forwarding boxes. The failure domain shrinks dramatically.

The network architecture simplifies as a result. Connecting a hundred thousand GPUs used to require three to four layers of switches. MRC's multi-plane design cuts that to two layers, reducing power, cost, and points of failure.

The protocol, which connects 100,000 GPUs with just two switch layers, is already deployed across all of OpenAI's largest Nvidia GB200 supercomputers, including the Stargate site in Abilene, Texas (built with Oracle) and Microsoft's Fairwater datacenter. Multiple OpenAI models have been trained using it.

The most concrete example: during a recent frontier model training run (the one powering ChatGPT and Codex), the team hot-swapped four core switches without needing to coordinate with the training team. Multiple link flaps occurred every minute — with no measurable impact on the training job. Previously, that kind of incident would have crashed the entire task.

Leave a Reply

Your email address will not be published. Required fields are marked *