OpenAI just open-sourced a networking protocol called MRC (Multipath Reliable Connection), developed alongside AMD, Broadcom, Intel, Microsoft, and Nvidia. The spec is now available through the Open Compute Project (OCP), the industry's largest open-source hardware standards body. AMD, Broadcom, Microsoft, and Nvidia all published companion blog posts.
https://twitter.com/OpenAI/status/2052025532485902368
Training large language models requires tens of thousands of GPUs to stay perfectly in sync. A single training step can involve millions of data transfers — if one packet arrives late, every GPU sits idle. The bigger the cluster, the more frequent the link flaps and failures.
Traditional networks have a problem: if a single link goes down, the entire training job can crash, forcing a rollback to the last checkpoint. Recouting paths through switches takes seconds or even tens of seconds. For OpenAI's Stargate project — its massive compute infrastructure initiative — the first bottleneck was the network.
https://twitter.com/OpenAI/status/2052025533937103102
MRC's innovation is packet spraying. Instead of sending an entire transfer down a single path, it shatters the data into pieces and sends them across hundreds of paths simultaneously. At the destination, they're reassembled by memory address.
Link failures are bypassed in microseconds — no need for switches to recalculate routing tables. OpenAI also ripped out BGP entirely, replacing it with SRv6 source routing: the sender specifies exactly which path each packet takes, turning switches into dumb forwarding boxes. The failure domain shrinks dramatically.
The network architecture simplifies as a result. Connecting a hundred thousand GPUs used to require three to four layers of switches. MRC's multi-plane design cuts that to two layers, reducing power, cost, and points of failure.
The protocol, which connects 100,000 GPUs with just two switch layers, is already deployed across all of OpenAI's largest Nvidia GB200 supercomputers, including the Stargate site in Abilene, Texas (built with Oracle) and Microsoft's Fairwater datacenter. Multiple OpenAI models have been trained using it.
The most concrete example: during a recent frontier model training run (the one powering ChatGPT and Codex), the team hot-swapped four core switches without needing to coordinate with the training team. Multiple link flaps occurred every minute — with no measurable impact on the training job. Previously, that kind of incident would have crashed the entire task.