OpenAI, working with AMD, Broadcom, Intel, Microsoft, and Nvidia, has developed a new networking protocol called MRC (Multipath Reliable Connection). It's already baked into the latest 800Gb/s NICs and is now open-sourced through the Open Compute Project (OCP). MRC is deployed across all of OpenAI's largest Nvidia GB200 supercomputers, including the Abilene cluster in Texas built with Oracle and Microsoft's Fairwater supercomputer, where it's being used to train cutting-edge models.
In traditional supercomputing networks, a single training step involves millions of data transfers, and even one delayed transfer can leave GPUs idling. MRC's core innovation is packet spraying: it breaks each transfer into chunks and sends them across hundreds of independent network paths simultaneously. In conventional setups, each transfer follows a fixed route, and when multiple transfers hit the same link, you get a traffic jam. MRC spreads the load evenly, so congestion in the core network virtually disappears.
Architecturally, MRC splits a single 800Gb/s port into eight 100Gb/s connections, each feeding into a separate network plane. That means a 64-port switch can handle 512 endpoints, allowing more than 100,000 GPUs to be fully connected with just two switch layers — traditional designs need three or four. MRC also replaces BGP dynamic routing with SRv6 source routing: the sender writes the entire path into the packet header, and switches just forward along static routes. This eliminates a whole class of failures caused by dynamic routing hiccups.
In production, even with multiple link flaps per minute, MRC has zero impact on training jobs. The team once rebooted four core switches without coordinating with the training crew — jobs recovered automatically. Fault response time drops from seconds (or tens of seconds) in conventional setups to microseconds.