OpenAI partners with Microsoft, AMD, Broadcom, Nvidia, and Intel researchers to detail the Multipath Reliable Connection (MRC) protocol to help scale compute
The Deep ViewNat Rubio-Licht
Context & Ripple Effects
OpenAI’s compute relationship with Microsoft has long centered on Azure-hosted infrastructure, including a dedicated large-scale supercomputer for distributed models. More recent coverage also places Microsoft’s OpenAI partnership alongside broader data-center investment and work on proprietary AI infrastructure.
The MRC work brings OpenAI, Microsoft, AMD, Broadcom, Nvidia, and Intel into a shared networking-focused effort. It follows adjacent attempts to reduce software and infrastructure dependence on any single AI-chip platform, including work around Triton.
First-order effects
The participating companies gain a common, documented protocol intended to make connections across large-scale compute systems more reliable and scalable, directly affecting how their AI infrastructure teams can design and operate distributed clusters.
OpenAI and Microsoft can apply the work to the infrastructure supporting large distributed models, while AMD, Broadcom, Nvidia, and Intel are positioned to align relevant hardware and networking implementations with the protocol.
Second-order effects
A shared approach can shift competition among AI-infrastructure suppliers toward implementation quality, performance, and integration rather than forcing customers into wholly proprietary connection stacks.
Cloud and data-center operators building mixed-vendor AI clusters could face lower integration friction if the protocol is broadly adopted, increasing pressure on suppliers to support interoperable networking paths.
Third-order effects
If adoption extends beyond the initial researchers, AI compute may evolve toward more standardized, composable cluster infrastructure, making it easier to combine processors, networking gear, and cloud capacity from multiple vendors.
The effort also underscores that scaling AI is becoming a systems problem—not solely a chip problem—though its structural impact depends on whether MRC becomes a broadly implemented standard rather than a limited collaboration.
The trend: AI infrastructure is moving toward cross-vendor standards for operating ever-larger distributed compute clusters, alongside continued competition in chips, networking, and cloud capacity.
Today we shared MRC ( https://openai.com/...), a networking protocol developed with @Microsoft, @nvidia, @AMD, @Broadcom, and @intel to improve how large AI training systems move data and recover from failures. This innovation has come full circle for me personally, it was
At AI scale, raw bandwidth breaks down. What matters is resilience, recovery, and consistency under load. AMD in collaboration with Microsoft and @OpenAI defines a proven solution approach to AI networking with MRC. Learn more now: https://www.amd.com/... [image]
We've partnered with @AMD, @Broadcom, @Intel, @Microsoft, and @NVIDIA, to release Multipath Reliable Connection (MRC), a new open networking protocol that helps large AI training clusters run faster and more reliably, with less wasted GPU time. https://openai.com/...
MRC is already deployed across all of OpenAI's largest supercomputers that we use to train frontier models, including our site with @Oracle Cloud Infrastructure (OCI) in Abilene, Texas, and in @Microsoft's Fairwater supercomputers. MRC is now available through the [video]
Gigascale AI needs networking built for scale, resilience and openness. NVIDIA Spectrum-X Ethernet now supports MRC, a new RDMA-based protocol that improves throughput, availability and failure recovery for large-scale AI training. Used by @OpenAI, @Microsoft, and @Oracle. Now [i…
Exclusive from me this morning: OpenAI and some of the industry's biggest names have come together to fix two of the biggest issues in networking: Congestion and failure. …
Today we shared MRC (https://lnkd.in/... a networking protocol developed with Microsoft, NVIDIA, AMD, Broadcom, and Intel to improve how large AI training systems move data and recover from failures. …