Sources: Microsoft, Meta, AWS, and Google recently cut some orders of Nvidia's Blackwell GB200 racks, as overheating and connection glitches lead to new delays
Some of Nvidia's biggest customers are facing new delays in getting its most advanced artificial intelligence chips up and running in data centers.
Context & Ripple Effects
Blackwell’s rollout had already been disrupted by a reported three-month-or-more chip delay tied to design flaws. By November, Nvidia was reportedly asking suppliers for repeated server-rack design changes to address overheating, placing the issue beyond the GPU itself.
The reported order reductions show that deployment readiness—not just chip availability—has become a constraint for Nvidia’s largest cloud customers as they bring advanced AI systems into data centers.
First-order effects
- Microsoft, Meta, AWS and Google reportedly reduce some GB200 rack commitments while they wait for overheating and connection problems to be resolved.
- Nvidia faces another delay in converting Blackwell demand into deployed systems, while affected customers’ planned data-center capacity comes online later than expected.
Second-order effects
- Rack and component suppliers may face changing schedules and further engineering work as Nvidia and customers address system-level thermal and connectivity issues.
- Large cloud operators may need to adjust AI infrastructure plans around the timing of usable Blackwell capacity, increasing the importance of execution certainty alongside accelerator performance.
Third-order effects
- If recurring system-integration issues persist, procurement for frontier AI hardware could put more weight on validated rack-level deployment rather than on GPU specifications alone.
- The episode points to rack engineering and thermal design becoming a strategic bottleneck in AI infrastructure, potentially favoring vendors and operators with tighter control over the full deployment stack.
The trend: AI compute competition is shifting from securing leading chips to reliably deploying complete, power- and thermally constrained systems at data-center scale.