An in-depth look at GPU supply and demand, particularly Nvidia H100s, how a reliance on TSMC causes bottlenecks, H100 clients, Nvidia's allocations, and more
Introduction — As of August 2023, it seems AI might be bottlenecked by the supply of GPUs. … Is There Really A Bottleneck? Twitter: @nixcraft , @thestalwart , @mhdempsey , @wholemarsblog , and @lpolovets Twitter: @nixcraft : Nvidia H100 GPUs: Supply and Demand https://t.co/... This is truly mind blowing. Obtaining those NVDIA GPUs can be quite challenging, even if they cost $100,000 per unit. Joe Weisenthal / @thestalwart : “...Nvidia could purely raise prices to find a clearing price, and are doing that to some extent. But it's important to know that ultimately H100 allocation comes down to who Nvidia prefers to give that allocation to.” https://t.co/... Michael Dempsey / @mhdempsey : Everything you've ever wanted to know about the current GPU shortage and supply/demand dynamics. https://t.co/... @wholemarsblog : Nvidia H100 GPUs: Supply and Demand “This post is an exploration of the supply and demand of GPUs, particularly Nvidia H100s. We're also releasing a song and music video on the same day as this post” https://t.co/... Leo Polovets / @lpolovets : Excellent deep dive on GPU supply and demand. Covers everything from startups' needs to public cloud capacity to the chemicals required for producing GPUs (!!). https://t.co/... https://t.co/...
Context & Ripple Effects
Generative-AI demand had already made Nvidia’s earlier A100 a critical input, with coverage describing the company’s dominant position in machine-learning GPUs. This account identifies where that scarcity becomes operational: manufacturing capacity, specialized supply-chain inputs, and Nvidia’s allocation choices.
The constraint mattered beyond a single buyer because H100 demand spanned startups and public-cloud providers. Subsequent reporting of large H100 purchases by Saudi Arabia and the UAE underscores how limited supply could be directed toward a small set of well-capitalized customers.
First-order effects
- Startups and cloud providers face constrained access and higher effective costs for H100 capacity, while Nvidia uses pricing and allocation rather than price alone to distribute scarce units.
- TSMC-linked production limits and required chemicals constrain Nvidia’s ability to quickly convert demand into shipments, giving selected allocation recipients near-term access advantages.
Second-order effects
- Cloud providers with allocations can turn scarce accelerators into differentiated AI capacity, while customers without them may delay deployments or seek alternative compute arrangements.
- The shortage raises the value of adjacent inputs—manufacturing capacity, networking, and other components required to deploy GPUs—rather than concentrating pressure solely on the accelerator itself; Nvidia had also been expanding data-center networking and DGX systems.
Third-order effects
- If allocation remains central to access, AI development may be shaped more by capital strength and supply-chain relationships than by model-building capability alone.
- The bottleneck could eventually shift toward alternative accelerators and software portability, but later MI300X-versus-H100 benchmarking suggests hardware specifications alone may not remove switching friction.
The trend: AI infrastructure is becoming a supply-chain-constrained market in which scarce accelerators and their complementary inputs determine who can scale compute first.