A close look at DeepSeek, which is estimated to have access to ~50K Hopper GPUs, a total server capex of ~$1.3B, and a GPU spend of $500M+ over its history
The DeepSeek Narrative Takes the World by Storm — DeepSeek took the world by storm. For the last week, DeepSeek has been the only topic …
SemiAnalysisDylan Patel
Context & Ripple Effects
DeepSeek’s viral moment followed coverage arguing that its commodity-hardware and open-source approach challenged assumptions behind AI hyperscaling. A contemporaneous estimate that it had spent well over $500M on GPUs had already shifted attention from model outputs to the infrastructure supporting them.
The new estimates put a more concrete scale on that infrastructure: DeepSeek may have substantial access to Hopper-class compute and a server buildout requiring major capital, even if its architecture differs from the conventional hyperscaler template.
First-order effects
The estimates recast DeepSeek from a seemingly low-cost model story into one also dependent on significant GPU access and server capital, sharpening scrutiny of its underlying compute position.
Nvidia and other GPU-market observers gain a clearer reason to distinguish between a challenge to model-development economics and an immediate reduction in accelerator demand.
Second-order effects
AI developers and investors will face greater pressure to separate efficiency gains from total infrastructure requirements when comparing frontier-model competitors.
Suppliers across the GPU, server, and memory chain may see demand narratives become more nuanced: efficient software can alter compute utilization without eliminating the need for large installed capacity.
Third-order effects
If similar deployments pair lower-cost model techniques with large GPU fleets, frontier AI competition may shift from raw training spend toward the ability to extract more output from constrained compute.
The episode reinforces a broader uncertainty in AI capex: efficiency advances can redistribute spending across the stack rather than produce a simple, proportional decline in hardware demand.
The trend: AI model efficiency is becoming a competitive lever, but its effect on infrastructure spending depends on how much compute developers can still assemble and utilize.
Interesting analysis of DeepSeek's use of NVIDIA GPUs. Good morning. Last night I was hanging out with a bunch of AI people and in real world tests R1 isn't as good as OpenAI's models, and that's before OpenAI's release of its O3 models, that are coming today, according to
A great chart by SemiAnalysis shows the spike in demand (price) for $NVDA H100 on AWS after DeepSeek. This is something that I also discussed in my DeepSeek article: With the reduction of LLMs, you will have more usage, which benefits the whole ecosystem. Good for $AMZN, $MSFT, […
The only people who didn't know the $6m was false were those who didn't know anything about the semiconductor industry. We worked on a handful of models with the sell side, @dylan522p was very precise, but everyone's model was drastically higher than $6m.
“at best describes the cost of the final training run only”. That's literally what they say. In the original paper. With letters, words, intelligible signs.
New report by leading semiconductor analyst Dylan Patel shows that DeepSeek spent over $1 billion on its compute cluster. The widely reported $6M number is highly misleading, as it excludes capex and R&D, and at best describes the cost of the final training run only.
This key chart from @SemiAnalysis_ appears to have been the key source for claims of “50,000 Hoppers” and more detailed disclosure on their CapEx buildup analysis ("$1.3B"). But the table has errors/inconsistencies. More significantly, key assumptions don't pass sanity checks. [i…
Excellent analysis of DeepSeek by @dylan522p, building on Anthropic CEO's post. My key takeaway is how Reinforcement Learning (AI teaching itself to reason) plays a critical role - and the widely quoted $6m excluded that compute. (link below) [image]
We did a lot of content regarding DeepSeek for clients this week but eventually had to do a newsletter post when my brother texted me about deep seal [image]
“Deepseeks ... main headline being the “$6M” training cost ... is wrong... ” “We are confident their hardware spend is well higher than $500M” Deepeek is not a tiny lab: https://semianalysis.com/...
One of my favourites has done it again. The best all article (and take!) on DeepSeek. Well done https://semianalysis.com/... (Only thing missing is the work they've done on data movement close to the metal and avoiding CUDA which fits very nicely to @tenstorrent's approach) 😉