An in-depth look at DeepSeek: DeepSeekMoE and DeepSeekMLA, cheap V3 training, the US chip ban, “distillation” from other models, impact on Nvidia, AGI, and more
It's Monday, January 27. Why haven't you written about DeepSeek yet? — I did! I wrote about R1 last Tuesday.
Context & Ripple Effects
DeepSeek entered the discussion with claims that its open-source V3 could rival US models while using fewer chips and costing $6M to train, as outlined in an earlier profile of its V3 claims. This analysis examines the architectural and sourcing choices behind those efficiency claims, and why they matter under US chip restrictions.
The story also foreshadows a longer contest over whether software efficiency can offset hardware constraints. Later coverage of an effort to reduce HBM use and V3.1’s adaptation for next-generation Chinese-made chips shows that DeepSeek’s optimization work became tied to a broader domestic hardware path.
First-order effects
- DeepSeek’s reported use of MoE, MLA, and low-cost V3 training reframes its model work as a challenge to the assumption that frontier capability necessarily requires ever-larger training budgets.
- Nvidia faces immediate investor and customer scrutiny over whether more efficient model development could temper demand expectations for the most compute-intensive training workloads.
Second-order effects
- Rival model developers are pushed to demonstrate not just benchmark performance but training and inference efficiency; DeepSeek later made a V3 update available under an MIT license, increasing the visibility of that comparison.
- US chip restrictions make architecture-level efficiency strategically important for Chinese AI developers, while also increasing incentives to tailor models to alternative hardware stacks.
Third-order effects
- If efficient architectures continue to close capability gaps, competition may shift from access to the largest training clusters toward the ability to optimize models, memory use, and deployment economics.
- The resulting market would not eliminate the value of advanced accelerators, but it could make AI hardware demand more sensitive to software efficiency and to geographically segmented supply chains.
The trend: DeepSeek is one data point in a shift from brute-force scaling toward efficiency-led AI development shaped by compute constraints and hardware geopolitics.