DeepSeek, OpenAI, and others building smaller, powerful models quickly and cheaply using “distillation” raises questions about LLMs' first-mover advantage
DeepSeek used technique to create smaller powerful models based on the technology of competitors such as Meta Bluesky: @prietschka . X: @trengriffin . Forums: r/artificial Bluesky: Paul Rietschka / @prietschka : Translation: “um, these companies may not be returning that $3000 kajilllion dollars they were given.” www.ft.com/content/c117... X: Tren Griffin / @trengriffin : I asked the kids running the lemonade stand whether they had a moat that could create some pricing power and whether they had read the FT article on DeepSeek this morning. Surprisingly, they believe superior operating skills will give them sustainable competitive advantage. [image] Forums: r/artificial : One-Minute Daily AI News 3/1/2025
Context & Ripple Effects
DeepSeek’s emergence had already put training economics under scrutiny: its open-source DeepSeek-V3 claims paired competitive performance with lower chip use and cost, while coverage described an approach built around open source and reinforcement learning.
The response was immediate enough that Meta reportedly created internal teams to study DeepSeek’s methods. Distillation adds a distinct pressure point: capable models can be adapted into smaller systems without reproducing every cost of the original model build.
First-order effects
- DeepSeek, OpenAI, and other model builders can use distillation to produce smaller, capable models faster and at lower cost, increasing the supply of models aimed at narrower deployment needs.
- The exclusivity of a leading model’s capabilities becomes harder to sustain when those capabilities can help train lower-cost alternatives, weakening first-mover advantage at the model layer.
Second-order effects
- Frontier labs face greater pressure to differentiate through continued research, product integration, and access to users rather than relying solely on an initial capability lead.
- Buyers gain more leverage as smaller models become credible options for cost- and latency-sensitive workloads, pushing providers to make performance-per-cost a more central competitive measure.
Third-order effects
- If the pattern persists, value in AI may shift further from owning a single frontier model toward distribution, application integration, and efficient inference infrastructure.
- The model market could become more layered: expensive frontier systems generate capabilities at the top, while a broader set of smaller models competes in downstream deployment.
The trend: Distillation is part of a broader shift from a winner-take-all race for the largest model toward competition over efficiently delivering capable models across more use cases.