/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

Analysis: from 2019 to 2025, gains in pretraining compute efficiency came mostly from data improvements rather than model improvements

Breaking down 6 years of pretraining progress into data vs model improvements  —  How much of the rapid progress in AI that we've seen over the last few years 1 …

Dwarkesh Podcast

Context & Ripple Effects

A 2020 assessment found efficiency-focused neural-network algorithms showed little progress over the preceding decade, while 2024 coverage mapped pretraining scaling alongside emerging post-training and inference-time approaches. This analysis supplies a concrete explanation for that imbalance: progress in getting more from pretraining compute has been led by data work rather than architecture work.

That matters as AI companies face constraints on conventional web training data and may move toward specialized models, a pressure outlined in coverage of shrinking conventional training-data supplies. Nvidia's 2025 emphasis on pre-training, post-training, and inference-time scaling also shows that the optimization agenda is already broadening beyond a single pretraining lever.

First-order effects

  • Model providers seeking greater capability per unit of pretraining compute have stronger evidence to prioritize data selection, quality, and improvement work over model-only efficiency changes.
  • Teams evaluating pretraining progress must separate data-driven gains from architectural gains rather than crediting all efficiency improvement to larger or altered models.

Second-order effects

  • Competition for differentiated training data and the systems used to improve it intensifies as conventional web datasets become less sufficient for frontier pretraining.
  • Hardware and infrastructure planning shifts toward a mixed scaling strategy: the pre-training, post-training, and inference-time scaling agenda matters more when model-side pretraining gains are comparatively limited.

Third-order effects

  • If this pattern persists, access to high-quality and improvable data becomes a more durable source of advantage in foundation-model development than model architecture alone.
  • The industry’s scaling narrative continues to diversify from pretraining-only scaling toward post-training and inference-time methods, as outlined in earlier coverage of changing scaling laws.

The trend: AI capability work is shifting from a pretraining-compute race toward a broader optimization stack in which data quality, post-training, and inference-time compute carry more weight.

Discussion

  • @dwarkesh_sp Dwarkesh Patel on x
    Dwarkesh Patel (@dwarkesh_sp) on X
  • @sarahookr Sara Hooker on x
    I literally had the debate at a party recently. Pretraining model size scaling is dead because transformers are saturated. Only gains now are data. The group found it to be a very controversial statement 😂
  • @arimorcos Ari Morcos on x
    Since 2019, it's been clear that the biggest gains to be had are from improving data. This motivated me to shift my entire research agenda to improving data quality and led directly to @datologyai. The answer, as always, is not better architectures, but rather better data.
  • @dorialexander Alexander Doria on x
    cool research program reconfirming the practical wisdom: data is capability is model.
  • Rohin Govindarajan Rohin Govindarajan on linkedin
    A very interesting read.  True for the model providers sure... but for the CIOs, CTOs and AI Heads out there, this is irrefutable evidence for Enteprise success as well - …