Analysis: from 2019 to 2025, gains in pretraining compute efficiency came mostly from data improvements rather than model improvements
Breaking down 6 years of pretraining progress into data vs model improvements — How much of the rapid progress in AI that we've seen over the last few years 1 …
Context & Ripple Effects
A 2020 assessment found efficiency-focused neural-network algorithms showed little progress over the preceding decade, while 2024 coverage mapped pretraining scaling alongside emerging post-training and inference-time approaches. This analysis supplies a concrete explanation for that imbalance: progress in getting more from pretraining compute has been led by data work rather than architecture work.
That matters as AI companies face constraints on conventional web training data and may move toward specialized models, a pressure outlined in coverage of shrinking conventional training-data supplies. Nvidia's 2025 emphasis on pre-training, post-training, and inference-time scaling also shows that the optimization agenda is already broadening beyond a single pretraining lever.
First-order effects
- Model providers seeking greater capability per unit of pretraining compute have stronger evidence to prioritize data selection, quality, and improvement work over model-only efficiency changes.
- Teams evaluating pretraining progress must separate data-driven gains from architectural gains rather than crediting all efficiency improvement to larger or altered models.
Second-order effects
- Competition for differentiated training data and the systems used to improve it intensifies as conventional web datasets become less sufficient for frontier pretraining.
- Hardware and infrastructure planning shifts toward a mixed scaling strategy: the pre-training, post-training, and inference-time scaling agenda matters more when model-side pretraining gains are comparatively limited.
Third-order effects
- If this pattern persists, access to high-quality and improvable data becomes a more durable source of advantage in foundation-model development than model architecture alone.
- The industry’s scaling narrative continues to diversify from pretraining-only scaling toward post-training and inference-time methods, as outlined in earlier coverage of changing scaling laws.
The trend: AI capability work is shifting from a pretraining-compute race toward a broader optimization stack in which data quality, post-training, and inference-time compute carry more weight.