A skeptical look at the new AI scaling “laws”, including post-train duration and “inference time compute”, and why they may fail to predict AI model performance
Scaling laws ain't what they used to be — “If anything we are seeing the emergence of a new scaling law …
Context & Ripple Effects
The article challenges the effort to extend the old pre-training scaling framework to post-training duration and inference-time compute. Related coverage mapped those newer regimes as distinct scaling trends, while AI labs were already emphasizing inference improvements alongside pre-training.
The debate matters because test-time compute was increasingly treated as a route to stronger benchmark results, including OpenAI’s o3 benchmark showing for test-time compute. This piece argues that such inputs may not be dependable stand-ins for real model performance.
First-order effects
- Claims that post-training time or inference-time compute can predict capability face a sharper burden of validation; researchers and model buyers must distinguish benchmark gains from broadly reliable performance.
- The article weakens the case for treating a single compute-related metric as a planning proxy when comparing model-development approaches.
Second-order effects
- Model providers pursuing inference-heavy systems may need to show where extra reasoning or search improves outcomes, rather than relying on the existence of a new scaling curve; later work on inference-time search illustrates the contested approach.
- Infrastructure decisions become harder to justify solely through projected capability gains, since more inference compute can raise operating demands without a guaranteed performance relationship.
Third-order effects
- If these proposed laws remain task- and method-dependent, AI progress may be governed less by one general scaling rule and more by combinations of training, evaluation, and inference design.
- That would shift competition toward proving useful performance per unit of inference cost, rather than simply maximizing training or test-time compute.
The trend: AI development is moving from a comparatively simple pre-training scaling narrative toward a contested, economically consequential mix of post-training and inference-time optimization.