/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

A skeptical look at the new AI scaling “laws”, including post-train duration and “inference time compute”, and why they may fail to predict AI model performance

Scaling laws ain't what they used to be  —  “If anything we are seeing the emergence of a new scaling law …

Marcus on AI Gary Marcus

Context & Ripple Effects

The article challenges the effort to extend the old pre-training scaling framework to post-training duration and inference-time compute. Related coverage mapped those newer regimes as distinct scaling trends, while AI labs were already emphasizing inference improvements alongside pre-training.

The debate matters because test-time compute was increasingly treated as a route to stronger benchmark results, including OpenAI’s o3 benchmark showing for test-time compute. This piece argues that such inputs may not be dependable stand-ins for real model performance.

First-order effects

  • Claims that post-training time or inference-time compute can predict capability face a sharper burden of validation; researchers and model buyers must distinguish benchmark gains from broadly reliable performance.
  • The article weakens the case for treating a single compute-related metric as a planning proxy when comparing model-development approaches.

Second-order effects

  • Model providers pursuing inference-heavy systems may need to show where extra reasoning or search improves outcomes, rather than relying on the existence of a new scaling curve; later work on inference-time search illustrates the contested approach.
  • Infrastructure decisions become harder to justify solely through projected capability gains, since more inference compute can raise operating demands without a guaranteed performance relationship.

Third-order effects

  • If these proposed laws remain task- and method-dependent, AI progress may be governed less by one general scaling rule and more by combinations of training, evaluation, and inference design.
  • That would shift competition toward proving useful performance per unit of inference cost, rather than simply maximizing training or test-time compute.

The trend: AI development is moving from a comparatively simple pre-training scaling narrative toward a contested, economically consequential mix of post-training and inference-time optimization.