/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

Meta, OpenAI, Microsoft, and other AI companies create their own internal benchmarks as new models approach or exceed 90% accuracy on existing public tests

Rapidly advancing technology is surpassing current methods of evaluating and comparing large language models

Financial Times Cristina Criddle

Context & Ripple Effects

Public tests are losing their ability to distinguish leading models as scores converge near the top of the scale. This makes evaluation itself a competitive capability for Meta, OpenAI, Microsoft, and peers rather than a shared external scoreboard.

The story anticipates coverage of harder evaluations designed to restore discrimination among frontier models and later concerns that many benchmarks lack consistent objectives or statistical comparability, as examined in an Oxford Internet Institute benchmark study.

First-order effects

  • Meta, OpenAI, Microsoft, and other model developers must rely more heavily on proprietary internal testing to identify improvements once public benchmarks stop separating models clearly.
  • Outside users and observers have less access to a common basis for comparing frontier-model claims, since the most decision-useful evaluations can remain internal.

Second-order effects

  • Benchmark creators and independent evaluators face pressure to produce harder, better-defined tests; the emerging set of more challenging AI evaluations is a direct response to saturated public measures.
  • Model vendors will increasingly need to substantiate performance through task-specific evidence, not just headline benchmark scores, raising the importance of evaluation design in enterprise buying.

Third-order effects

  • If proprietary evaluation becomes the norm, model comparison may shift from standardized leaderboard competition toward proof of performance on particular workloads, costs, and deployment conditions.
  • The pattern reinforces AI industrialization: evaluation becomes embedded in the model-development and product-delivery stack, though independent tests remain important for cross-vendor accountability.

The trend: Frontier AI competition is moving beyond broad public benchmarks toward proprietary and workload-specific measures of useful performance.

Discussion

  • @spirosmargaris Spiros Margaris on x
    AI groups rush to redesign model testing and create new benchmarks https://www.ft.com/... @CristinaCriddle @FT