Meta, OpenAI, Microsoft, and other AI companies create their own internal benchmarks as new models approach or exceed 90% accuracy on existing public tests
Rapidly advancing technology is surpassing current methods of evaluating and comparing large language models
Context & Ripple Effects
Public tests are losing their ability to distinguish leading models as scores converge near the top of the scale. This makes evaluation itself a competitive capability for Meta, OpenAI, Microsoft, and peers rather than a shared external scoreboard.
The story anticipates coverage of harder evaluations designed to restore discrimination among frontier models and later concerns that many benchmarks lack consistent objectives or statistical comparability, as examined in an Oxford Internet Institute benchmark study.
First-order effects
- Meta, OpenAI, Microsoft, and other model developers must rely more heavily on proprietary internal testing to identify improvements once public benchmarks stop separating models clearly.
- Outside users and observers have less access to a common basis for comparing frontier-model claims, since the most decision-useful evaluations can remain internal.
Second-order effects
- Benchmark creators and independent evaluators face pressure to produce harder, better-defined tests; the emerging set of more challenging AI evaluations is a direct response to saturated public measures.
- Model vendors will increasingly need to substantiate performance through task-specific evidence, not just headline benchmark scores, raising the importance of evaluation design in enterprise buying.
Third-order effects
- If proprietary evaluation becomes the norm, model comparison may shift from standardized leaderboard competition toward proof of performance on particular workloads, costs, and deployment conditions.
- The pattern reinforces AI industrialization: evaluation becomes embedded in the model-development and product-delivery stack, though independent tests remain important for cross-vendor accountability.
The trend: Frontier AI competition is moving beyond broad public benchmarks toward proprietary and workload-specific measures of useful performance.