A look at the more challenging AI evaluations emerging in response to the rapid progress of models, including FrontierMath, Humanity's Last Exam, and RE-Bench
Despite their expertise, AI developers don't always know what their most advanced systems are capable of—at least, not at first. X: @tharin_p and @tharin_p X: @tharin_p : My latest piece for @TIME contextualises o3's benchmark results with a look at the new wave of evals shaping the field, including @EpochAIResearch's FrontierMath, @METR_Evals's RE-Bench, and @scale_AI / @ai_risks's Humanity's Last Exam: @tharin_p : Merry Christmas everyone here is 2500 words on AI evals — more interesting than it sounds!
Context & Ripple Effects
Recent results such as o3's 87.5% score on ARC Prize's semi-private evaluation have underscored how quickly established tests can lose discriminatory power. Coverage also showed major labs building their own internal benchmarks as public measures approach ceiling performance.
FrontierMath, RE-Bench, and Humanity's Last Exam form part of the response: evaluations designed to reveal capabilities and high-end behavior that developers may not identify immediately from routine testing.
First-order effects
- Model developers gain harder external tests for probing advanced systems, while evaluation groups become more central sources of evidence about capability limits and behavior.
- Benchmark results become less easily summarized by legacy test scores, particularly for systems such as o3 that are presented as reasoning-oriented models.
Second-order effects
- Labs that rely on internal benchmarks face pressure to demonstrate performance on independently developed evaluations, rather than treating proprietary tests as sufficient evidence.
- More demanding evaluations raise the cost and complexity of model assessment, shifting attention from headline accuracy rates toward the conditions, task types, and compute configurations behind results.
Third-order effects
- If capability gains continue to outpace existing tests, evaluation will become an ongoing measurement discipline rather than a one-time pre-release comparison—an institutional role shared by labs and specialist evaluators.
- The gap between what developers initially observe and what models can do strengthens the case for operational assurance, though the corpus does not establish a common standard for it yet.
The trend: Frontier AI is moving from broad benchmark competition toward continuous, adversarial evaluation of increasingly capable reasoning systems.