GPT-6 Astra scores 62.7% on ARC-AGI-3 with the standard harness and 99.9% with a new provider adapter harness; Claude Opus 5 scored 30.2%, and GPT-5.6 Sol 7.8%
Summary — GPT-6 Astra scores 62.7% for $26K on ARC-AGI-3 Semi-Private with our Standard harness, and 99.9% for $19K with a Provider Adapter harness.
Astra's benchmark results arrive alongside OpenAI's claims of stronger computer-use performance, tying abstract reasoning evaluations more closely to the reliability buyers seek from agents that act across software tools.
First-order effects
ARC Prize's standard-harness result gives GPT-6 Astra a clear published lead over Claude Opus 5 and GPT-5.6 Sol on ARC-AGI-3, establishing the standard configuration as the cleanest like-for-like comparison among the reported scores.
The gap between Astra's standard-harness and Provider Adapter results makes the harness itself an immediate determinant of both measured task completion and reported task cost.
Second-order effects
Enterprise evaluators comparing OpenAI and Anthropic will need to test the full model-and-harness stack rather than treat a headline benchmark score as a portable property of the underlying model.
Claude Opus 5's 30.2% standard-harness result becomes a concrete competitive target, while OpenAI has an incentive to make provider-specific integration advantages available in its agent tooling.
Third-order effects
If providers continue to post sharply different results across harnesses, frontier-model benchmarks will increasingly measure operational systems—model, state handling, and tool orchestration—rather than a model in isolation.
Cost-per-task reporting is likely to become a stronger procurement criterion alongside accuracy, because the two Astra configurations pair radically different completion rates with different reported costs.
The trend: Agent evaluation is shifting from single-model leaderboards toward costed, end-to-end reliability tests in which the execution harness is part of the product.
GPT-6 Astra is the new SOTA on ARC-AGI-3 It's a qualitatively large leap towards AGI and the pace of progress is frankly surprising. That said, we lack evidence to call this AGI yet. While we are still studying the human capability gaps, we believe open-ended invention is unsolve…
Reflections on Astra from a benchmark perspective: 1. No harness was used in the making of these scores Up until this point we've seen multiple groups report results above 90% on ARC-AGI-3. Examples include PRO-LONG, Tycho, Schema, Prime Agent, see community leaderboard [1] for t…
For ARC-AGI-3 we tested GPT-6 Astra with both our Standard harness (the model decides which notes to carry forward) as well as a new provider adapter harness (which preserves opaque reasoning between requests and uses compaction). Going forward we will test all new models on ARC-…
In addition to achieving state-of-the-art scores on ARC-AGI-3, GPT-6 Astra achieved a record 95.0% on ARC-AGI-2 at $1.12/task, and tied Fable 5's high score of 98.5% at $0.28/task. Full results: https://arcprize.org/...
Astra creates a dense compact symbolic world model to complete ARC-AGI-3 environments. For example, in environment s5i5, Astra: - Recorded the current level, hub orientation, and mechanism lengths: “L8: hub q2 (8↓). Lengths: 14=1...” - It mapped operations to exact controls:
Before launching ARC-AGI-3, we tested approximately 500 members of the general public to establish a human baseline for action efficiency, or simply, how quickly did people solve each environment? This gives us an efficiency metric to compare AI to humans. With the provider adapt…
GPT-6 Astra by @OpenAI achieves SOTA on ARC-AGI: - Astra scores 63% on ARC-AGI-3, 99% via a new provider adapter harness - It surpasses human performance on 96% of ARC-AGI-3 levels - It builds the most precise symbolic model of novel environments we've seen Our analysis:
Astra is useful for so many things, but I'm particularly excited to see how it transforms areas like entrepreneurship, scientific discovery, and how small teams tackle big problems.
goddamn... Astra absolutely nuked ARC-AGI-3 GPT-5.6 Sol (Max): 8% GPT-6 Astra (Max): 63% Astra + provider adapter: 98.6% even ignoring the 98.6% harness score, the jump from 8% to 63% on the standard setup is insane
When we released ARC 3, I got asked, “when do you think a frontier model will saturate it?”, and I answered “in about a year, though it depends on how much it gets explicitly targeted
Benchmarking AI systems is a continual process that co-evolves with the models. New benchmarks challenge AI capabilities with emerging questions to shape the directions and feedback signal of the research process. Then they adapt as models progress, targeting the residual between…
GPT-6 Astra represents a step-function change in model capability for interactive reasoning problems. It scores 66% on ARC-AGI-3 using our standard harness, and nearly 100% with a continuous conversation harness and custom compaction, at a cost of roughly $360 per game. In fact, …
GPT-6 Astra is state-of-the-art on FrontierMath Tier 4, ARC-AGI 3, and TerminalBench-4.0. GPT-6 Astra is also a major advance for scientific discovery, with state-of-the-art performance on Terminal-Bench Science 0.1 and HealthBench Pro.
You can quote me on this: OpenAI crushed Anthropic's IPO. They put the newly released, state-of-the-art Fable 5.1 to shame. And they knew exactly what they were doing. Just look at the numbers. Its not even close.
most interesting gpt-6 post so far. it saturates ARC-AGI 3, which is something given that the previous best score was 30%. “we found it performing highly efficient, on-the-fly symbolic world modeling for each game and level.” — x.com/fchollet/sta... arcprize.org/blog/astra [i…
Not a typo! Lotta caveats here it seems, but even the 62.7 score is an insane leap forward. ARC-AGI-3 is uniquely difficult. arcprize.org/blog/astra [embedded post]