/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

← → days · ↑ ↓ browse · Enter similar · o open

GPT-6 Astra scores 62.7% on ARC-AGI-3 with the standard harness and 99.9% with a new provider adapter harness; Claude Opus 5 scored 30.2%, and GPT-5.6 Sol 7.8%

Summary  — GPT-6 Astra scores 62.7% for $26K on ARC-AGI-3 Semi-Private with our Standard harness, and 99.9% for $19K with a Provider Adapter harness.

ARC Prize Greg Kamradt

Context & Ripple Effects

ARC Prize's earlier benchmark cycle set a low base: most leading models scored about 1% on ARC-AGI-2, versus a 60% human result, in the test's initial published comparison. By July 2026, OpenAI had already shown that its Responses API harness materially lifted GPT-5.6 Sol's ARC-AGI-3 result, making evaluation setup a central part of the performance story.

Astra's benchmark results arrive alongside OpenAI's claims of stronger computer-use performance, tying abstract reasoning evaluations more closely to the reliability buyers seek from agents that act across software tools.

First-order effects

  • ARC Prize's standard-harness result gives GPT-6 Astra a clear published lead over Claude Opus 5 and GPT-5.6 Sol on ARC-AGI-3, establishing the standard configuration as the cleanest like-for-like comparison among the reported scores.
  • The gap between Astra's standard-harness and Provider Adapter results makes the harness itself an immediate determinant of both measured task completion and reported task cost.

Second-order effects

  • Enterprise evaluators comparing OpenAI and Anthropic will need to test the full model-and-harness stack rather than treat a headline benchmark score as a portable property of the underlying model.
  • Claude Opus 5's 30.2% standard-harness result becomes a concrete competitive target, while OpenAI has an incentive to make provider-specific integration advantages available in its agent tooling.

Third-order effects

  • If providers continue to post sharply different results across harnesses, frontier-model benchmarks will increasingly measure operational systems—model, state handling, and tool orchestration—rather than a model in isolation.
  • Cost-per-task reporting is likely to become a stronger procurement criterion alongside accuracy, because the two Astra configurations pair radically different completion rates with different reported costs.

The trend: Agent evaluation is shifting from single-model leaderboards toward costed, end-to-end reliability tests in which the execution harness is part of the product.

Discussion

  • @mikeknoop Mike Knoop on x
    GPT-6 Astra is the new SOTA on ARC-AGI-3 It's a qualitatively large leap towards AGI and the pace of progress is frankly surprising. That said, we lack evidence to call this AGI yet. While we are still studying the human capability gaps, we believe open-ended invention is unsolve…
  • @gregkamradt Greg Kamradt on x
    Reflections on Astra from a benchmark perspective: 1. No harness was used in the making of these scores Up until this point we've seen multiple groups report results above 90% on ARC-AGI-3. Examples include PRO-LONG, Tycho, Schema, Prime Agent, see community leaderboard [1] for t…
  • @arcprize @arcprize on x
    For ARC-AGI-3 we tested GPT-6 Astra with both our Standard harness (the model decides which notes to carry forward) as well as a new provider adapter harness (which preserves opaque reasoning between requests and uses compaction). Going forward we will test all new models on ARC-…
  • @arcprize @arcprize on x
    In addition to achieving state-of-the-art scores on ARC-AGI-3, GPT-6 Astra achieved a record 95.0% on ARC-AGI-2 at $1.12/task, and tied Fable 5's high score of 98.5% at $0.28/task. Full results: https://arcprize.org/...
  • @arcprize @arcprize on x
    Astra creates a dense compact symbolic world model to complete ARC-AGI-3 environments. For example, in environment s5i5, Astra: - Recorded the current level, hub orientation, and mechanism lengths: “L8: hub q2 (8↓). Lengths: 14=1...” - It mapped operations to exact controls:
  • @arcprize @arcprize on x
    Before launching ARC-AGI-3, we tested approximately 500 members of the general public to establish a human baseline for action efficiency, or simply, how quickly did people solve each environment? This gives us an efficiency metric to compare AI to humans. With the provider adapt…
  • @arcprize @arcprize on x
    GPT-6 Astra by @OpenAI achieves SOTA on ARC-AGI: - Astra scores 63% on ARC-AGI-3, 99% via a new provider adapter harness - It surpasses human performance on 96% of ARC-AGI-3 levels - It builds the most precise symbolic model of novel environments we've seen Our analysis:
  • @gdb Greg Brockman on x
    Astra is useful for so many things, but I'm particularly excited to see how it transforms areas like entrepreneurship, scientific discovery, and how small teams tackle big problems.
  • @haider1 Haider on x
    goddamn... Astra absolutely nuked ARC-AGI-3 GPT-5.6 Sol (Max): 8% GPT-6 Astra (Max): 63% Astra + provider adapter: 98.6% even ignoring the 98.6% harness score, the jump from 8% to 63% on the standard setup is insane
  • @fchollet François Chollet on x
    When we released ARC 3, I got asked, “when do you think a frontier model will saturate it?”, and I answered “in about a year, though it depends on how much it gets explicitly targeted
  • @fchollet François Chollet on x
    Benchmarking AI systems is a continual process that co-evolves with the models. New benchmarks challenge AI capabilities with emerging questions to shape the directions and feedback signal of the research process. Then they adapt as models progress, targeting the residual between…
  • @tobi Tobi Lutke on x
    Singularity
  • @fchollet François Chollet on x
    Many of you will ask, “if it saturates ARC 3, is it AGI?
  • @willdepue Will Depue on x
    it's over boys. i think we're out of ARCs to saturate
  • @fchollet François Chollet on x
    GPT-6 Astra represents a step-function change in model capability for interactive reasoning problems. It scores 66% on ARC-AGI-3 using our standard harness, and nearly 100% with a continuous conversation harness and custom compaction, at a cost of roughly $360 per game. In fact, …
  • @openai @openai on x
    GPT-6 Astra is state-of-the-art on FrontierMath Tier 4, ARC-AGI 3, and TerminalBench-4.0. GPT-6 Astra is also a major advance for scientific discovery, with state-of-the-art performance on Terminal-Bench Science 0.1 and HealthBench Pro.
  • @openai @openai on x
    Astra is our most aligned model, with substantial improvements in understanding user intent.
  • @kimmonismus @kimmonismus on x
    You can quote me on this: OpenAI crushed Anthropic's IPO. They put the newly released, state-of-the-art Fable 5.1 to shame. And they knew exactly what they were doing. Just look at the numbers. Its not even close.
  • r/singularity r on reddit
    Not only does Astra saturate ARC-AGI-3, it does so using fewer moves than the average human
  • @davidcrespo @davidcrespo on bluesky
    most interesting gpt-6 post so far.  it saturates ARC-AGI 3, which is something given that the previous best score was 30%. “we found it performing highly efficient, on-the-fly symbolic world modeling for each game and level.”  —  x.com/fchollet/sta...  arcprize.org/blog/astra [i…
  • @jonkeegan.com Jon Keegan on bluesky
    Not a typo!  Lotta caveats here it seems, but even the 62.7 score is an insane leap forward.  ARC-AGI-3 is uniquely difficult. arcprize.org/blog/astra [embedded post]