/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

GPT-6 Astra scores 62.7% on ARC-AGI-3 with the standard harness and 99.9% with a new provider adapter harness; Claude Opus 5 scored 30.2%, and GPT-5.6 Sol 7.8%

Summary  — GPT-6 Astra scores 62.7% for $26K on ARC-AGI-3 Semi-Private with our Standard harness, and 99.9% for $19K with a Provider Adapter harness.

ARC Prize Greg Kamradt

Context & Ripple Effects

ARC-AGI’s difficulty was established in 2025, when GPT-4.5 and Claude 3.7 Sonnet scored about 1% on ARC-AGI-2 while humans reached 60%. The intervening record also showed that evaluation scaffolding matters: OpenAI said its Responses API harness tripled GPT-5.6 Sol’s ARC-AGI-3 score while reducing output tokens.

Astra’s result turns that harness issue into the central comparison. ARC Prize tested both a standard setup, where the model manages its own carried-forward notes, and a provider adapter that preserves opaque reasoning; the score gap means the benchmark is measuring a combined model-and-integration system as well as the base model.

First-order effects

  • GPT-6 Astra leads Claude Opus 5 and GPT-5.6 Sol on ARC-AGI-3’s standard harness, giving OpenAI a materially stronger result on the benchmark’s directly comparable configuration.
  • The provider adapter produces a near-saturated result at a lower reported cost than the standard run, but establishes a separate evaluation condition that cannot be read as a like-for-like model score.

Second-order effects

  • Anthropic and other frontier-model providers face pressure to publish harness-specific ARC-AGI results, since provider-managed state handling can alter both measured capability and task cost.
  • Enterprise buyers evaluating Astra through the API must distinguish the model’s standard-harness performance from the performance of an OpenAI-supplied integration layer, particularly for agent workflows that depend on persistent task state.

Third-order effects

  • If provider adapters become common, frontier benchmarks will increasingly evaluate complete agent stacks—model, memory, tooling, and orchestration—rather than treating the base model as the sole unit of comparison.
  • That shift makes reproducible, independently specified harnesses more important: a single headline score will be less useful unless the execution environment and cost are reported alongside it.

The trend: AI evaluation is moving from model-only scores toward harness-adjusted measures of what a provider’s full agent system can accomplish per useful task.

Discussion

  • @openai @openai on x
    GPT-6 Astra is state-of-the-art on FrontierMath Tier 4, ARC-AGI 3, and TerminalBench-4.0. GPT-6 Astra is also a major advance for scientific discovery, with state-of-the-art performance on Terminal-Bench Science 0.1 and HealthBench Pro.
  • @tobi Tobi Lutke on x
    Singularity
  • @fchollet François Chollet on x
    When we released ARC 3, I got asked, “when do you think a frontier model will saturate it?”, and I answered “in about a year, though it depends on how much it gets explicitly targeted
  • @arcprize @arcprize on x
    GPT-6 Astra by @OpenAI achieves SOTA on ARC-AGI: - Astra scores 63% on ARC-AGI-3, 99% via a new provider adapter harness - It surpasses human performance on 96% of ARC-AGI-3 levels - It builds the most precise symbolic model of novel environments we've seen Our analysis:
  • @arcprize @arcprize on x
    For ARC-AGI-3 we tested GPT-6 Astra with both our Standard harness (the model decides which notes to carry forward) as well as a new provider adapter harness (which preserves opaque reasoning between requests and uses compaction). Going forward we will test all new models on ARC-…
  • @willdepue Will Depue on x
    it's over boys. i think we're out of ARCs to saturate
  • @fchollet François Chollet on x
    Many of you will ask, “if it saturates ARC 3, is it AGI?
  • @gregkamradt Greg Kamradt on x
    Reflections on Astra from a benchmark perspective: 1. No harness was used in the making of these scores Up until this point we've seen multiple groups report results above 90% on ARC-AGI-3. Examples include PRO-LONG, Tycho, Schema, Prime Agent, see community leaderboard [1] for t…
  • @kimmonismus @kimmonismus on x
    You can quote me on this: OpenAI crushed Anthropic's IPO. They put the newly released, state-of-the-art Fable 5.1 to shame. And they knew exactly what they were doing. Just look at the numbers. Its not even close.
  • @arcprize @arcprize on x
    Astra creates a dense compact symbolic world model to complete ARC-AGI-3 environments. For example, in environment s5i5, Astra: - Recorded the current level, hub orientation, and mechanism lengths: “L8: hub q2 (8↓). Lengths: 14=1...” - It mapped operations to exact controls:
  • @haider1 Haider on x
    goddamn... Astra absolutely nuked ARC-AGI-3 GPT-5.6 Sol (Max): 8% GPT-6 Astra (Max): 63% Astra + provider adapter: 98.6% even ignoring the 98.6% harness score, the jump from 8% to 63% on the standard setup is insane
  • @arcprize @arcprize on x
    Before launching ARC-AGI-3, we tested approximately 500 members of the general public to establish a human baseline for action efficiency, or simply, how quickly did people solve each environment? This gives us an efficiency metric to compare AI to humans. With the provider adapt…
  • @fchollet François Chollet on x
    Benchmarking AI systems is a continual process that co-evolves with the models. New benchmarks challenge AI capabilities with emerging questions to shape the directions and feedback signal of the research process. Then they adapt as models progress, targeting the residual between…
  • @openai @openai on x
    Astra is our most aligned model, with substantial improvements in understanding user intent.
  • @mikeknoop Mike Knoop on x
    GPT-6 Astra is the new SOTA on ARC-AGI-3 It's a qualitatively large leap towards AGI and the pace of progress is frankly surprising. That said, we lack evidence to call this AGI yet. While we are still studying the human capability gaps, we believe open-ended invention is unsolve…
  • @arcprize @arcprize on x
    In addition to achieving state-of-the-art scores on ARC-AGI-3, GPT-6 Astra achieved a record 95.0% on ARC-AGI-2 at $1.12/task, and tied Fable 5's high score of 98.5% at $0.28/task. Full results: https://arcprize.org/...
  • @fchollet François Chollet on x
    GPT-6 Astra represents a step-function change in model capability for interactive reasoning problems. It scores 66% on ARC-AGI-3 using our standard harness, and nearly 100% with a continuous conversation harness and custom compaction, at a cost of roughly $360 per game. In fact, …
  • @gdb Greg Brockman on x
    Astra is useful for so many things, but I'm particularly excited to see how it transforms areas like entrepreneurship, scientific discovery, and how small teams tackle big problems.
  • @jonkeegan.com Jon Keegan on bluesky
    Not a typo!  Lotta caveats here it seems, but even the 62.7 score is an insane leap forward.  ARC-AGI-3 is uniquely difficult. arcprize.org/blog/astra [embedded post]
  • @davidcrespo @davidcrespo on bluesky
    most interesting gpt-6 post so far.  it saturates ARC-AGI 3, which is something given that the previous best score was 30%. “we found it performing highly efficient, on-the-fly symbolic world modeling for each game and level.”  —  x.com/fchollet/sta...  arcprize.org/blog/astra [i…
  • r/singularity r on reddit
    Not only does Astra saturate ARC-AGI-3, it does so using fewer moves than the average human
  • r/codex r on reddit
    GPT-6 Astra Benchmarks
  • @haydenfield Hayden Field on bluesky
    “[If] we look back and say, ‘When was it, really, that AGI was created?’  I think it's going to be about this time, and I think it might be about this model,” Greg Brockman said today about GPT-6 Astra, adding, “For me personally, I do think we're there.” www.theverge.com/ai-arti…