The ARC Prize Foundation unveils ARC-AGI-3, an AI benchmark with simple video-game-like scenarios designed to measure on-the-fly reasoning, not memory recall
The influential AI researcher François Chollet has long argued that the field measures intelligence incorrectly …
Fast CompanyMark Sullivan
Context & Ripple Effects
ARC-AGI-3 extends the foundation’s effort to separate adaptable problem-solving from pattern recall. Its ARC-AGI-2 predecessor stumped most tested models despite substantially stronger reported human performance, while an earlier high-compute ARC-AGI-1 result illustrated how benchmark scores can depend on evaluation setup and available compute.
The new game-like format joins a broader push toward harder evaluations, alongside newer tests designed to challenge rapidly improving models. It matters because it shifts attention from static-answer performance toward whether a system can infer rules and adapt during a task.
First-order effects
AI developers and evaluators gain a new reasoning-focused test whose scenarios are intended to reward adapting to unfamiliar rules rather than retrieving learned patterns.
ARC Prize Foundation raises the bar for claims that a model’s benchmark performance reflects generalizable, on-the-fly reasoning.
Second-order effects
Model builders may face pressure to report performance across more interactive and contamination-resistant evaluations, not just established static benchmarks.
Buyers and researchers assessing AI agents get a more relevant—but still narrow—signal for tasks that require learning a workflow or responding to changing conditions.
Third-order effects
If interactive reasoning tests become influential, AI evaluation could move from leaderboard optimization toward evidence of reliable adaptation under novel conditions.
The field may increasingly distinguish high scores produced through training-data coverage or heavy inference-time computation from capabilities that transfer to new tasks; whether ARC-AGI-3 establishes that distinction will depend on broad, reproducible use.
The trend: AI benchmarks are evolving from tests of answer accuracy toward evaluations of adaptive reasoning and operational behavior in unfamiliar environments.
ARC-AGI-3 scores agents on how close they are to human action efficiency. All ARC-AGI-3 environments were solved by at least 2 human testers out of 10 (most of the time it was 5+). We use the action count of the 2nd best tester (to avoid outlier performance) as our human
We created an in-house game studio and built 135 novel environments from scratch No instructions, Core Knowledge Priors-only In order to beat these games, AI must: • Explore the environment • Form hypotheses • Execute a plan • Learn and adapt [image]
Keep in mind: ARC-AGI is *not* a final exam that you pass to claim AGI. Including ARC-AGI-3. The benchmarks target the residual gap between what's hard for AI and what's easy for humans. It's meant to be a tool to measure AGI progress and to drive researchers towards the most
ARC-AGI-3 took me a few tries, but it is definitely human winnable. I am curious how much of the very initially very low performance of frontier models is harness, vision, and tools, versus how much are limitations of LLMs. I guess we will find out! https://arcprize.org/...
if we saturate arc 3 this year and there's no meaningful shift in the economy, it's clear benchmarks have become a gimmick between labs and providers to hill climb, market, make bread. whilst it's exciting to see a benchmark the models perform so poorly at. and i love this
ARC-AGI-3 and ARC Prize 2026 are now live with $2,000,000 in prizes! As of today, version 3 is the world's only unsaturated agentic intelligence benchmark. Humans score 100% and frontier AI scores ~0%. Play here: https://arcprize.org/... While no single version of ARC is
ARC-AGI-3 benchmark: - 100% solvable by humans - 1% solvable by AI Everybody keep building benchmarks that agents utterly fail at! Proud this was a Laude Slingshot; will fund other benchmarks that reset SotA to 1%: https://www.laude.org/...
ARC-AGI-3 is an important benchmark. However, I have a major issue with the “Human score 100%” statement. How many humans have tested all 1000 puzzles? How were people selected? This was not published for previous ARCs either. In one case, the human score was based on I think 2
At the moment, ARC-AGI-3 is the only unsaturated agentic AI benchmark. Sub-1% scores from frontier models on the private test set. If you want to be among the first to know when an AGI breakthrough happens, monitor the ARC-AGI-3 leaderboard. Any sudden score jump will mean som…
Announcing ARC-AGI-3 The only unsaturated agentic intelligence benchmark in the world Humans score 100%, AI <1% This human-AI gap demonstrates we do not yet have AGI Most benchmarks test what models already know, ARC-AGI-3 tests how they learn [image]
The 25 public ARC-AGI-3 games On average, they are easier for humans and AI However, the difficulty ranges, there are very easy games and games which are more difficult Easy for AI: https://arcprize.org/... Hard for AI: https://arcprize.org/...
It's alive! This 3rd version of ARC-AGI represents an incredible amount of work from the ARC Prize team. Hundreds of games. Thousands of levels. Go build agents!
Today we're launching ARC-AGI-3 135 Novel Environments (nearly 1K levels) we build by hand It is the only unsaturated agent benchmark in the world Each game is 100% human solvable, AI scores <1% This gap between human and AI performance proves we do not have AGI Agents today [vid…
ARC-AGI-3 is out now! We've designed the benchmark to evaluate agentic intelligence via interactive reasoning environments. Beating ARC-AGI-3 will be achieved when an AI system matches or exceeds human-level action efficiency on all environments, upon seeing them for the first …