/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

← → days · ↑ ↓ browse · Enter similar · o open

Nvidia says its general-purpose coding agent system AVO scored 100% across all 25 environments in the ARC-AGI-3 public set, completing all 183 levels

NVIDIA Technical Blog Terry Chen

Context & Ripple Effects

When the ARC Prize Foundation launched ARC-AGI-3 in March, it was built specifically to test on-the-fly reasoning in unfamiliar video-game-like scenarios rather than memory recall — a benchmark meant to be hard. Six months later, Nvidia reports its general-purpose coding agent system AVO cleared every one of the 25 public environments and all 183 levels.

The claim lands amid a fast-moving agent narrative: Karpathy observed in February that coding agents had leapt forward since December, completing complex projects with minimal oversight. It also fits an old Nvidia pattern — pairing flagship compute platforms with headline results, from the A100's 20x generational jump to the Pegasus autonomy platform — where benchmark wins double as proof points for the underlying stack.

First-order effects

  • Nvidia now holds a claimed 100% result on the benchmark explicitly designed to resist memorization, giving AVO a marquee credential in the general-purpose agent race and Nvidia a showcase for its own tooling.
  • The ARC Prize Foundation's public set is effectively saturated by this result, shifting attention to whatever held-out or harder evaluation it fields next.

Second-order effects

  • Rival agent builders — coding-assistant vendors and labs chasing the same 'minimal oversight' milestone Karpathy flagged — face pressure to publish comparable ARC-AGI-3 numbers or cede the reasoning-benchmark narrative to Nvidia.
  • Benchmark operators and enterprise buyers alike will discount self-reported vendor results, pushing third-party verification and independent replication into the center of how agent claims get judged.

Third-order effects

  • If general-purpose agents can genuinely solve novel environments rather than recall training patterns, evaluation itself becomes a moving target — benchmarks will need faster rotation, and marketing advantage will accrue to whoever pairs models with the full compute-to-agent stack, as Nvidia has done since the A100 era.

The trend: General-purpose AI agents are moving from narrow task competence toward demonstrated novel-environment reasoning, with chipmakers using benchmark saturation to market integrated stacks.

Discussion

  • @nvidiaai @nvidiaai on x
    Our general-purpose coding agent just scored 100% on the ARC-AGI-3 interactive reasoning benchmark. NVIDIA AVO completed all 183 levels across all 25 public environments, figuring out what to do with no instructions, explicit rules, or stated goals.
  • @fakepsyho @fakepsyho on x
    this is non-news, it doesn't even report the token usage also, blog post is complete slop
  • @deryatr_ Derya Unutmaz on x
    What!? ARC-AGI-3 is cooked too! This is what happens in singularity: intelligence benchmarks fall faster than they can be created! Big congratulations to NVIDIA for the landmark achievement with their coding agents!
  • @alexocheema Alex Cheema on x
    It's the harness, guys! It's the harness!
  • @angaisb_ Angel on x
    ARC-AGI-3 solved Can't wait to watch the ARC-AGI supporters cry about why this doesn't count along with a million other excuses I hope ARC-AGI-4 goes back to being as good as the first ones and is actually useful for something
  • @nvidiaai @nvidiaai on x
    NVIDIA AVO continuously inspects, plans, implements, and evaluates, using memory, tools, and execution feedback to build on what it learns along the way. This allows the system to sustain progress across long-running tasks rather than starting over with each model context. Read
  • @jfpuget @jfpuget on x
    Another thing I am associated with. NVIDIA AVO, our research coding agent, gets 100% on ARC AGI3 with very little input and tooling. It can basically look at past grids, or subsets of these, represented as text. What is interesting is that AVO was not designed at all for ARC
  • @lyalindotcom Dmitry Lyalin on x
    Then all at once.
  • Bing Xu Bing Xu on linkedin
    Finally, the AVO blog is out!  It's been half a year.  Terry Chen and I have built four generations of Self-Improving Agents at NVIDIA since 2024. …
  • @timkellogg.me Mr. Tim on bluesky
    NVIDIA has their own agent harness that saturates ARC-AGI-3, AVO  —  developer.nvidia.com/blog/nvidia- ...  [image]
  • r/accelerate r on reddit
    Nvidia: harness more important than model
  • r/NVDA_Stock r on reddit
    Amazing new Nvidia Coding Agent