Nvidia says its general-purpose coding agent system AVO scored 100% across all 25 environments in the ARC-AGI-3 public set, completing all 183 levels
Context & Ripple Effects
When the ARC Prize Foundation launched ARC-AGI-3 in March, it was built specifically to test on-the-fly reasoning in unfamiliar video-game-like scenarios rather than memory recall — a benchmark meant to be hard. Six months later, Nvidia reports its general-purpose coding agent system AVO cleared every one of the 25 public environments and all 183 levels.
The claim lands amid a fast-moving agent narrative: Karpathy observed in February that coding agents had leapt forward since December, completing complex projects with minimal oversight. It also fits an old Nvidia pattern — pairing flagship compute platforms with headline results, from the A100's 20x generational jump to the Pegasus autonomy platform — where benchmark wins double as proof points for the underlying stack.
First-order effects
- Nvidia now holds a claimed 100% result on the benchmark explicitly designed to resist memorization, giving AVO a marquee credential in the general-purpose agent race and Nvidia a showcase for its own tooling.
- The ARC Prize Foundation's public set is effectively saturated by this result, shifting attention to whatever held-out or harder evaluation it fields next.
Second-order effects
- Rival agent builders — coding-assistant vendors and labs chasing the same 'minimal oversight' milestone Karpathy flagged — face pressure to publish comparable ARC-AGI-3 numbers or cede the reasoning-benchmark narrative to Nvidia.
- Benchmark operators and enterprise buyers alike will discount self-reported vendor results, pushing third-party verification and independent replication into the center of how agent claims get judged.
Third-order effects
- If general-purpose agents can genuinely solve novel environments rather than recall training patterns, evaluation itself becomes a moving target — benchmarks will need faster rotation, and marketing advantage will accrue to whoever pairs models with the full compute-to-agent stack, as Nvidia has done since the A100 era.
The trend: General-purpose AI agents are moving from narrow task competence toward demonstrated novel-environment reasoning, with chipmakers using benchmark saturation to market integrated stacks.