The Arc Prize Foundation says its new ARC-AGI-2 test stumps most AI models; humans get 60% of the questions right but GPT-4.5 and Claude 3.7 Sonnet score ~1%
[image] François Chollet / @fchollet : Unlike ARC-AGI-1, this new version is not easily brute-forced. Current top AI approaches score 0-4%. All base LLMs (GPT-4.5, Claude 3.7 Sonnet, Gemini 2, etc.) score 0%. Single-CoT reasoning models (Claude Thinking, R1, o3-mini...) score 0-1%. So you can't solve these tasks via memorization alone. You need the ability to recombine concepts on the fly - you need test-time adaptation... Peter Wildeford / @peterwildeford : Very interesting - the new ARC prize finds tasks where even advanced reasoning models don't do anywhere close to human level. [image]
Today, we're releasing ARC-AGI-2. It's an AI benchmark designed to measure general fluid intelligence, not memorized skills - a set of never-seen-before tasks that humans find easy, but current AI struggles with. It keeps the same format as ARC-AGI-1, while significantly increa…
Today we are announcing ARC-AGI-2, an unsaturated frontier AGI benchmark that challenges AI reasoning systems (same relative ease for humans). Grand Prize: 85%, ~$0.42/task efficiency Current Performance: * Base LLMs: 0% * Reasoning Systems: <4% [image]
ARC-AGI just got a big update. This is in my opinion the most interesting benchmark in AI because it demonstrates outside the box thinking. While math and programming benchmarks are more utilitarian, abstraction opens up whole new possibilities we humans can't even conceive of.
The $1,000,000 @arcprize 2025 competition is back! And introducing ARC-AGI-2 the only unbeaten benchmark (we're aware of) that remains easy for humans but now even harder for AI. New ideas are still needed to reach AGI. We've got lots of great updates for 2025 — [image]
Unlike ARC-AGI-1, this new version is not easily brute-forced. Current top AI approaches score 0-4%. All base LLMs (GPT-4.5, Claude 3.7 Sonnet, Gemini 2, etc.) score 0%. Single-CoT reasoning models (Claude Thinking, R1, o3-mini...) score 0-1%. So you can't solve these tasks v…