/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

The Arc Prize Foundation says its new ARC-AGI-2 test stumps most AI models; humans get 60% of the questions right but GPT-4.5 and Claude 3.7 Sonnet score ~1%

[image] François Chollet / @fchollet : Unlike ARC-AGI-1, this new version is not easily brute-forced.  Current top AI approaches score 0-4%.  All base LLMs (GPT-4.5, Claude 3.7 Sonnet, Gemini 2, etc.) score 0%.  Single-CoT reasoning models (Claude Thinking, R1, o3-mini...) score 0-1%.  So you can't solve these tasks via memorization alone.  You need the ability to recombine concepts on the fly - you need test-time adaptation... Peter Wildeford / @peterwildeford : Very interesting - the new ARC prize finds tasks where even advanced reasoning models don't do anywhere close to human level. [image]

TechCrunch Maxwell Zeff

Discussion

  • @fchollet François Chollet on x
    Today, we're releasing ARC-AGI-2.  It's an AI benchmark designed to measure general fluid intelligence, not memorized skills - a set of never-seen-before tasks that humans find easy, but current AI struggles with.  It keeps the same format as ARC-AGI-1, while significantly increa…
  • @arcprize @arcprize on x
    Today we are announcing ARC-AGI-2, an unsaturated frontier AGI benchmark that challenges AI reasoning systems (same relative ease for humans). Grand Prize: 85%, ~$0.42/task efficiency Current Performance: * Base LLMs: 0% * Reasoning Systems: <4% [image]
  • @benpielstick Ben Pielstick on x
    ARC-AGI just got a big update. This is in my opinion the most interesting benchmark in AI because it demonstrates outside the box thinking. While math and programming benchmarks are more utilitarian, abstraction opens up whole new possibilities we humans can't even conceive of.
  • @mikeknoop Mike Knoop on x
    The $1,000,000 @arcprize 2025 competition is back! And introducing ARC-AGI-2 the only unbeaten benchmark (we're aware of) that remains easy for humans but now even harder for AI. New ideas are still needed to reach AGI. We've got lots of great updates for 2025 — [image]
  • @fchollet François Chollet on x
    Unlike ARC-AGI-1, this new version is not easily brute-forced.  Current top AI approaches score 0-4%.  All base LLMs (GPT-4.5, Claude 3.7 Sonnet, Gemini 2, etc.) score 0%.  Single-CoT reasoning models (Claude Thinking, R1, o3-mini...) score 0-1%.  So you can't solve these tasks v…
  • @peterwildeford Peter Wildeford on x
    Very interesting - the new ARC prize finds tasks where even advanced reasoning models don't do anywhere close to human level. [image]