/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

← → days · ↑ ↓ browse · Enter similar · o open

AI language models like GPT-3 can achieve up to 97% accuracy on some Winograd schemas, but understanding language doesn't equate to understanding the world

It's simple enough for AI to seem to comprehend data, but devising a true test of a machine's knowledge has proved difficult.

Quanta Magazine Melanie Mitchell

Context & Ripple Effects

This piece lands mid-arc in the debate over what large models actually know. Early skepticism about superficial knowledge in OpenAI's GPT-2 gave way, after GPT-3's release, to researchers treating it as an unexpected step toward machines that grasp human language.

Quanta's point here cuts against that optimism: GPT-3 scoring up to 97% on some Winograd schemas shows the tests measure pattern completion as much as comprehension, and devising a genuine test of machine knowledge has itself proved difficult.

First-order effects

  • Winograd schema scores stop functioning as evidence of understanding — evaluators citing near-perfect results on them are measuring a saturated benchmark, not comprehension.

Second-order effects

  • Benchmark designers respond by building harder evaluations, a path that leads to tests like ARC-AGI-2, where humans score 60% while leading models score around 1%.

Third-order effects

  • If each model generation saturates its predecessors' tests, the industry's credibility problem shifts from 'can it pass' to 'what does passing mean' — visible in GPT-4's gains in precision and image input alongside continued hallucination.

The trend: Language-model evaluation is locked in a saturation cycle: each benchmark that models ace gets retired in favor of harder ones, because fluency keeps outrunning demonstrated world knowledge.