AI language models like GPT-3 can achieve up to 97% accuracy on some Winograd schemas, but understanding language doesn't equate to understanding the world
It's simple enough for AI to seem to comprehend data, but devising a true test of a machine's knowledge has proved difficult.
Context & Ripple Effects
This piece lands mid-arc in the debate over what large models actually know. Early skepticism about superficial knowledge in OpenAI's GPT-2 gave way, after GPT-3's release, to researchers treating it as an unexpected step toward machines that grasp human language.
Quanta's point here cuts against that optimism: GPT-3 scoring up to 97% on some Winograd schemas shows the tests measure pattern completion as much as comprehension, and devising a genuine test of machine knowledge has itself proved difficult.
First-order effects
- Winograd schema scores stop functioning as evidence of understanding — evaluators citing near-perfect results on them are measuring a saturated benchmark, not comprehension.
Second-order effects
- Benchmark designers respond by building harder evaluations, a path that leads to tests like ARC-AGI-2, where humans score 60% while leading models score around 1%.
Third-order effects
- If each model generation saturates its predecessors' tests, the industry's credibility problem shifts from 'can it pass' to 'what does passing mean' — visible in GPT-4's gains in precision and image input alongside continued hallucination.
The trend: Language-model evaluation is locked in a saturation cycle: each benchmark that models ace gets retired in favor of harder ones, because fluency keeps outrunning demonstrated world knowledge.