A critical look at OpenAI's GPT-2 text generation AI and why the knowledge acquired by systems like the GPT-2 has been superficial and unreliable
OpenAI's GPT-2 has been discussed everywhere from The New Yorker to The Economist. What does it really tell us about natural and artificial intelligence? Tweets: @sapinker , @talyarkoni , @plevy , @benambridge , @bengoertzel , @stevestuwill , @ajmackiel , and @pauldaugh Tweets: Steven Pinker / @sapinker : For decades I've argued that a key test of innateness hypotheses (e.g., Chomsky on language) is whether an AI could learn to speak & think like people with no innate structure. @garymarcus argues that GPT-2 is now providing that test. https://thegradient.pub/... Tal Yarkoni / @talyarkoni : a few years from now, when GPT-9 is happily writing novels, I fully expect to see @GaryMarcus writing opinion pieces complaining that it isn't *real* intelligence until it displays context-appropriate emotions or fine motor control https://twitter.com/... Pierre Levy / @plevy : «GPT-2 has been a monumental experiment in Locke's hypothesis, and it has failed. (...) Even with massive data sets and enormous compute, the knowledge that it acquires has been superficial and unreliable. » by @GaryMarcus via @kmesch1 #AI https://thegradient.pub/... Ben Ambridge / @benambridge : The whole point of usage based approaches to language is that syntactic generalization is constrained by communicative function - so of COURSE a system with no communicative goals produces garbage https://twitter.com/... Ben Goertzel / @bengoertzel : @GaryMarcus gives a clear explanation, with examples, of what GPT2 and other transformer NNs lack in terms of fundamental comprehension ... https://thegradient.pub/... Steve Stewart-Williams / @stevestuwill : “Upon careful inspection, it becomes apparent that [GPT-2, a language-generating AI] has no idea what it is talking about... Rather than supporting the Lockean, blank-slate view, GPT-2 appears to be an accidental counter-evidence to that view.” https://thegradient.pub/... https://twitter.com/... Alexander Mackiel / @ajmackiel : Throwing more computing power into a purely empiricist process won't create the kind of genuine understanding a system like the human brain has: “GPT-2 is both a triumph for empiricism, and...a clear sign that it is time to consider investing in different approaches."@GaryMarcus https://twitter.com/... Paul Daugherty / @pauldaugh : “Current systems can regurgitate knowledge, but they can't really understand in a developing story, who did what to whom, where, when, and why; they have no real sense of time, or place, or causality.” Read this by @GaryMarcus #AI https://thegradient.pub/...
Context & Ripple Effects
When Gary Marcus published this critique in The Gradient, GPT-2 was being celebrated everywhere from The New Yorker to The Economist, and his argument — that its acquired knowledge is superficial and unreliable — drew supporting responses from academics including Steven Pinker, Tal Yarkoni, Ben Ambridge, and Ben Goertzel. He framed the system as a live test of Chomsky-style innateness hypotheses: whether raw text prediction alone can yield human-like language and thought.
What makes the piece worth rereading is how the subsequent record plays out. ChatGPT marked a genuine step change in use cases, but GPT-4 still hallucinates despite added precision and image input, GPT-5 landed as incremental rather than transformative, and skeptics now argue builders shouldn't assume the tech is world-changing at all (the 'could be a dud' case). The reliability gap Marcus flagged in 2020 is the through-line.
First-order effects
- Marcus and his supporting commentators force the debate onto a specific question — does next-token prediction produce understanding or surface statistics — rather than letting OpenAI's demos stand as evidence of general intelligence.
- OpenAI faces a framing problem at the height of GPT-2's publicity cycle: the most-cited academic response to its flagship model is a critique of what the model actually knows.
Second-order effects
- Every later OpenAI release gets measured against the reliability bar this critique set — GPT-4's residual hallucinations and GPT-5's underwhelming reception are both read as scorecards for the pure-scaling thesis, not just product reviews.
- Buyers and analysts separate fluency from trustworthiness when evaluating generative systems, with error tolerance becoming an explicit evaluation criterion alongside capability benchmarks.
Third-order effects
- If the pattern holds — each generation more fluent but still unreliable — the industry's center of gravity shifts from scaling claims toward verification, grounding, and trust, the position argued in The Road to AI We Can Trust.
- The innateness-versus-learning question Marcus posed stays unresolved, meaning cognitive-science critiques remain a standing constraint on how AGI progress gets claimed and audited.
The trend: Each frontier-model release reopens the same dispute Marcus opened with GPT-2 — whether scale alone closes the gap between fluent output and reliable knowledge — and the answer so far keeps the critique alive.