A look at seven rebuttals to Apple's paper on limitations of Large Reasoning Models, and why none make a compelling case
‘superintelligence’ or the ‘illusion of thinking’? arXiv.org e-Print archive : Comment on The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem Complexity Bluesky: Paul Rietschka / @prietschka : “If people like Sam Altman are sweating, it's because they should.” — Lol. Claus Wilke / @clauswilke.com : Gary Marcus replies to the critiques of the Apple paper on reasoning models. — garymarcus.substack.com/p/seven- repl... @onisillos : The paradox of AI in a nutshell: — “One is left simply having to test everything, all the time, with little guarantees of anything.” — garymarcus.substack.com/p/seven- repl... @its-tom-williams : There's been an interesting follow-up here from AI expect @garymarcus.bsky.social, who rebutted the “roughly seven different efforts at rebuttal” he'd seen people make against the Apple paper. X: Gary Marcus / @garymarcus : All the replies to the Apple paper fall short, but explaining why is just too long for a post on X. But perfect for a short newsletter. At the link in the bottom left of the image, I dissect 7 major replies. Some are clever; some are not. None change things fundamentally, at [image] LinkedIn: John Thompson : I took some time this afternoon to carefully read the Apple research paper - The Illusion of Thinking. My conviction and belief that Large Reasoning Models (LRMs) do not think is even stronger. … Matthew Brown : Apple's new research, “The Illusion of Thinking,” offers a sobering look at the limits of Large Reasoning Models (LRMs) — models that generate step-by-step “thoughts” before giving an answer. … Forums: Hacker News : Seven replies to the viral Apple reasoning paper and why they fall short r/LocalLLaMA : Comment on The Illusion of Thinking: Recent paper from Apple contain glaring flaws in the original study's experimental design, from not considering token limit to testing unsolvable puzzles.
Context & Ripple Effects
This is the latest turn in a recurring dispute over whether strong performance on reasoning benchmarks demonstrates formal reasoning. Apple’s earlier work argued that language-model behavior was better explained by pattern matching, while its more recent study of large reasoning-model limits on classic puzzles sharpened that claim.
Marcus’s review of seven responses matters because it shifts the focus from the paper’s headline conclusion to the validity of its evaluation design. The disagreement follows a longer discussion of how models are trained to present reasoning and how that presentation should be evaluated when assessing general reasoning.
First-order effects
- Apple’s paper receives a more sustained public defense from Marcus, while critics’ objections—especially around experimental design and token limits—remain central to interpreting its conclusions.
- Developers and buyers evaluating Large Reasoning Models get a clear warning not to treat apparent step-by-step outputs or benchmark successes as sufficient evidence of robust reasoning.
Second-order effects
- Model providers face added pressure to disclose evaluation conditions and demonstrate performance across problem complexity, rather than relying on selected reasoning examples.
- The dispute makes independent, adversarial testing more important for organizations deciding where reasoning models can be trusted without extensive verification.
Third-order effects
- If such findings continue to withstand methodological challenge, AI competition may shift from broad claims of reasoning ability toward measured reliability on bounded tasks and explicit human-checking workflows.
- The larger uncertainty is not whether models can solve useful problems, but which evaluation methods can distinguish genuine generalization from performance that breaks under changed task conditions.
The trend: The story is one data point in the push to replace persuasive “reasoning” demonstrations with tougher, reproducible tests of model reliability and limits.