Hugging Face cofounder Thomas Wolf says current AI development paradigms won't yield outside-the-box problem solving that leads to true scientific breakthroughs
AI company founders have a reputation for making bold claims about the technology's potential to reshape fields, particularly the sciences. Bluesky: @prietschka and @martinsfp X: @thom_wolf , @thom_wolf , @littmath , and @thom_wolf Forums: Hacker News , r/technology , and Slashdot Bluesky: Paul Rietschka / @prietschka : It's hilarious this is coming from Huggingface, a company that does...what, exactly? — It incinerates cash at a fantastical rate, that's for certain, but I don't see much more. [embedded post] @martinsfp : Yes! “Wolf thinks that AI labs are building what are essentially ‘very obedient students’ — not scientific revolutionaries in any sense of the phrase” [embedded post] X: Thomas Wolf / @thom_wolf : I shared a controversial take the other day at an event and I decided to write it down in a longer format: I'm afraid AI won't give us a “compressed 21st century”. The “compressed 21st century” comes from Dario's “Machine of Loving Grace” and if you haven't read it, you probably should, it's a noteworthy essay. In a nutshell the paper claims that, over a year or two, we'll have a “country of Einsteins sitting in a data center”, and it will result in a compressed 21st century during which all the scientific discoveries of the 21st century will happen in the span of only 5-10 years. I read this essay twice. The first time I was totally amazed: AI will change everything in science in 5 years, I thought! A few days later I came back to it and, re-reading it, I realized that much of it seemed like wishful thinking at best. What we'll actually get, in my opinion, is “a country of yes-men on servers” (if we just continue on current trends). Let me explain the difference with a small part of my personal story. … Thomas Wolf / @thom_wolf : Interesting thread from Daniel Litt on what it takes to do real research in math for an AI and a human I guess we had similar thoughts this week as it's really aligned with my AI-Einstein short essay: https://thomwolf.io/... Daniel Litt / @littmath : In this thread I want to share some thoughts about the FrontierMath benchmark, on which, according to OpenAI, some frontier models are scoring ~20%. This is benchmark consisting of difficult math problems with numerical answers. What does it measure, and what doesn't it measure? [image] Thomas Wolf / @thom_wolf : Many are proposing “move 37” as evidence that AI has already reached Einstein-level intelligence, so I'd like to expand on this specific example. Move 37, while impressive, is still essentially a straight-A student answer to the question posed by the rules of the game of Go. Forums: Hacker News : AI is becoming ‘yes-men on servers’ r/technology : Hugging Face's chief science officer worries AI is becoming ‘yes-men on servers’ Msmash / Slashdot : AI Isn't Creating New Knowledge, Hugging Face Co-Founder Says
Context & Ripple Effects
Wolf’s intervention lands amid evidence that frontier-model progress is constrained by more than ambition: reporting on OpenAI’s next model described delays, compute costs, and limits on high-quality training data. It directly challenges the assumption that scaling existing approaches naturally leads to scientific originality.
Related coverage has also separated AI’s useful, bounded capabilities from claims of humanlike intelligence in the debate over AI as a normal technology. That distinction matters as labs increasingly attach scientific and societal significance to their systems.
First-order effects
- Hugging Face and other AI labs face sharper scrutiny over whether current models can generate genuinely novel scientific hypotheses rather than improve performance on established tasks.
- The claim reframes benchmark- and scale-led progress as insufficient evidence for the “AI-Einstein” narrative, raising the bar for scientific-breakthrough assertions.
Second-order effects
- Labs pursuing scientific customers will have greater incentive to demonstrate discovery workflows and validation, not only stronger general-purpose model results.
- Constraints already reported around training data and compute make the case for alternative research approaches more salient, rather than treating additional scale as the sole path forward.
Third-order effects
- If this critique is borne out, frontier AI could differentiate into systems optimized for dependable cognitive assistance and a separate, harder-to-prove pursuit of open-ended scientific reasoning.
- The gap between capable automation and autonomous discovery may become central to how institutions evaluate AI-lab claims, investment priorities, and policy narratives.
The trend: The story is one data point in the shift from treating scale-driven model gains as a proxy for general intelligence to demanding evidence of novel reasoning and real-world scientific value.