GPT-5's release was underwhelming, offering incremental improvements and failing to meet expectations, showing that pure scaling simply isn't the path to AGI
and he's not alone Maximilian Schreiner / The Decoder : GPT-5 is here and Gary Marcus is not impressed Laura Varley / Silicon Republic : Altman admits GPT-5 currently ‘way dumber’ amid rough roll-out Matthew Bishop / The Observer : ChatGPT update is big step foward, says OpenAI Noah Bovenizer / The Stack : OpenAI threatens rate limits but restores “legacy” models after outcry Luiza Jarovsky, PhD / Luiza's Newsletter : GPT-5 And The Mirage of AGI · OpenAI's New Open-Weight Models · And More Marcus Schuler / Implicator.ai : GPT-5 exposes scaling limits, accelerates shift to specialized models Mathieu Acher : GPT-5 and GPT-5 Thinking as Other LLMs in Chess: Illegal Move After 4th Turn Realtime Techpocalypse Newsletter : GPT-5 Should Be Ashamed of Itself Bluesky: Lincoln Michel / @thelincoln : Yes, but this would require tech journalists who didn't (9 times out of 10) parrot the same ridiculous claims. E.g., someone like NYT's Kevin Roose is never going to hold CEOs to account lest he be held to account for mindlessly hyping NFTs and the Metaverse — garymarcus.substack.com/p/gpt-5- over... … Mastodon: @leighelse@mastodon.nz : Open AI's release of GPT5 shows a substantial gap between artificial intelligence promises and reality. — https://garymarcus.substack.com/ ... 'Altman's reputation should by now be completely burned. … X: Ethan Mollick / @emollick : The issue with GPT-5 in a nutshell is that unless you pay for model switching & know to use GPT-5 Thinking or Pro, when you ask “GPT-5” you sometimes get the best available AI & sometimes get one of the worst AIs available and it might even switch within a single conversation. [image] @swarley75042565 : @GaryMarcus Dario was on every podcast a year ago saying that AGI was imminent in 2027 and that everything shows that the scaling laws hold and all we need is a bigger cluster Gary Marcus / @garymarcus : If you want to hear fantasies about the future, listen to @sama. If you want to listen to someone with a strong, documented track record of actually predicting the future, come to me. Gary Marcus / @garymarcus : AI, which has never invented any science of significance, is going to be smarter than any kid born this decade? I call bullshit. Looks to me like a smokescreen to distract from the disappointment of GPT-5. Daniel Eth / @daniel_271828 : GPT-5 didn't live up to OpenAI's hype, but it is *exactly* in line with extrapolations from prior AI advancements. Go ahead and discount future statements from OpenAI/Altman, but you should still expect the fast AI progress that we've been seeing to continue [image] Gary Marcus / @garymarcus : LLM alone NGMI acronyms to live by Gary Marcus / @garymarcus : Incredible how well these March 2024 predictions held up, and not just for 2024, but more than halfway into 2025. Yes, we have many more GPT-4 level models, and yes we have o3 etc, but GPT-5 was disappointing, and took far far longer than most people expected. And we STILL Gary Marcus / @garymarcus : and he was wrong, as should now be evident Ewan Morrison / @mrewanmorrison : The essential read on why the flop of GPT5 means so much more than AI companies admit and could be a wake-up-call for all who bought into the hype. Yes, Gary Marcus was right. [image] Aidan McLaughlin / @aidan_mclau : gpt-5 is above trend if you or someone you know has updated to “agi over bro” after its release, i have no idea what model of the future you were working with extrapolate this and we have models doing month-long projects in 2027 [image] Luiza Jarovsky, PhD / @luizajarovsky : Sam Altman: With GPT-5, you'll have a PhD-level expert in any area you need Me: Draw a map of North America, highlighting countries, states, and capitals GPT 5: *Sam Altman forgot to mention that the PhD-level expert used ChatGPT to cheat on all their geography classes... [image] Yuchen Jin / @yuchenj_uw : GPT-5 failed twice. Scaling laws are coming to an end. Open-source AI will have the Mandate of Heaven. Quanquan Gu / @quanquangu : You're only half right. Yes, GPT-5 failed twice. No, scaling laws aren't over. Do them right, and the game goes on. @scaling01 : things got so bad, that Gary Marcus unblocked me, dm'd me, apologized and quoted me twice in his latest article get your shit together OpenAI [image] Dr. Émile P. Torres / @xriskology : Fantastic article on how GPT-5 is “overdue, overhyped, and underwhelming.” People at OpenAI have been mocking Gary Marcus for years (lolz), but turns out that Marcus is right about the inherent limitations of LLMS. Last laugh goes to Marcus! Read it here: https://garymarcus.substack.com/ ... [image] @the_ai_skeptic : Once again @GaryMarcus nails it. The bigger question for me is why none of the big influencers in the tech space, or even tech writers for mainstream news outlets, are covering the disappointing debut of GPT-5, or the embarrassing user backlash. https://garymarcus.substack.com/ ... [image] Gary Marcus / @garymarcus : Check it out. Almost everyone in the major media is missing the real story around GPT-5. The real story is about how so many people (even big fans of OpenAI) were disappointed. And it's about how that may well spell the end of scaling mania. And it's about how the premature release of GPT-5 may prove to be OpenAI's biggest blunder... Ben Ansell / @benwansell : Brutal analysis of ChatGPT5 from @GaryMarcus. This was a big moment for OpenAI and so far a dud. Since US economy is largely being kept afloat by AI investment, this could be inflection point. Hold onto your hats. https://garymarcus.substack.com/ ... Jeremy Howard / @jeremyphoward : Now that the era of the scaling “law” is coming to a close, I guess every lab will have their Llama 4 moment. Grok had theirs. OpenAI just had theirs too. Gary Marcus / @garymarcus : My work here is truly done. Nobody with intellectual integrity can still believe that pure scaling will get us to AGI. GPT-5 may be a moderate quantitative improvement (and it may be cheaper) but it still fails in all the same qualitative ways as its predecessors, on chess, on reasoning, in vision; even sometimes on counting and basic math. Hallucinations linger. Dozens of shots on goal (Grok, Claude, Gemini) etc have invariably faced the same problems... Colin Fraser / @colin_fraser : Oh brother [Screenshot of ChatGPT 5 solving the 5.9 = x + 5.11 equation wrong] Burny / @burny_tech : GPT-5 isn't scoring well on SimpleBench, relatively speaking [image] Anoop / @anoop_331 : @GaryMarcus Gotta agree, O3 was a shit good model, from O3 to gpt-5, whether it's a model or router or whatever it is, it was an utter disappointment, especially given the kind of hype towards its release. Egide Murisa / @egidemurisa : @GaryMarcus 4o was a breakthrough. Great alternative to day to day Google search. Made me believe GPT-5 was gonna be AGI. Expectations were high! I wonder what impact this will have on the AI hardware industry. @mgonto : The saddest thing on my day is that @GaryMarcus is right. I hate it! @scaling01 : OpenAI: “We are going to fix model naming and make it less confusing” also OpenAI: [image] Cameron Williams / @wasgo : @GaryMarcus Found a visually simple research failure. In it's rankings, Eagle, Owl, Frog and Rabbit don't exist. Bat, Spider, Rat, Ant, Lizard are missing. Possibly Polar Bear became Bear, Lynx became Cat, though the alternatives of Crocodile, Boar and Coyote don't exist. [image] @explodemeow102 : @GaryMarcus How should one choose? [image] Gary Marcus / @garymarcus : Sorry, but this tweet did not age well. [Quotes Sam Altman's X post with Death Star rising from GPT-5 launch day] @scaling01 : Markets disappointed by GPT-5 OpenAI getting crushed on Polymarket [image] LinkedIn: Hugi Aegisberg : Now that we see that GPT5 is really just a small incremental improvement on earlier generations, perhaps it is time for the AGI-around-the-corner crowd … Gary Marcus : Already described as “marvelous”, this is one of my finest pieces. — GPT-5: Overdue, overhyped, and underwhelming. And that's not the worst of it. … Chris Richardson : A long but entertaining and worthwhile read. — Part 1 discusses the underwhelming release of GPT-5 and the growing disconnect between the hype and the reality. … Forums: Hacker News : GPT-5: Overdue, overhyped and underwhelming. And that's not the worst of it Hacker News : Wait, isn't the Bernoulli effect thing they're demoing now wrong? I thought that was a … r/artificial : GPT-5: Overdue, overhyped and underwhelming. And that's not the worst of it. r/agi : GPT-5: Overdue, overhyped and underwhelming. And that's not the worst of it.
Context & Ripple Effects
OpenAI had already positioned GPT-4.5 as a research preview rather than a frontier model, despite describing it as its most knowledgeable model. That made the next flagship release a test of whether larger general-purpose models could deliver a clearer step-change.
Early hands-on coverage found competent but not dramatically ahead performance alongside aggressive pricing. The reported rollout problems, user backlash over legacy-model access, and disappointment therefore matter as much for product trust as for the AGI narrative.
First-order effects
- OpenAI must manage immediate customer dissatisfaction: it restored legacy models after outcry, while reported in-conversation model switching makes output quality less predictable for users who cannot pay to choose models.
- The gap between elevated expectations and incremental reported gains weakens the release’s immediate persuasive power with users and market observers, even as OpenAI continues to offer the new model.
Second-order effects
- Rival model providers can compete more directly on consistency, controllability, price, and task-specific performance rather than needing to match a perceived generational leap.
- Enterprise buyers are likely to place more weight on workload-level evaluation and fallback options; GPT-4.5’s earlier positioning below reasoning models had already signaled that a single general model may not lead every task.
Third-order effects
- If repeated flagship releases yield uneven practical gains, AI products may shift toward portfolios of specialized models and explicit routing rather than a single scaled model presented as a path to AGI.
- The episode raises the evidentiary bar for broad AGI claims: benchmark or capability messaging will need to translate into stable, user-visible performance before it can sustain product and market expectations.
The trend: This is one data point in the shift from scaling-led frontier-model narratives toward reliable, specialized, and economically competitive AI systems.