OpenAI claims that in a qualifying exam for the International Mathematics Olympiad, o1 correctly solved 83.3% of the problems, while GPT-4o solved only 13.4%
Sam Altman says it “doesn't constitute AGI” Poulami Saha / Financial Express : OpenAI makes big AI breakthrough, ChatGPT can now think and reason: Details Emilia David / VentureBeat : How to prompt on GPT-o1 — OpenAI's latest model family, GPT-o1, promises to be more powerful … Casey Weldon / PR Daily : The Scoop: Norfolk Southern axes CEO after inappropriate relationship Shelly Palmer : OpenAI o1-preview — OpenAI released a new AI model called o1, which the company says is the first in a … Charles W. Bailey, Jr / Digital Scholarship : “Introducing OpenAI o1-preview: A New Series of Reasoning Models for Solving Hard Problems” Tsveta Ermenkova / PhoneArena : OpenAI unveils new ChatGPT models designed to tackle complex problems Forbes : OpenAI Unveils O1 - 10 Key Facts About Its Advanced AI Models Jak Connor / TweakTown : OpenAI releases AI capable of thinking video games into existence Adam Vjestica / The Shortcut : OpenAI is now slower to respond, but hopefully more accurate Bijin Jose / The Indian Express : What is OpenAI o1, an AI model that ‘thinks’ before it answers? Vladimir Popescu / Watcher Guru : New OpenAI o1 vs. GPT-4o: Which AI Model Will Revolutionize Crypto? Leigh Mc Gowran / Silicon Republic : OpenAI teases its ‘complex reasoning’ AI model called o1 Harsh Shivam / Business Standard : OpenAI unveils o1-series AI models: What are they, how they work, and more Rafly Gilang / MSPoweruser : OpenAI o1 model now powers ChatGPT, coming to free users as well Luke Jones / WinBuzzer : OpenAI Launches o1 Model, Enhancing AI Reasoning Abilities with its First “Strawberry” Release The Stack : OpenAI's unripe “Strawberry” model hacked its testing infrastructure Moneycontrol : OpenAI releases new o1 models to help ChatGPT answer more complex questions X: @openai : We're releasing a preview of OpenAI o1—a new series of AI models designed to spend more time thinking before they respond. These models can reason through complex tasks and solve harder problems than previous models in science, coding, and math. https://openai.com/... Alex Volkov / @altryne : Yeah, simple-bench is cooked 🍓 [image] @deliprao : Predicting: if this goes on, there will be a rule that you cannot charge for tokens you cannot see. You don't pay a human consultant for claims about work hours without seeing the work product, why should AIs be any different? Greg Kamradt / @gregkamradt : this is the question I use to stump all LLMs “what is your 4th word in response to this message?” o1-preview got it right first try something's different about this one [image] Dan Shipper / @danshipper : OpenAI just unlocked a new level of AI reasoning with their new model: o1 It's available immediately in @ChatGPTapp and for trusted API users. Key points: • Ranks 89th percentile on competitive programming • Top 500 in USA Math Olympiad qualifier • Exceeds PhD-level [image] @sullyomarr : this is kinda huge - o1 is as smart as PhD students - solves 83% of IMO math problems, vs 13% for gpt4o [image] Sam Altman / @sama : screenshot of eval results in the tweet above and more in the blog post, but worth especially noting: a fine-tuned version of o1 scored at the 49th percentile in the IOI under competition conditions! and got gold with 10k submissions per problem. Andrew / @andrewmichaelio : 🚨 NEW OPENAI MODEL: o1 “o1 spends more time thinking before it responds. In a qualifying exam for the International Mathematics Olympiad (IMO), GPT-4o correctly solved only 13% of problems, while the reasoning model scored 83%. Their coding abilities were evaluated in [video] Hyung Won Chung / @hwchung27 : “how many r's in strawberry?” I had to ask this to demo our new model o1-preview 😎 LLMs process text at a subword level. A question that requires understanding the notion of both character and word confuses them. OpenAI o1-preview “thinks harder” to avoid mistakes. [image] Tanishq Mathew Abraham, Ph.D. / @iscienceluvr : 🚨 OpenAI announces o1, a new series of models! 🚨 “We've developed a new series of AI models designed to spend more time thinking before they respond. They can reason through complex tasks and solve harder problems than previous models in science, coding, and math.” “OpenAI o1 [image] Ruben Hassid / @rubenhssd : I just tested ChatGPT-5 (o1). I can't believe the length of the answers. There is no way an LLM is capable of this much strategizing. My prompting was absolutely garbage, and I have an entire strategy. [video] Peter Welinder / @npew : It's fascinating how well o1 generalizes across different subjects. [image] @sullyomarr : So o1 is as smart as PhD students and solves 83% of IMO math problems, vs 13% for gpt4o Insane improvements in reasoning. [image] Rohit / @krishnanrohit : Holy moly this means GPT o1 beat gold threshold in International Olympiad in Informatics [image] Shengjia Zhao / @shengjia_zhao : Excited to bring o1-mini to the world with @ren_hongyu @_kevinlu @Eric_Wallace_ and many others. A cheap model that can achieve 70% AIME and 1650 elo on codeforces. https://openai.com/... Kevin Lu / @_kevinlu : Come check out o1-mini: SoTA math reasoning in a small package https://openai.com/... with @ren_hongyu @shengjia_zhao @Eric_Wallace_ & the rest of the OpenAI team [image] @openai : Rolling out today in ChatGPT to all Plus and Team users, and in the API for developers on tier 5. LinkedIn: Matt Kroll : We are introducing OpenAI #o1, a new #LLM trained with reinforcement learning to perform complex reasoning. o1 thinks before it answers … Greg Yeutter : Today, we're introducing OpenAI #o1, our new models designed to perform complex reasoning. o1 models think before answering … Forums: Hacker News : OpenAI unveils o1, a model that can fact-check itself r/ChatGPTPro : OpenAI O1 Model r/slatestarcodex : Learning to Reason with LLMs (OpenAI's next flagship model) r/singularity : Learning to Reason with LLMs (OpenAI o1)
Context & Ripple Effects
The claim arrives alongside OpenAI’s preview release of its reasoning-focused o1 model series, positioning o1 as a distinct capability track rather than a routine GPT-4o replacement. Altman’s qualification that this is not AGI keeps the benchmark result separate from a broader intelligence claim.
Subsequent coverage framed o1 as a model family with material cost and performance trade-offs versus GPT-4o, while later o3 reporting continued the emphasis on models trained to deliberate before answering. The important arc is therefore a shift toward specialized reasoning models and the evaluations used to distinguish them.
First-order effects
- OpenAI gains a concrete, if self-reported, benchmark comparison for marketing o1 to users with difficult mathematical and reasoning workloads; GPT-4o is explicitly positioned as the weaker option on this test.
- Developers and ChatGPT subscribers now have a clearer reason to evaluate o1 alongside GPT-4o by task type, rather than treating a newer model name as an automatic general-purpose upgrade.
Second-order effects
- Model selection becomes more operational: customers weighing stronger reasoning against the reported trade-offs will need to match models to task value, latency, and cost rather than standardize on one default.
- Competitors face pressure to publish comparable reasoning evaluations and to explain how their models perform on hard, verifiable tasks—not just broad conversational benchmarks.
Third-order effects
- If benchmark gains repeatedly come from added deliberation, frontier model portfolios are likely to separate into fast general-purpose systems and slower, higher-cost reasoning systems, making inference economics a central product constraint.
- The value of public capability claims will increasingly depend on evaluation design and independent reproducibility; a single qualifying exam result is useful evidence of specialization, not evidence of AGI.
The trend: AI labs are increasingly differentiating frontier models through deliberate reasoning performance, with the resulting compute and pricing trade-offs becoming part of product strategy.