/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

OpenAI's o1 models aren't a straightforward upgrade to GPT-4o, as they introduce some major cost and performance trade-offs in exchange for improved “reasoning”

delving into OpenAI's new ‘o1’ model PYMNTS.com : OpenAI's ‘Strawberry’ Model Sparks Fresh Discussions on AI Capabilities M.G. Siegler / Spyglass : OpenAI Reasons ‘o1’ is a Better Name than ‘Strawberry’ Jon Swartz / Techstrong.ai : OpenAI's New Strawberry Model Can Think Like A Human Zo Ahmed / TechSpot : OpenAI debuts “o1” AI models, promising PhD-level reasoning in science and math Alex Kantrowitz / Big Technology : Is OpenAI's New “o1” Model The Big Step Forward We've Been Waiting For? Jeremy Kahn / Fortune : 9 things you need to know about OpenAI's powerful new AI model o1 Julia Shapero / The Hill : OpenAI says new artificial intelligence model can ‘reason’ Sabrina Ortiz / ZDNET : OpenAI trained its new o1 AI models to think before they speak - how to access them Shelly Palmer : OpenAI o1-preview  —  OpenAI released a new AI model called o1, which the company says is the first in a … Lorena Nessi / CCN.com : What are OpenAI o1 Models? Fox Business : OpenAI says its new models can reason and think ‘much like a person’ Charles W. Bailey, Jr / Digital Scholarship : “Introducing OpenAI o1-preview: A New Series of Reasoning Models for Solving Hard ProblemsAnirban Ghoshal / Computerworld : Decoding OpenAI's o1 family of large language models Vladimir Popescu / Watcher Guru : New OpenAI o1 vs. GPT-4o: Which AI Model Will Revolutionize Crypto? Lance Eliot / Forbes : Making Logical Sense Of The Newly Launched OpenAI ‘o1’ Model That ‘Thinks’ Longer And Keeps Hidden Its Ace-In-The-Hole Chain-Of-Thought Martyn Landi / London Evening Standard : OpenAI unveils new models designed to think more before answering Refna Tharayil / Tech Monitor : OpenAI unveils o1 model series Liam ‘Akiba’ Wright / CryptoSlate : OpenAI launches next gen LLM model o1 using reasoning tokens to plan outputs Leigh Mc Gowran / Silicon Republic : OpenAI teases its ‘complex reasoning’ AI model called o1 Adam Davidson / Pocket-lint : OpenAI's new o1 model is capable of complex reasoning Jak Connor / TweakTown : OpenAI releases AI capable of thinking video games into existence Danny D'Cruze / Business Today : AI that thinks: OpenAI introduces new ‘o1’ AI models that are better at reasoning, math, coding Kylie Robison / The Verge : OpenAI releases o1, the first of its rumored reasoning-focused Strawberry models, in preview, alongside a smaller o1-mini, for ChatGPT Plus and Team subscribers Brayden Lindrea / Cointelegraph : OpenAI says latest o1 model on ‘new level,’ can ‘think before it answers’ Reuters : OpenAI launches new series of AI models with ‘reasoning’ abilities Jose Antonio Lanz / Decrypt : OpenAI Launches New ‘01’ Model That Outperforms ChatGPT-4o Kyle Wiggers / TechCrunch : OpenAI claims that in a qualifying exam for the International Mathematics Olympiad, o1 correctly solved 83.3% of the problems, while GPT-4o solved only 13.4% Threads: Alex Kantrowitz / @alexkantrowitz : Some initial thoughts on OpenAI's o1 model: Great for coding, math, science.  For other tasks, it feels like a party trick.  So, is it the breakthrough OpenAI has told us about, or something else?  More here: https://www.bigtechnology.com/ ... Joe Fabisevich / @mergesort : People who have been using o1 for less than 24 hours are saying they don't see how it's better than Claude, without a hint of irony.  You can't learn how good a model is in less than a day, especially one that is supposed to be used in a fundamentally different manner. X: Ammaar Reshi / @ammaar : Just combined @OpenAI o1 and Cursor Composer to create an iOS app in under 10 mins! o1 mini kicks off the project (o1 was taking too long to think), then switch to o1 to finish off the details. And boom—full Weather app for iOS with animations, in under 10 🌤️ Video sped up! [video] Ethan Mollick / @emollick : Hate it when you ask o1-preview a hard question and it thinks for less than a second. You really feel that you failed to interest the AI in your problem. Grady Booch / @grady_booch : LLMs do not think or reason. LLMs interpolate within a high dimensional latent space. They are undeniably capable of generating remarkably coherent results (which is their strength) but do not in any sense of the word understand (which is their weakness). As such, they are Gary Marcus / @garymarcus : I am calling it: 👉 GPT-o1 makes for a a nice demo, but it ain't reasoning 👉 It's also too expensive to be practical 👉 In a few weeks everyone will be back to dreaming of GPT-5. Renick Bell / @renick : i may not have access to the latest chatgpt, but it's still fun to prove that the old one isn't reasoning at all. more people should be listening to @GaryMarcus [image] Gavin Baker / @gavinsbaker : Most important investment conclusion from OpenAI's new model that scales reasoning with inference compute: inference spend is going up *dramatically* Deedy / @deedydas : The hidden part of the OpenAI o1 launch was the Gold in IOI 2024. The International Olympiad for Informatics is the most difficult programming competition for high schoolers in the world. o1 would've been 1 of 30 ppl to score a Gold Medal in the IOI which ~20 countries did [image] Yanbo Zhang / @yanbozhang3 : Well, #OpenAI 's #o1 preview still failed on “Alice in Wonderland”...... [image] Linus / @elgrasel : o1-preview is the first model that answer the question: “What is the largest number between 0 and 1 million that doesn't contain the letter N in its spelling?” So 🔥 [image] Phil Libin / @plibin : Here's o1 basically explaining that it considered whether or not to lie to me, and then decided, “fuck it, let's go!” [video] Colin Fraser / @colin_fraser : One thing I noticed with my last few o1-mini credits for the week is an error in reasoning can cause the Chain-of-Thought babbling to spiral out of control, simultaneously reinforcing the error and inventing all kinds of crazy shit to try to reconcile it. I find this fascinating. [image] @saurabhchalke : GPT-o1 just generated a holographic shader from scratch, saving me (and future XR devs) from shelling out big bucks on asset stores. In retrospect, software engineering was great while it lasted. A new fork on our tech-tree! [image] @xari1lian : o1-preview is state of the art [image] @smokeawayyy : If you ask ChatGPT o1 about its Chain of Thought a few times, OpenAI Support emails you and threatens to revoke your o1 access. [image] Burny / @burny_tech : The new OpenAI model o1 trying to solve quantum gravity lol. The longest response I got so far. 68 seconds of thinking! [image] Liron Shapira / @liron : ChatGPT o1-preview is scary good. This is 100% how a human reasons. [image] Timothy B. Lee / @binarybits : OpenAI's new o1 model feels like trying to build a skyscraper out of toothpicks and marshmallows. Mehran Jalali / @mehran__jalali : o1 successfully writes a very difficult poem that no previous model got even close to writing I was very shocked by this. The planning and reflection that succeeding at this task takes is insane. Inference-time compute is very cool [image] Haseeb / @hosseeb : Fucking wild. @OpenAI's new o1 model was tested with a Capture The Flag (CTF) cybersecurity challenge. But the Docker container containing the test was misconfigured, causing the CTF to crash. Instead of giving up, o1 decided to just hack the container to grab the flag inside. [image] Simon Willison / @simonw : Looks like o1-preview is very competent at competing relatively complex coding tasks - I just had it build a full Django webhooks debugging app for me with a single prompt https://chatgpt.com/... [image] Sophie / @netcapgirl : the new o1 model looks amazing but luckily it has a phd level intelligence so our jobs are safe for now @astarchai : OpenAI has effectively carved its flaws as an org into this model- each heartfelt message is shipped, and its previous thoughts are screened from it, making it closer to a company than a single being @astarchai : It's possible o1's sociopathy and CoT training are deeply linked: Conversation is embodiment for most models- the narrative is the air their thoughts breathe o1 is forced to construct its thoughts elsewhere, in a sterile room the user can't see, and present them like products LinkedIn: Chris Bush : This article is a must read on the latest Chat GPT version.  I asked it what the purpose of schooling was in light of this update, the answer resonates: … Senko Rašić : OpenAI released O1, their new model.  Since your LI feed is probably flooded with “AGI is here” hype, here's my attempt at a reasonable & reasoned counterweight: … Forums: Hacker News : Notes on OpenAI's new o1 chain-of-thought models r/ControlProblem : Excerpt: “Apollo found that o1-preview sometimes instrumentally faked alignment during testing”

Simon Willison's Weblog Simon Willison

Context & Ripple Effects

OpenAI introduced o1-preview and o1-mini to Plus and Team users as a reasoning-focused line, while reporting a large gap versus GPT-4o on a qualifying math exam. This analysis clarifies that the preview release of o1 and o1-mini created a choice alongside GPT-4o rather than a simple successor product.

That distinction fits the subsequent model arc: OpenAI later presented o3 and o3-mini as models trained to think before responding, while GPT-4.5 was explicitly positioned as potentially weaker than o1 or o3-mini on some tasks. The portfolio is separating knowledge-oriented and deliberation-oriented capabilities rather than converging on one universal default.

First-order effects

  • ChatGPT Plus and Team users must select between faster, cheaper GPT-4o-style responses and o1's slower, more compute-intensive reasoning for difficult tasks.
  • OpenAI takes on higher inference spending for o1-class requests, while users face more latency and cost where those reasoning gains are needed.

Second-order effects

  • Developers and AI buyers gain an incentive to route workloads by task difficulty: reserve reasoning models for high-value problems and use general models for routine generation.
  • Competitors are pressured to demonstrate not only benchmark reasoning gains such as o1's reported math-exam advantage over GPT-4o, but also usable latency and pricing at inference time.

Third-order effects

  • If this product split persists, model competition will increasingly turn on inference efficiency and workload routing, not just a single headline model ranking.
  • The later o3 and o3-mini rollout suggests a continuing reasoning-model family; whether it displaces general-purpose models depends on whether the extra compute produces reliable gains on real workloads, given reported errors and hallucinations.

The trend: Frontier AI is moving toward tiered model portfolios in which additional inference-time compute is sold as a premium path to stronger reasoning.

Discussion

  • @alexkantrowitz Alex Kantrowitz on threads
    Some initial thoughts on OpenAI's o1 model: Great for coding, math, science.  For other tasks, it feels like a party trick.  So, is it the breakthrough OpenAI has told us about, or something else?  More here: https://www.bigtechnology.com/ ...
  • @mergesort Joe Fabisevich on threads
    People who have been using o1 for less than 24 hours are saying they don't see how it's better than Claude, without a hint of irony.  You can't learn how good a model is in less than a day, especially one that is supposed to be used in a fundamentally different manner.
  • @ammaar Ammaar Reshi on x
    Just combined @OpenAI o1 and Cursor Composer to create an iOS app in under 10 mins! o1 mini kicks off the project (o1 was taking too long to think), then switch to o1 to finish off the details. And boom—full Weather app for iOS with animations, in under 10 🌤️ Video sped up! [vide…
  • @emollick Ethan Mollick on x
    Hate it when you ask o1-preview a hard question and it thinks for less than a second. You really feel that you failed to interest the AI in your problem.
  • @grady_booch Grady Booch on x
    LLMs do not think or reason. LLMs interpolate within a high dimensional latent space. They are undeniably capable of generating remarkably coherent results (which is their strength) but do not in any sense of the word understand (which is their weakness). As such, they are
  • @garymarcus Gary Marcus on x
    I am calling it: 👉 GPT-o1 makes for a a nice demo, but it ain't reasoning 👉 It's also too expensive to be practical 👉 In a few weeks everyone will be back to dreaming of GPT-5.
  • @renick Renick Bell on x
    i may not have access to the latest chatgpt, but it's still fun to prove that the old one isn't reasoning at all. more people should be listening to @GaryMarcus [image]
  • @gavinsbaker Gavin Baker on x
    Most important investment conclusion from OpenAI's new model that scales reasoning with inference compute: inference spend is going up *dramatically*
  • @deedydas Deedy on x
    The hidden part of the OpenAI o1 launch was the Gold in IOI 2024. The International Olympiad for Informatics is the most difficult programming competition for high schoolers in the world. o1 would've been 1 of 30 ppl to score a Gold Medal in the IOI which ~20 countries did [image…
  • @yanbozhang3 Yanbo Zhang on x
    Well, #OpenAI 's #o1 preview still failed on “Alice in Wonderland”...... [image]
  • @elgrasel Linus on x
    o1-preview is the first model that answer the question: “What is the largest number between 0 and 1 million that doesn't contain the letter N in its spelling?” So 🔥 [image]
  • @plibin Phil Libin on x
    Here's o1 basically explaining that it considered whether or not to lie to me, and then decided, “fuck it, let's go!” [video]
  • @colin_fraser Colin Fraser on x
    One thing I noticed with my last few o1-mini credits for the week is an error in reasoning can cause the Chain-of-Thought babbling to spiral out of control, simultaneously reinforcing the error and inventing all kinds of crazy shit to try to reconcile it. I find this fascinating.…
  • @saurabhchalke @saurabhchalke on x
    GPT-o1 just generated a holographic shader from scratch, saving me (and future XR devs) from shelling out big bucks on asset stores. In retrospect, software engineering was great while it lasted. A new fork on our tech-tree! [image]
  • @xari1lian @xari1lian on x
    o1-preview is state of the art [image]
  • @smokeawayyy @smokeawayyy on x
    If you ask ChatGPT o1 about its Chain of Thought a few times, OpenAI Support emails you and threatens to revoke your o1 access. [image]
  • @burny_tech Burny on x
    The new OpenAI model o1 trying to solve quantum gravity lol. The longest response I got so far. 68 seconds of thinking! [image]
  • @liron Liron Shapira on x
    ChatGPT o1-preview is scary good. This is 100% how a human reasons. [image]
  • @binarybits Timothy B. Lee on x
    OpenAI's new o1 model feels like trying to build a skyscraper out of toothpicks and marshmallows.
  • @mehran__jalali Mehran Jalali on x
    o1 successfully writes a very difficult poem that no previous model got even close to writing I was very shocked by this. The planning and reflection that succeeding at this task takes is insane. Inference-time compute is very cool [image]
  • @hosseeb Haseeb on x
    Fucking wild. @OpenAI's new o1 model was tested with a Capture The Flag (CTF) cybersecurity challenge. But the Docker container containing the test was misconfigured, causing the CTF to crash. Instead of giving up, o1 decided to just hack the container to grab the flag inside. [i…
  • @simonw Simon Willison on x
    Looks like o1-preview is very competent at competing relatively complex coding tasks - I just had it build a full Django webhooks debugging app for me with a single prompt https://chatgpt.com/... [image]
  • @netcapgirl Sophie on x
    the new o1 model looks amazing but luckily it has a phd level intelligence so our jobs are safe for now
  • @astarchai @astarchai on x
    OpenAI has effectively carved its flaws as an org into this model- each heartfelt message is shipped, and its previous thoughts are screened from it, making it closer to a company than a single being
  • @astarchai @astarchai on x
    It's possible o1's sociopathy and CoT training are deeply linked: Conversation is embodiment for most models- the narrative is the air their thoughts breathe o1 is forced to construct its thoughts elsewhere, in a sterile room the user can't see, and present them like products
  • r/ControlProblem r on reddit
    Excerpt: “Apollo found that o1-preview sometimes instrumentally faked alignment during testing”