/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

OpenAI's o1 models are dramatically better at reasoning than previous LLMs, but they struggle with spatial reasoning and are far from human-level intelligence

Timothy B Lee / Understanding AI :

Understanding AI Timothy B Lee

Context & Ripple Effects

o1 arrived with a sharp benchmark contrast: OpenAI said it solved far more problems on an IMO qualifying exam than GPT-4o, a claim covered in the early math-benchmark comparison. This assessment adds an important boundary to that performance narrative by distinguishing reasoning gains from broader capability.

Related coverage also framed o1 as a break from a simple model-generation upgrade, with cost and performance trade-offs behind its reasoning approach. That makes its spatial weakness consequential for deciding which workloads can safely benefit from the new capability.

First-order effects

  • OpenAI gains a clearer differentiation point in tasks where extended reasoning helps, while users must treat spatially dependent work as a material limitation rather than assuming a general upgrade.
  • Teams evaluating o1 face a more selective deployment choice: improved reasoning may justify use on suitable problems, but it does not establish human-level competence across tasks.

Second-order effects

  • Competing model providers are pushed to demonstrate not just headline reasoning benchmarks but performance across capability types, including spatial tasks and reliability-sensitive use cases.
  • The reported trade-offs make model selection more workload-specific, favoring evaluations that weigh reasoning improvements against cost and performance rather than a single overall ranking.

Third-order effects

  • If reasoning-focused systems continue to improve unevenly, the market is likely to segment around specialized strengths instead of converging quickly on one broadly human-equivalent model.
  • Progress will increasingly be judged by capability coverage and failure modes, not only by standout test results; later concerns over higher hallucination rates in newer reasoning models illustrate why that broader scrutiny matters.

The trend: AI development is shifting from general next-model upgrades toward reasoning-oriented systems whose value depends on task-specific gains, costs, and persistent blind spots.

Discussion

  • Vox Kelsey Piper on x
    What it means that new AIs can “reason”
  • @benjedwards Benj Edwards on x
    I've created a new fruit-based benchmark for LLMs: “How many Rs are *not* in the word strawberry?” 😁 See how o1-preview fares vs. GPT-4o in the screenshots below (cc:@goodside) [image]
  • @garrisonlovely Garrison Lovely on x
    I agree that progress clustered around GPT-4 level models for a while (and thought it might be evidence of a wall), and agree with this Timothy's assessment here. Also good on him for saying so! It's very common for people to paint themselves into corners and not budge.
  • @binarybits Timothy B. Lee on x
    Over the last 9 months I developed a suite of reasoning puzzles to test the capabilities of new frontier models. o1 aced every single one of them, forcing me to come up with new ones.
  • @jaesf Jacob Eliosoff on x
    Sometimes you have to step back and acknowledge that even just THREE years ago, I would have thought it was TOTALLY INSANE that a chatbot would be able to solve these problems, never mind so soon. Anyway check out TBL's fuller writeup linked further down in the thread.
  • @binarybits Timothy B. Lee on x
    Here's a problem that's challenging because it requires trial and error. GPT-4o gets stuck and gives up. o1-preview got the right answer. [image]
  • @binarybits Timothy B. Lee on x
    For example, I tested the models on long word problems like this. GPT-4o can keep track of how many marbles are in each jar up to about 50 steps, but gets confused by 70. o1-preview gets it right up to about 200 steps. [image]
  • @binarybits Timothy B. Lee on x
    I've developed a bit of a reputation as an “AI skeptic,” but I think I was just accurately reporting on the slow pace of LLM progress following GPT-4. o1 is a totally different story. It's by far the biggest jump in performance since GPT-4.
  • @sporadicalia @sporadicalia on x
    o1 is absolutely melting my mind and it should be melting yours too we live at the most exciting time in human history
  • @binarybits Timothy B. Lee on x
    The one big blind spot I found is that o1 is bad at spatial reasoning. o1 can't accept images yet, but I gave it a word problem describing a set of streets like this. The brown boxes show streets that are closed. o1-preview recommended the following invalid route. [image]
  • @mpopv Matt Popovich on x
    I had honestly completely forgotten that it was an early version of o1 that sparked the board coup. Looks quite silly/odd in retrospect now that it's been released to a relatively muted response from the safety crowd compared to previous high-profile releases. [image]