/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

OpenAI's o1 models are dramatically better at reasoning than previous LLMs, but they struggle with spatial reasoning and are far from human-level intelligence

Timothy B Lee / Understanding AI :

Understanding AI Timothy B Lee

Discussion

  • Vox Kelsey Piper on x
    What it means that new AIs can “reason”
  • @benjedwards Benj Edwards on x
    I've created a new fruit-based benchmark for LLMs: “How many Rs are *not* in the word strawberry?” 😁 See how o1-preview fares vs. GPT-4o in the screenshots below (cc:@goodside) [image]
  • @garrisonlovely Garrison Lovely on x
    I agree that progress clustered around GPT-4 level models for a while (and thought it might be evidence of a wall), and agree with this Timothy's assessment here. Also good on him for saying so! It's very common for people to paint themselves into corners and not budge.
  • @binarybits Timothy B. Lee on x
    Over the last 9 months I developed a suite of reasoning puzzles to test the capabilities of new frontier models. o1 aced every single one of them, forcing me to come up with new ones.
  • @jaesf Jacob Eliosoff on x
    Sometimes you have to step back and acknowledge that even just THREE years ago, I would have thought it was TOTALLY INSANE that a chatbot would be able to solve these problems, never mind so soon. Anyway check out TBL's fuller writeup linked further down in the thread.
  • @binarybits Timothy B. Lee on x
    Here's a problem that's challenging because it requires trial and error. GPT-4o gets stuck and gives up. o1-preview got the right answer. [image]
  • @binarybits Timothy B. Lee on x
    For example, I tested the models on long word problems like this. GPT-4o can keep track of how many marbles are in each jar up to about 50 steps, but gets confused by 70. o1-preview gets it right up to about 200 steps. [image]
  • @binarybits Timothy B. Lee on x
    I've developed a bit of a reputation as an “AI skeptic,” but I think I was just accurately reporting on the slow pace of LLM progress following GPT-4. o1 is a totally different story. It's by far the biggest jump in performance since GPT-4.
  • @sporadicalia @sporadicalia on x
    o1 is absolutely melting my mind and it should be melting yours too we live at the most exciting time in human history
  • @binarybits Timothy B. Lee on x
    The one big blind spot I found is that o1 is bad at spatial reasoning. o1 can't accept images yet, but I gave it a word problem describing a set of streets like this. The brown boxes show streets that are closed. o1-preview recommended the following invalid route. [image]
  • @mpopv Matt Popovich on x
    I had honestly completely forgotten that it was an early version of o1 that sparked the board coup. Looks quite silly/odd in retrospect now that it's been released to a relatively muted response from the safety crowd compared to previous high-profile releases. [image]