OpenAI's o1 models are dramatically better at reasoning than previous LLMs, but they struggle with spatial reasoning and are far from human-level intelligence
Timothy B Lee / Understanding AI :
Understanding AI Timothy B Lee
Related Coverage
Discussion
-
Vox
Kelsey Piper
on x
What it means that new AIs can “reason”
-
@benjedwards
Benj Edwards
on x
I've created a new fruit-based benchmark for LLMs: “How many Rs are *not* in the word strawberry?” 😁 See how o1-preview fares vs. GPT-4o in the screenshots below (cc:@goodside) [image]
-
@garrisonlovely
Garrison Lovely
on x
I agree that progress clustered around GPT-4 level models for a while (and thought it might be evidence of a wall), and agree with this Timothy's assessment here. Also good on him for saying so! It's very common for people to paint themselves into corners and not budge.
-
@binarybits
Timothy B. Lee
on x
Over the last 9 months I developed a suite of reasoning puzzles to test the capabilities of new frontier models. o1 aced every single one of them, forcing me to come up with new ones.
-
@jaesf
Jacob Eliosoff
on x
Sometimes you have to step back and acknowledge that even just THREE years ago, I would have thought it was TOTALLY INSANE that a chatbot would be able to solve these problems, never mind so soon. Anyway check out TBL's fuller writeup linked further down in the thread.
-
@binarybits
Timothy B. Lee
on x
Here's a problem that's challenging because it requires trial and error. GPT-4o gets stuck and gives up. o1-preview got the right answer. [image]
-
@binarybits
Timothy B. Lee
on x
For example, I tested the models on long word problems like this. GPT-4o can keep track of how many marbles are in each jar up to about 50 steps, but gets confused by 70. o1-preview gets it right up to about 200 steps. [image]
-
@binarybits
Timothy B. Lee
on x
I've developed a bit of a reputation as an “AI skeptic,” but I think I was just accurately reporting on the slow pace of LLM progress following GPT-4. o1 is a totally different story. It's by far the biggest jump in performance since GPT-4.
-
@sporadicalia
@sporadicalia
on x
o1 is absolutely melting my mind and it should be melting yours too we live at the most exciting time in human history
-
@binarybits
Timothy B. Lee
on x
The one big blind spot I found is that o1 is bad at spatial reasoning. o1 can't accept images yet, but I gave it a word problem describing a set of streets like this. The brown boxes show streets that are closed. o1-preview recommended the following invalid route. [image]
-
@mpopv
Matt Popovich
on x
I had honestly completely forgotten that it was an early version of o1 that sparked the board coup. Looks quite silly/odd in retrospect now that it's been released to a relatively muted response from the safety crowd compared to previous high-profile releases. [image]