OpenAI's o1 models are dramatically better at reasoning than previous LLMs, but they struggle with spatial reasoning and are far from human-level intelligence
o1 arrived with a sharp benchmark contrast: OpenAI said it solved far more problems on an IMO qualifying exam than GPT-4o, a claim covered in the early math-benchmark comparison. This assessment adds an important boundary to that performance narrative by distinguishing reasoning gains from broader capability.
Related coverage also framed o1 as a break from a simple model-generation upgrade, with cost and performance trade-offs behind its reasoning approach. That makes its spatial weakness consequential for deciding which workloads can safely benefit from the new capability.
First-order effects
OpenAI gains a clearer differentiation point in tasks where extended reasoning helps, while users must treat spatially dependent work as a material limitation rather than assuming a general upgrade.
Teams evaluating o1 face a more selective deployment choice: improved reasoning may justify use on suitable problems, but it does not establish human-level competence across tasks.
Second-order effects
Competing model providers are pushed to demonstrate not just headline reasoning benchmarks but performance across capability types, including spatial tasks and reliability-sensitive use cases.
The reported trade-offs make model selection more workload-specific, favoring evaluations that weigh reasoning improvements against cost and performance rather than a single overall ranking.
Third-order effects
If reasoning-focused systems continue to improve unevenly, the market is likely to segment around specialized strengths instead of converging quickly on one broadly human-equivalent model.
Progress will increasingly be judged by capability coverage and failure modes, not only by standout test results; later concerns over higher hallucination rates in newer reasoning models illustrate why that broader scrutiny matters.
The trend: AI development is shifting from general next-model upgrades toward reasoning-oriented systems whose value depends on task-specific gains, costs, and persistent blind spots.
I've created a new fruit-based benchmark for LLMs: “How many Rs are *not* in the word strawberry?” 😁 See how o1-preview fares vs. GPT-4o in the screenshots below (cc:@goodside) [image]
I agree that progress clustered around GPT-4 level models for a while (and thought it might be evidence of a wall), and agree with this Timothy's assessment here. Also good on him for saying so! It's very common for people to paint themselves into corners and not budge.
Over the last 9 months I developed a suite of reasoning puzzles to test the capabilities of new frontier models. o1 aced every single one of them, forcing me to come up with new ones.
Sometimes you have to step back and acknowledge that even just THREE years ago, I would have thought it was TOTALLY INSANE that a chatbot would be able to solve these problems, never mind so soon. Anyway check out TBL's fuller writeup linked further down in the thread.
For example, I tested the models on long word problems like this. GPT-4o can keep track of how many marbles are in each jar up to about 50 steps, but gets confused by 70. o1-preview gets it right up to about 200 steps. [image]
I've developed a bit of a reputation as an “AI skeptic,” but I think I was just accurately reporting on the slow pace of LLM progress following GPT-4. o1 is a totally different story. It's by far the biggest jump in performance since GPT-4.
The one big blind spot I found is that o1 is bad at spatial reasoning. o1 can't accept images yet, but I gave it a word problem describing a set of streets like this. The brown boxes show streets that are closed. o1-preview recommended the following invalid route. [image]
I had honestly completely forgotten that it was an early version of o1 that sparked the board coup. Looks quite silly/odd in retrospect now that it's been released to a relatively muted response from the safety crowd compared to previous high-profile releases. [image]