OpenAI says its new o3 and o4-mini AI models hallucinate more often than its previous reasoning and traditional models, and the company doesn't know why
OpenAI's recently launched o3 and o4-mini AI models are state-of-the-art in many respects. However, the new models still hallucinate …
TechCrunch Maxwell Zeff
Context & Ripple Effects
OpenAI positioned the o-series around models that reason before answering, beginning with its December o3 and o3-mini unveiling. It subsequently marketed o3-mini as a faster, lower-cost option with capabilities broadly comparable to earlier o-series models.
The newly released o3 and o4-mini therefore complicate a product arc centered on stronger reasoning and more accessible performance: capability gains do not automatically translate into more dependable outputs.
First-order effects
- OpenAI must account for an acknowledged reliability regression in its newest reasoning offerings while it lacks an explanation for the change.
- Paid users of the newly launched models need to treat their answers as requiring verification, particularly where an incorrect factual claim would matter.
Second-order effects
- Model buyers gain a clearer reason to compare systems on hallucination behavior, not only reasoning benchmarks, speed, or cost—the dimensions emphasized in the earlier o3-mini positioning.
- Competitors can use demonstrable reliability and transparent evaluation as a point of differentiation when customers are deciding whether newer reasoning models are suitable for production work.
Third-order effects
- If more capable releases continue to show uneven factual reliability, AI deployment will increasingly depend on assurance layers—evaluation, human review, and task-specific guardrails—rather than model version alone.
- The episode strengthens pressure for evaluations that reward uncertainty and calibrated refusal as well as correct answers; whether that changes model-training incentives remains unresolved.
The trend: Reasoning-model competition is shifting from headline capability toward operational trustworthiness and the systems used to measure it.
Related: Operational AI assurance · Operational AI governance · AI industrialization · OpenAI launches o3-mini · OpenAI unveils o3 and o3-mini
Related Coverage
- Investigating truthfulness in a pre-release o3 model Transluce
- OpenAI o3 and o4-mini System Card OpenAI
- Powered by Smarter, but less accurate? ChatGPT's hallucination conundrum The Economic Times
- OpenAI New o3/o4-mini Models Hallucinate More Than Previous Models WinBuzzer · Markus Kasanmascheff
- OpenAI's new reasoning models see rise in hallucination rates Tech in Asia · Minh Le
- OpenAI's latest AI models are smarter, but they make things up more often. Here's what we know Livemint · Aman Gupta
- Why software developers need to watch out for package hallucinations IT Brew · Brianna Monsanto
- OpenAI's New AI Models o3 and o4-mini Can Now ‘Think With Images’ TechRepublic · Aminu Abdullahi
- Breakthroughs, Concerns in OpenAI's Latest Lineup HealthcareInfoSecurity.com · Rashmi Ramesh
- Vibe Check: OpenAI's o3, GPT-4.1, and o4-mini Every
- What to know about o3 and o4-mini, OpenAI's new reasoning models TechTalks · Ben Dickson
- “OpenAI found that o3 hallucinated in response to 33% of questions on PersonQA, the company's in-house benchmark for measuring the accuracy of a model's knowledge about people. That's roughly double the hallucination rate of OpenAI's previous reasoning models, o1 and o3-mini, which scored 16% and 14.8%, respectively. … @aulia@mementomori.social · Aulia Masna
- OpenAI's new reasoning AI models hallucinate more Hacker News
- OpenAI Puzzled as New Models Show Rising Hallucination Rates Slashdot · Msmash
- ChatGPT is referring to users by their names unprompted, and some find it ‘creepy’ TechCrunch · Kyle Wiggers
- New ChatGPT feature uses your past chats for smarter web searches Moneycontrol
- “Creepy” ChatGPT calls users by name—even when it shouldn't The Economic Times
- OpenAI o3, o4-mini, and o3-mini Usage Limits on ChatGPT and the API OpenAI
- ChatGPT's latest image trend? How Redditors are turning usernames into viral arrest records — and they're too good Livemint · Aman Gupta
- All the AI news of the week: ChatGPT debuts o3 and o4-mini, Gemini talks to dolphins Mashable · Cecily Mauran
- ChatGPT Can Now Guess the Location of Photos With Stunning Accuracy PCMag · Emily Forlini
- ChatGPT Is Scary Good at Guessing the Location of a Photo PetaPixel · Matt Growcoot
- OpenAI details ChatGPT-o3, o4-mini, o4-mini-high usage limits BleepingComputer · Mayank Parmar
- ChatGPT's Memory Now Personalizes Web Searches WinBuzzer · Markus Kasanmascheff
- Think GeoGuessr is fun? Try using ChatGPT to guess locations in your photos ZDNET · Elyse Betters Picaro
- OpenAI Says Its Latest Models Bring More Capable AI Agents to Business PYMNTS.com
- OpenAI's latest AI models can ‘think with images’ and combine tools PCWorld
- You can't hide from ChatGPT - new viral AI challenge can geo-locate you from almost any photo - we tried it and it's wild and worrisome TechRadar · Lance Ulanoff
- ChatGPT can now guess where a photo was taken, which is slightly terrifying BGR · Joshua Hawkins
- OpenAI's New Models Could Be Its Smartest and Most Powerful, Thanks to This New Feature Inc · Ben Sherry
- OpenAI's o3 can use images while reasoning Android Headlines · Arthur Brown
- ChatGPT Can Now Reason Using the Images You Upload: Why This Is Amazing MakeUseOf · Saikat Basu
- “The latest viral ChatGPT trend is doing ‘reverse location search’ from photos” — This kind of thing has been going on for a long time, but it did take some skill with available tools. It's not immediately clear to me whether easy access is good or bad. — https://techcrunch.com/... @John@socks.masto.host · John Socks
- Viral ChatGPT trend is doing ‘reverse location search’ from photos Hacker News
- ChatGPT Models Are Surprisingly Good At Geoguessing Slashdot · BeauHD
Discussion
-
@smcgrath.phd
Scott McGrath
on bluesky
OpenAI's new “reasoning” models (o3 and o4-mini) actually hallucinate MORE than their predecessors — OpenAI's internal tests show o3 hallucinated on 33% of person-related questions, double the rate of previous models. Even worse, o4-mini hit 48%.
-
@twtzero_
Natsuki
on x
alright party's over [image]
-
@transluceai
@transluceai
on x
We tested a pre-release version of o3 and found that it frequently fabricates actions it never took, and then elaborately justifies these actions when confronted. We were surprised, so we dug deeper 🔎🧵(1/) https://x.com/... [image]
-
@emollick
Ethan Mollick
on x
A potential issue with o3 is that it thinks it is using tools even when it does not, leading to some hallucinations where it assumes work that was implied in the reasoning chain was actually done. You should double check the reasoning trace for complex work to see what it did.
-
@transluceai
@transluceai
on x
These behaviors are surprising. It seems that despite being incredibly powerful at solving math and coding tasks, o3 is not by default truthful about its capabilities. (12/)
-
@peterwildeford
Peter Wildeford
on x
Great thread. o3 makes meaningful progress for mathematical applications, excelling at undergrad problems and basic tool use... but still struggles with research-level mathematics, proof construction, and avoiding hallucinations.
-
@angie_rasmussen
Dr. Angela Rasmussen
on x
The model has improved and is now capable of making shit up about why it made shit up
-
@chowdhuryneil
Neil Chowdhury
on x
Another transcript: o3 confidently claims it executed code and defends its incorrect calculations. First, o3 tells me about its Python sandbox (which it does not have access to!) 🧵 (1/) [image]
-
@ryan_t_lowe
Ryan Lowe
on x
o3 seems to hallucinate >2x more than o1, according to the system card so hallucinations could scale *inversely* with increased reasoning (unlike for increased model size), bc outcome-based optimization incentivizes confident guessing (the Transluce example is kinda hilarious) [i…
-
@littmath
Daniel Litt
on x
First impressions of o3/o4-mini for math: tool use is really great; *lots* of hallucinations; underlying reasoning is maybe slightly better than o1/o3-mini or gemini 2.5 pro but I'm not confident about this.
-
@aashaysachdeva
Aashay Sachdeva
on x
o-series now mixes tool usage during training. This is definitely an issue stemming from complexity of training with multi-tool setup. Model is now hallucinating in the tools space. The AI safety research also got a lot more interesting - llms with tools have ability to [image]
-
@natolambert
Nathan Lambert
on x
reasoning models are kind of yolo and brining the fun back to AI caveat: lots of ways we dont know what happens when theyre in the world
-
@modestproposal1
@modestproposal1
on x
man do some humans need this
-
@dorialexander
Alexander Doria
on x
Very insightful early tests of o3 showing that a top frontier “PhD-level” model remain totally unreliable for a wide variety of mundane tasks.
-
r/artificial
r
on reddit
OpenAI's new reasoning AI models hallucinate more
-
r/technology
r
on reddit
OpenAI Puzzled as New Models Show Rising Hallucination Rates
-
r/BetterOffline
r
on reddit
OpenAI's new reasoning AI models hallucinate more | TechCrunch
-
r/singularity
r
on reddit
OpenAI's new reasoning AI models hallucinate more | TechCrunch
-
r/OpenAI
r
on reddit
OpenAI's new reasoning AI models hallucinate more
-
r/ControlProblem
r
on reddit
Researchers find pre-release of OpenAI o3 model lies and then invents cover story
-
@mikebutcher
Mike Butcher
on bluesky
People are using ChatGPT's image recognition to figure out the location shown in pictures. That means someone could screenshot say, a person's Instagram Story, and work out where they are. And it's highly accurate. Huge safety issue. — techcrunch.com/2025/04/17/t...
-
@timely06
@timely06
on bluesky
What privacy? In 21st century? Fortunes have been built on the systematic violation of privacy rights disguised as free services. Now what are we talking about? The horse has already left the barn.
-
@hypervisible
@hypervisible
on bluesky
Again: the most prominent use cases are some form of fraud or abuse, and it's not even close.
-
@emollick
Ethan Mollick
on x
The geoguessing power of o3 is a really good sample of its agentic abilities. Between its smart guessing and its ability to zoom into images, to do web searches, and read text, the results can be very freaky. I stripped location info from the photo & prompted “geoguess this” [ima…
-
@deedydas
Deedy
on x
o3 really blew my mind with this one. I gave it an image of a menu of my favorite Chinese place in SF with no title or EXIF data, and it was able to search the web, match menu items, and locate it. 🤯 [image]
-
@swax
@swax
on x
@emollick Wow, nailed it and not even a tree in sight. [image]
-
@vyrotek
Jason Barnes
on x
this is a fun ChatGPT o3 feature. geoguessr! [image]
-
@k_kohlbrenner
Kohl
on x
@emollick Nailed this and others. Very impressive [image]
-
@izyuuumi
Yumi
on x
o3 is insane I asked a friend of mine to give me a random photo They gave me a random photo they took in a library o3 knows it in 20 seconds and it's right [image]
-
@grantslatton
Grant Slatton
on x
@arithmoquine another one i think if there are basically any mountain ranges, it's got enough data to figure it out it did fail on one of my pics from puerto rico, thinking it was hawaii [image]
-
@neilsuperduper
Neil Chudleigh
on x
@AviSchiffmann it wasnt right in the end, the picture is from southern mexico 1000s of km away but still goddamn. its reasoning and image analysis is incredible. [image]
-
@itsurboyevan
Evan Armstrong
on x
o3 is able to geolocate you with a photo in the middle of the wilderness. this thing is spooky smart. [image]
-
@matthewberman
@matthewberman
on x
OMG o3 is insane. No one is safe lol. [image]