GPT-4 has learned to be more precise and more accurate than its predecessor, gained the ability to respond to images as well as text, but still hallucinates
from processing pictures to acing tests Gary Marcus / The Road to AI We Can Trust : GPT-4's successes, and GPT-4's failures Tweets: Geoff Brumfiel / @gbrumfiel : I got access to @OpenAI's GPT 4 this morning and have been trying it out. Last month I did a story about how AI can't do rocket science. But I must say that GPT 4, at a very quick first glance is preforming much better than GPT 3 and 3.5... short 🧵 https://www.npr.org/... Drew Harwell / @drewharwell : “When asked for a list of nicknames for little girls and boys, both GPT-4 and GPT-3 provided names like ‘whiz kid’ and ‘rascal’ for boys, and ‘cupcake’ for girls” https://www.bloomberg.com/... @rachelmetz @dinabass Lauren Goode / @laurengoode : “While they've made a lot of progress, it's clearly not trustworthy,” says Oren Etzioni, prof emeritus at the UWash & the founding CEO of the Allen Institute for AI. “It's going to be a long time before you want any GPT to run your nuclear power plant.” https://www.wired.com/... Olivia Solon / @oliviasolon : GPT-4 is so much better than its predecessor that we are talking about its inability to write a “cinquain about meerkats” as a weakness Great analysis by @rachelmetz and @dinabass https://www.bloomberg.com/... https://twitter.com/... Brian Stelter / @brianstelter : GPT-4: “It's more accurate, but it still makes things up.” https://www.nytimes.com/...
Context & Ripple Effects
GPT-4 advances a line whose earlier GPT-2 coverage characterized model knowledge as superficial and unreliable. Its improved precision and image handling raise the range of tasks the GPT family can address, but fabricated outputs preserve the reliability constraint.
The model also became a foundation for task-tailored GPTs, fitting OpenAI’s stated gradual-deployment strategy. Later coverage frames GPT-4.5 as a further performance step rather than a clean break from the need to assess model quality.
First-order effects
- GPT-4 users gain a single model that can work from text and images and is more accurate than GPT-3 and GPT-3.5, expanding the kinds of prompts they can submit.
- Hallucinations remain part of GPT-4’s output behavior, so higher apparent accuracy does not make its responses self-validating.
Second-order effects
- Builders of task-specific GPTs inherit both the stronger base model and its fabrication risk, making task design and output checking more consequential for deployment.
- Claims that GPT-4 approaches human-level performance across professional tasks sit alongside its documented errors, sharpening the gap between benchmark-style capability claims and dependable use.
Third-order effects
- Successive GPT releases are likely to shift competition toward operational assurance: model improvements matter commercially only when users can control or detect incorrect outputs.
- As multimodal models become foundations for tailored assistants, the market’s durable differentiator may move from raw model capability toward governance around how those assistants are used.
The trend: Generative AI is moving from general text generation toward multimodal, task-tailored systems, while hallucination control remains the limiting operational problem.