/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

GPT-4 has learned to be more precise and more accurate than its predecessor, gained the ability to respond to images as well as text, but still hallucinates

from processing pictures to acing tests Gary Marcus / The Road to AI We Can Trust : GPT-4's successes, and GPT-4's failures Tweets: Geoff Brumfiel / @gbrumfiel : I got access to @OpenAI's GPT 4 this morning and have been trying it out. Last month I did a story about how AI can't do rocket science. But I must say that GPT 4, at a very quick first glance is preforming much better than GPT 3 and 3.5... short 🧵 https://www.npr.org/... Drew Harwell / @drewharwell : “When asked for a list of nicknames for little girls and boys, both GPT-4 and GPT-3 provided names like ‘whiz kid’ and ‘rascal’ for boys, and ‘cupcake’ for girls” https://www.bloomberg.com/... @rachelmetz @dinabass Lauren Goode / @laurengoode : “While they've made a lot of progress, it's clearly not trustworthy,” says Oren Etzioni, prof emeritus at the UWash & the founding CEO of the Allen Institute for AI. “It's going to be a long time before you want any GPT to run your nuclear power plant.” https://www.wired.com/... Olivia Solon / @oliviasolon : GPT-4 is so much better than its predecessor that we are talking about its inability to write a “cinquain about meerkats” as a weakness Great analysis by @rachelmetz and @dinabass https://www.bloomberg.com/... https://twitter.com/... Brian Stelter / @brianstelter : GPT-4: “It's more accurate, but it still makes things up.” https://www.nytimes.com/...

New York Times

Context & Ripple Effects

GPT-4 advances a line whose earlier GPT-2 coverage characterized model knowledge as superficial and unreliable. Its improved precision and image handling raise the range of tasks the GPT family can address, but fabricated outputs preserve the reliability constraint.

The model also became a foundation for task-tailored GPTs, fitting OpenAI’s stated gradual-deployment strategy. Later coverage frames GPT-4.5 as a further performance step rather than a clean break from the need to assess model quality.

First-order effects

  • GPT-4 users gain a single model that can work from text and images and is more accurate than GPT-3 and GPT-3.5, expanding the kinds of prompts they can submit.
  • Hallucinations remain part of GPT-4’s output behavior, so higher apparent accuracy does not make its responses self-validating.

Second-order effects

  • Builders of task-specific GPTs inherit both the stronger base model and its fabrication risk, making task design and output checking more consequential for deployment.
  • Claims that GPT-4 approaches human-level performance across professional tasks sit alongside its documented errors, sharpening the gap between benchmark-style capability claims and dependable use.

Third-order effects

  • Successive GPT releases are likely to shift competition toward operational assurance: model improvements matter commercially only when users can control or detect incorrect outputs.
  • As multimodal models become foundations for tailored assistants, the market’s durable differentiator may move from raw model capability toward governance around how those assistants are used.

The trend: Generative AI is moving from general text generation toward multimodal, task-tailored systems, while hallucination control remains the limiting operational problem.

Discussion

  • @laurengoode Lauren Goode on x
    “While they've made a lot of progress, it's clearly not trustworthy,” says Oren Etzioni, prof emeritus at the UWash & the founding CEO of the Allen Institute for AI. “It's going to be a long time before you want any GPT to run your nuclear power plant.” https://www.wired.com/...
  • @brianstelter Brian Stelter on x
    GPT-4: “It's more accurate, but it still makes things up.” https://www.nytimes.com/...
  • @chafkin Max Chafkin on x
    in a bunch of these side by side comparisons gpt 3.5 seems the same or better than supposedly “advanced” system https://twitter.com/...
  • @chafkin Max Chafkin on x
    this explainer doesn't make much of a case that gpt-4 is better than gpt-3. and yet https://www.nytimes.com/... https://twitter.com/...
  • @john_bailey John Bailey on x
    GPT-4 with some impressive pass rates of a variety of exams. “We tested GPT-4 on a diverse set of benchmarks, including simulating exams that were originally designed for humans. We did no specific training for these exams.” https://cdn.openai.com/... https://twitter.com/...
  • @npew Peter Welinder on x
    GPT-4 can read images. Still in research preview, but the team is working hard on getting it ready for broader access. https://twitter.com/...
  • @whet Whet Moser on x
    dang, the thing where gpt 4 takes a picture of the inside of a fridge and suggests a meal is pretty wild https://www.nytimes.com/...
  • @shiraovide Shira Ovide on x
    Both of these AI jokes are bad. I feel ok about humans. https://www.nytimes.com/... https://twitter.com/...