/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

Researchers show that ChatGPT 3.5 outperforms ChatGPT 4 in many tasks, including solving math problems, highlighting the issue of “drift” in improving AI models

Josh Zumbrun / Wall Street Journal :

Wall Street Journal Josh Zumbrun

Context & Ripple Effects

OpenAI introduced GPT-4 as surpassing ChatGPT on advanced reasoning, making this finding a meaningful qualification of the assumption that a newer flagship model is uniformly better. The reported gap instead puts task-level evaluation alongside broad capability claims.

Later coverage reinforces that performance is uneven by task and vintage: a GPT-3.5 coding assessment found stronger results on problems predating 2021 than on newer ones. Model updates, including GPT-4 Turbo’s revised response style, can therefore alter the practical behavior customers depend on.

First-order effects

  • Teams using ChatGPT for math or other affected workflows cannot treat GPT-4 as an automatic upgrade over GPT-3.5; they need to test the versions against their own task sets.
  • OpenAI’s headline reasoning comparison becomes less sufficient for buyers deciding which model to deploy, because reported performance varies across tasks.

Second-order effects

  • Application vendors face pressure to add regression tests and potentially route different requests to different model versions rather than standardize on a single newest release.
  • Benchmark results and product-update messaging become less decisive for customers, increasing the value of evaluations tied to a customer’s specific workload and quality threshold.

Third-order effects

  • If drift persists, frontier-model competition will be measured less by a simple version ladder and more by reliability, reproducibility, and performance on defined use cases.
  • The market may move toward task-specific model portfolios and evaluation infrastructure, with the economically relevant metric becoming useful output rather than nominal model generation.

The trend: Generative-AI adoption is shifting from choosing the newest general model to continuously validating the model-version combination that performs best for each production task.

Discussion

  • @chr1sa Chris Anderson on x
    This article doesn't give a good reason for ChatGPT accuracy getting worse other than “drift”. I'd guess that as the service got popular the cost of running it rose, so they had to quantize or prune the model to get it to run faster at inference time https://www.wsj.com/...
  • @robinhanson Robin Hanson on x
    “phenomenon known to AI developers as drift, where attempts to improve one part of the enormously complex AI models make other parts of the models perform worse.” Sounds like software rot. https://www.wsj.com/...