/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

A comparison of GPT-4o, Claude 3.7 Sonnet, Gemini 2.0 Flash, Llama 4, and Copilot: Claude won overall, having the most consistent answers and no hallucinations

We challenged AI helpers to decode legal contracts, simplify medical research, speed-read a novel and make sense of Trump speeches. Bluesky: @emilprotalinski . X: @kylebrussell , @ryp__ , and @eilperin Bluesky: Emil Protalinski / @emilprotalinski : The generative AI space moves so quickly that it makes comparisons like this useless.  —  GPT4-o is from November 2024.  —  Claude 3.7 is from February 2025.  —  Gemini 2.0 is from February 2025.  —  All three of these have newer, much better versions.  —  The Washington Post should have updated before publishing. … X: Kyle Russell / @kylebrussell : the AI influencer folks absolutely lapping folks at older pubs at the basics of comparing the current offerings Robert Young Pelton / @ryp__ : As a creative person, I am not a fan of ripoff AI. As a writer, it's laughable; as a tool for nonprofessional analysis, it's excellent. Here is a ranking. Note that only Claude is drug-free. https://www.washingtonpost.com/ ... [image] Juliet Eilperin / @eilperin : .@geoffreyfowler and other @washingtonpost reporters gave 5 chatbots a reading test, on everything from literature to legal documents. See how they did - the answer might surprise you https://www.washingtonpost.com/ ...

Washington Post Geoffrey A. Fowler

Context & Ripple Effects

This is the latest in a run of consumer-facing model comparisons: a 2024 test of several assistants ranked Perplexity first overall, while earlier GPT-4 coverage emphasized that improved accuracy did not eliminate hallucinations. The new result instead rewards consistency across practical reading and interpretation tasks.

Its shelf life is unusually short. The comparison itself was criticized for testing versions that had already been superseded, even as OpenAI had positioned GPT-4.5 as a newer research-preview model after GPT-4o.

First-order effects

  • Claude 3.7 Sonnet gains a favorable third-party signal for users choosing an assistant for contract, research, long-document, and political-speech interpretation tasks; the test found it most consistent and reported no hallucinations.
  • GPT-4o, Gemini 2.0 Flash, Llama 4, and Copilot are placed behind Claude in this specific evaluation, but the reported version gap limits how broadly that ranking can be applied.

Second-order effects

  • Model vendors face more pressure to demonstrate reliability on end-user workflows rather than rely on broad capability claims; the prior cross-assistant comparison that ranked Perplexity first shows how quickly public league tables can change.
  • Buyers evaluating AI assistants will need to treat brand-level rankings as version- and task-specific, making their own validation more important where errors in legal or medical summarization matter.

Third-order effects

  • If rapid model replacement continues, generalized “best chatbot” tests will become less durable, shifting competition toward repeatable, version-dated evaluations tied to useful workflow outcomes.
  • The durable differentiator may be reliability per task, not a single aggregate leaderboard result—a shift toward [[a:concepts#ai-cost-per-useful-task|AI cost per useful task]] and [[a:concepts#workflow-native-ai|workflow-native AI]] purchasing criteria.

The trend: Consumer AI competition is moving from headline model releases toward continual, task-specific evidence of reliability and usefulness.

Discussion

  • @emilprotalinski Emil Protalinski on bluesky
    The generative AI space moves so quickly that it makes comparisons like this useless.  —  GPT4-o is from November 2024.  —  Claude 3.7 is from February 2025.  —  Gemini 2.0 is from February 2025.  —  All three of these have newer, much better versions.  —  The Washington Post sho…
  • @kylebrussell Kyle Russell on x
    the AI influencer folks absolutely lapping folks at older pubs at the basics of comparing the current offerings
  • @ryp__ Robert Young Pelton on x
    As a creative person, I am not a fan of ripoff AI. As a writer, it's laughable; as a tool for nonprofessional analysis, it's excellent. Here is a ranking. Note that only Claude is drug-free. https://www.washingtonpost.com/ ... [image]
  • @eilperin Juliet Eilperin on x
    .@geoffreyfowler and other @washingtonpost reporters gave 5 chatbots a reading test, on everything from literature to legal documents. See how they did - the answer might surprise you https://www.washingtonpost.com/ ...