A comparison of GPT-4o, Claude 3.7 Sonnet, Gemini 2.0 Flash, Llama 4, and Copilot: Claude won overall, having the most consistent answers and no hallucinations
We challenged AI helpers to decode legal contracts, simplify medical research, speed-read a novel and make sense of Trump speeches. Bluesky: @emilprotalinski . X: @kylebrussell , @ryp__ , and @eilperin Bluesky: Emil Protalinski / @emilprotalinski : The generative AI space moves so quickly that it makes comparisons like this useless. — GPT4-o is from November 2024. — Claude 3.7 is from February 2025. — Gemini 2.0 is from February 2025. — All three of these have newer, much better versions. — The Washington Post should have updated before publishing. … X: Kyle Russell / @kylebrussell : the AI influencer folks absolutely lapping folks at older pubs at the basics of comparing the current offerings Robert Young Pelton / @ryp__ : As a creative person, I am not a fan of ripoff AI. As a writer, it's laughable; as a tool for nonprofessional analysis, it's excellent. Here is a ranking. Note that only Claude is drug-free. https://www.washingtonpost.com/ ... [image] Juliet Eilperin / @eilperin : .@geoffreyfowler and other @washingtonpost reporters gave 5 chatbots a reading test, on everything from literature to legal documents. See how they did - the answer might surprise you https://www.washingtonpost.com/ ...
Context & Ripple Effects
This is the latest in a run of consumer-facing model comparisons: a 2024 test of several assistants ranked Perplexity first overall, while earlier GPT-4 coverage emphasized that improved accuracy did not eliminate hallucinations. The new result instead rewards consistency across practical reading and interpretation tasks.
Its shelf life is unusually short. The comparison itself was criticized for testing versions that had already been superseded, even as OpenAI had positioned GPT-4.5 as a newer research-preview model after GPT-4o.
First-order effects
- Claude 3.7 Sonnet gains a favorable third-party signal for users choosing an assistant for contract, research, long-document, and political-speech interpretation tasks; the test found it most consistent and reported no hallucinations.
- GPT-4o, Gemini 2.0 Flash, Llama 4, and Copilot are placed behind Claude in this specific evaluation, but the reported version gap limits how broadly that ranking can be applied.
Second-order effects
- Model vendors face more pressure to demonstrate reliability on end-user workflows rather than rely on broad capability claims; the prior cross-assistant comparison that ranked Perplexity first shows how quickly public league tables can change.
- Buyers evaluating AI assistants will need to treat brand-level rankings as version- and task-specific, making their own validation more important where errors in legal or medical summarization matter.
Third-order effects
- If rapid model replacement continues, generalized “best chatbot” tests will become less durable, shifting competition toward repeatable, version-dated evaluations tied to useful workflow outcomes.
- The durable differentiator may be reliability per task, not a single aggregate leaderboard result—a shift toward [[a:concepts#ai-cost-per-useful-task|AI cost per useful task]] and [[a:concepts#workflow-native-ai|workflow-native AI]] purchasing criteria.
The trend: Consumer AI competition is moving from headline model releases toward continual, task-specific evidence of reliability and usefulness.