/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

An analysis of responses by ChatGPT, Claude, Copilot, Gemini, and Perplexity to some real-life questions and everyday tasks; Perplexity ranked first overall

Wall Street Journal

Context & Ripple Effects

This comparison comes as assistants move from novelty tools into practical work surfaces: earlier coverage documented employees using ChatGPT and peers to accelerate tasks, while also exposing persistent reliability limits such as confidently wrong arithmetic answers.

Its result is a point-in-time usability benchmark rather than a settled hierarchy. A later cross-model test instead found Claude strongest on consistency and hallucinations, underscoring how quickly comparative leadership can change as products are updated.

First-order effects

  • Perplexity gains a third-party validation point for everyday-use positioning, while ChatGPT, Claude, Copilot, and Gemini face an unfavorable comparison on the tested tasks.
  • Users and buyers evaluating general-purpose assistants have another reason to distinguish products by task performance rather than treating the major chatbots as interchangeable.

Second-order effects

  • Rivals are pressured to improve the practical workflow elements that shape evaluations—answer quality, reliability, and usefulness on ordinary requests—rather than competing solely on model branding.
  • Independent comparisons become more consequential in user acquisition: the corpus shows subsequent benchmarks can shift the leader, as in a later test that favored Claude's consistency.

Third-order effects

  • If repeatable task-based testing becomes a common buying input, consumer AI competition is likely to fragment by use case instead of consolidating around one perceived best chatbot.
  • The durable constraint is trust: prior reporting on misinformation in ChatGPT programming answers suggests benchmark wins will matter most when products can demonstrate dependable performance across domains.

The trend: General-purpose AI assistants are increasingly competing on measured everyday usefulness and reliability, with leadership liable to rotate as models and product experiences change.

Discussion

  • @ociubotaru Oleg Ciubotaru on threads
    Well this is quite an intro for a paper... Adding to my reading list.  “Even without any narrative or industry- specific information, the LLM outperforms financial analysts in its ability to predict earnings changes.  The LLM exhibits a relative advantage over human analysts in s…
  • @jaspar.bsky.social @jaspar.bsky.social on bluesky
    literally every time I've fact-checked perplexity it's gotten major details wrong [embedded post]
  • @arrington @arrington on x
    Yep, @perplexity_ai wins.
  • @karadapena Kara Dapena on x
    WSJ tested five top AI chatbots in real-world, everyday tasks. Did the famous ChatGPT come out on top? Or was it bested by Google, Microsoft or one of the AI startups? Have a look. https://www.wsj.com/... via @WSJ w/@dalvin_brown @JoannaStern @wjrothman @CloudberryRoque [image]
  • @khalifahmanaa Khalifa Manaa on x
    Comparison between Perplexity and the others is a bit funny, but overall my experience are aligned with this. Perplexity is great as an answering machine. For everything else chatGPT, Gemini & Claude. I don't know what Microsoft did with co-pilot but it's somehow bad with
  • @firstadopter Tae Kim on x
    @AravSrinivas @perplexity_ai A scrappy startup already has a more effective product for daily life than the well-known 700+ person ChatGPT maker and tech giants. Tens of billions of funding aren't everything. Innovation and competition are grand. https://www.wsj.com/... [image]
  • @melindabchu1 Melinda B. Chu on x
    @AravSrinivas I like Perplexity a lot, but it shouldn't have been on this survey. It's not the same as the general purpose LLMs; has less overall capabilities. Perplexity should be compared it to just Google since you guys are just Search, which you talk and post about a lot.
  • @aravsrinivas Aravind Srinivas on x
    Perplexity has been ranked the number one AI chatbot in a survey run by Wall Street Journal, ahead of ChatGPT, Gemini. Microsoft Copilot is the least preferred. [image]
  • @aravsrinivas Aravind Srinivas on x
    Perplexity has been ranked the number one AI chatbot in a survey run by Wall Street Journal, ahead of ChatGPT, Gemini. Microsoft Copilot is the least preferred. [image]
  • @denisyarats Denis Yarats on x
    now @WSJ runs LLM evals too 🤯 https://www.wsj.com/... [image]