/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

An analysis of ChatGPT, Claude, Copilot, Gemini, and Perplexity's responses to some real-life questions and everyday tasks: Perplexity ranked first overall

We tested OpenAI's ChatGPT against Microsoft's Copilot and Google's Gemini, along with Perplexity and Anthropic's Claude.  Here's how they ranked.

Wall Street Journal

Context & Ripple Effects

This comparison extends an earlier benchmark in which ChatGPT led Gemini-powered Bard across several task types, while the margin had already narrowed. The result makes everyday-task performance—not just model capability claims—a visible point of differentiation among consumer assistants.

Later coverage reinforces that benchmark leadership can move: a subsequent cross-model evaluation favored Claude for consistency, while usage-share data showed a more contested assistant market. The significance is less a permanent ranking than evidence that users and buyers have multiple credible options to compare.

First-order effects

  • Perplexity gains an independent, task-based validation against ChatGPT, Copilot, Gemini and Claude, potentially strengthening its positioning with users evaluating assistants on practical usefulness.
  • The other providers face a clear comparative reference point in a category where perceived answer quality and task completion are central to product choice.

Second-order effects

  • Assistant vendors have added incentive to optimize for repeatable real-world workflows and to communicate performance in user-facing terms, rather than relying solely on broad model announcements.
  • For users and enterprise buyers, side-by-side testing lowers the cost of considering specialist or smaller assistants alongside platform-backed products, increasing pressure on incumbents to retain engagement.

Third-order effects

  • If rankings continue to vary by test and task, the market is likely to fragment around use cases and workflow fit rather than settle on a single universally best assistant.
  • Independent evaluations may become more influential as assistants compete to become an everyday work surface, though any one test remains a limited snapshot of fast-changing products.

The trend: Consumer AI competition is shifting from a single-model race toward recurring, task-specific comparisons among assistants embedded in different products and workflows.

Discussion

  • @ociubotaru Oleg Ciubotaru on threads
    Well this is quite an intro for a paper... Adding to my reading list.  “Even without any narrative or industry- specific information, the LLM outperforms financial analysts in its ability to predict earnings changes.  The LLM exhibits a relative advantage over human analysts in s…
  • @jaspar.bsky.social @jaspar.bsky.social on bluesky
    literally every time I've fact-checked perplexity it's gotten major details wrong [embedded post]
  • @khalifahmanaa Khalifa Manaa on x
    Comparison between Perplexity and the others is a bit funny, but overall my experience are aligned with this. Perplexity is great as an answering machine. For everything else chatGPT, Gemini & Claude. I don't know what Microsoft did with co-pilot but it's somehow bad with
  • @melindabchu1 Melinda B. Chu on x
    @AravSrinivas I like Perplexity a lot, but it shouldn't have been on this survey. It's not the same as the general purpose LLMs; has less overall capabilities. Perplexity should be compared it to just Google since you guys are just Search, which you talk and post about a lot.
  • @aravsrinivas Aravind Srinivas on x
    Perplexity has been ranked the number one AI chatbot in a survey run by Wall Street Journal, ahead of ChatGPT, Gemini. Microsoft Copilot is the least preferred. [image]
  • @karadapena Kara Dapena on x
    WSJ tested five top AI chatbots in real-world, everyday tasks. Did the famous ChatGPT come out on top? Or was it bested by Google, Microsoft or one of the AI startups? Have a look. https://www.wsj.com/... via @WSJ w/@dalvin_brown @JoannaStern @wjrothman @CloudberryRoque [image]
  • @firstadopter Tae Kim on x
    @AravSrinivas @perplexity_ai A scrappy startup already has a more effective product for daily life than the well-known 700+ person ChatGPT maker and tech giants. Tens of billions of funding aren't everything. Innovation and competition are grand. https://www.wsj.com/... [image]
  • @aravsrinivas Aravind Srinivas on x
    Perplexity has been ranked the number one AI chatbot in a survey run by Wall Street Journal, ahead of ChatGPT, Gemini. Microsoft Copilot is the least preferred. [image]
  • @denisyarats Denis Yarats on x
    now @WSJ runs LLM evals too 🤯 https://www.wsj.com/... [image]