An analysis of ChatGPT, Claude, Copilot, Gemini, and Perplexity's responses to some real-life questions and everyday tasks: Perplexity ranked first overall
We tested OpenAI's ChatGPT against Microsoft's Copilot and Google's Gemini, along with Perplexity and Anthropic's Claude. Here's how they ranked.
Wall Street Journal
Context & Ripple Effects
This comparison extends an earlier benchmark in which ChatGPT led Gemini-powered Bard across several task types, while the margin had already narrowed. The result makes everyday-task performance—not just model capability claims—a visible point of differentiation among consumer assistants.
Perplexity gains an independent, task-based validation against ChatGPT, Copilot, Gemini and Claude, potentially strengthening its positioning with users evaluating assistants on practical usefulness.
The other providers face a clear comparative reference point in a category where perceived answer quality and task completion are central to product choice.
Second-order effects
Assistant vendors have added incentive to optimize for repeatable real-world workflows and to communicate performance in user-facing terms, rather than relying solely on broad model announcements.
For users and enterprise buyers, side-by-side testing lowers the cost of considering specialist or smaller assistants alongside platform-backed products, increasing pressure on incumbents to retain engagement.
Third-order effects
If rankings continue to vary by test and task, the market is likely to fragment around use cases and workflow fit rather than settle on a single universally best assistant.
Independent evaluations may become more influential as assistants compete to become an everyday work surface, though any one test remains a limited snapshot of fast-changing products.
The trend: Consumer AI competition is shifting from a single-model race toward recurring, task-specific comparisons among assistants embedded in different products and workflows.
Well this is quite an intro for a paper... Adding to my reading list. “Even without any narrative or industry- specific information, the LLM outperforms financial analysts in its ability to predict earnings changes. The LLM exhibits a relative advantage over human analysts in s…
Comparison between Perplexity and the others is a bit funny, but overall my experience are aligned with this. Perplexity is great as an answering machine. For everything else chatGPT, Gemini & Claude. I don't know what Microsoft did with co-pilot but it's somehow bad with
@AravSrinivas I like Perplexity a lot, but it shouldn't have been on this survey. It's not the same as the general purpose LLMs; has less overall capabilities. Perplexity should be compared it to just Google since you guys are just Search, which you talk and post about a lot.
Perplexity has been ranked the number one AI chatbot in a survey run by Wall Street Journal, ahead of ChatGPT, Gemini. Microsoft Copilot is the least preferred. [image]
WSJ tested five top AI chatbots in real-world, everyday tasks. Did the famous ChatGPT come out on top? Or was it bested by Google, Microsoft or one of the AI startups? Have a look. https://www.wsj.com/... via @WSJ w/@dalvin_brown @JoannaStern @wjrothman @CloudberryRoque [image]
@AravSrinivas @perplexity_ai A scrappy startup already has a more effective product for daily life than the well-known 700+ person ChatGPT maker and tech giants. Tens of billions of funding aren't everything. Innovation and competition are grand. https://www.wsj.com/... [image]
Perplexity has been ranked the number one AI chatbot in a survey run by Wall Street Journal, ahead of ChatGPT, Gemini. Microsoft Copilot is the least preferred. [image]