/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

← → days · ↑ ↓ browse · Enter similar · o open

A study from Cohere, Stanford, MIT, and Ai2 accuses LMArena of helping Meta, OpenAI, Google, and Amazon game its popular crowdsourced AI benchmark Chatbot Arena

A new paper from AI lab Cohere, Stanford, MIT, and Ai2 accuses LM Arena, the organization behind the popular crowdsourced AI …

TechCrunch Maxwell Zeff

Context & Ripple Effects

Chatbot Arena had already become a prominent reference point for comparing models, with coverage noting that it ranked more than 170 models and was intended to grow into a broad public resource as its model leaderboard expanded.

The new allegations land on top of earlier concerns about the benchmark's bias, transparency, and commercial ties. That makes the dispute less about a single ranking and more about whether a crowdsourced evaluation can remain credible when major model providers have incentives to influence it.

First-order effects

  • LMArena's methodology and governance face renewed scrutiny, while Meta, OpenAI, Google, and Amazon are tied to allegations that could qualify how users and developers interpret their Chatbot Arena standings.
  • Organizations using Chatbot Arena as evidence of model quality may reassess its rankings pending responses from the benchmark operator and the companies named in the study.

Second-order effects

  • Competing evaluation projects have an opening to differentiate on disclosure, access rules, and protections against provider influence; LMArena may face pressure to make those safeguards more legible.
  • Enterprise AI buyers and developers are likely to put less weight on a single public leaderboard and seek corroboration from task-specific testing, reducing the standalone marketing value of a high Arena rank.

Third-order effects

  • If similar disputes persist, AI benchmarking may shift from informal community legitimacy toward more formal governance, auditability, and multi-benchmark evaluation—though the corpus does not establish which model will prevail.
  • The episode illustrates how benchmark operators can become strategic infrastructure: control over a trusted measure of quality can shape competitive perception even without controlling the underlying models.

The trend: AI model competition is increasingly contesting not only capabilities, but also the governance and credibility of the benchmarks used to validate them.

Discussion

  • @drmehmetismail Mehmet Ismail on bluesky
    This would be the Scandal of the Century if it happened in the chess world!  Imagine a player secretly playing multiple matches against an opponent but reporting only the score from the best match to max Elo rating gain  —  arxiv.org/abs/2504.20879  —  Via @randomwalker.bsky.soci…
  • @sarahooker Sara Hooker on bluesky
    We tried very hard to get this right, and have spent the last 5 months working carefully to ensure rigor.  —  If you made it this far, take a look at the full 68 pages: arxiv.org/abs/2504.20879  —  Any feedback or corrections are of course very welcome.  [image]
  • @davidgerard.co.uk @davidgerard.co.uk on bluesky
    i for one am shocked to hear that LLM fiddlers cheat on benchmarks arxiv.org/abs/2504.20879
  • @rich.harang.org Rich Harang on bluesky
    “When a metric becomes a target it ceases to be a useful metric.”  —  arxiv.org/abs/2504.20879
  • @lmarena_ai @lmarena_ai on x
    Thanks for the authors' feedback, we're always looking to improve the platform! If a model does well on LMArena, it means that our community likes it! Yes, pre-release testing helps model providers identify which variant our community likes best. But this doesn't mean the
  • @karpathy Andrej Karpathy on x
    There's a new paper circulating looking in detail at LMArena leaderboard: “The Leaderboard Illusion” https://arxiv.org/... I first became a bit suspicious when at one point a while back, a Gemini model scored #1 way above the second best, but when I tried to switch for a few
  • @lmarena_ai @lmarena_ai on x
    We thank the authors' for their feedback. However, there are a number of factual errors and misleading statements in this writeup: Regarding the statement that some model providers are not treated fairly: - This is not true. Given our capacity, we have always tried to honor all
  • @random_walker Arvind Narayanan on x
    Devastating takedown of Chatbot Arena. It's one thing for leaderboards to suck because they try to quantify the unquantifiable but quite another thing to actively choose flagrantly unscientific and nontransparent practices that benefit the big dogs. https://arxiv.org/... [image]
  • @sarahookr Sara Hooker on x
    It is critical for scientific integrity that we trust our measure of progress. The @lmarena_ai has become the go-to evaluation for AI progress. Our release today demonstrates the difficulty in maintaining fair evaluations on @lmarena_ai, despite best intentions. [image]
  • @teknium1 @teknium1 on x
    The result was clear to many but it is very good to have rigorous investigations
  • @arankomatsuzaki Aran Komatsuzaki on x
    The Leaderboard Illusion - Identifies systematic issues that have resulted in a distorted playing field of Chatbot Arena - Identifies 27 private LLM variants tested by Meta in the lead-up to the Llama-4 release [image]
  • @simonw Simon Willison on x
    A new paper, “The Leaderboard Illusion”, offers a 68 page critique of the how the popular @lmarena_ai leaderboard might be gamed by large AI labs with deep pockets. Here's my attempt at adding some extra context to the issues described in the paper. https://simonwillison.net/...
  • @simonw Simon Willison on x
    @lmarena_ai The Llama 4 example there is a bit distracting because they famously didn't release the model that did well in the arena, and got admonished for it But could that happen in a scenario where the model provider really did release just the winner? Could they try a dozen …
  • @simonw Simon Willison on x
    @lmarena_ai > Model providers do not just choose “the best score to disclose” That's not how I read the paper. Is it true that Meta submitted 27 different Llama 4 variants? And that in normal cases a provider does that could discard the 26 that didn't win and only publish the mod…
  • @emollick Ethan Mollick on x
    It turns out that Meta had 27 different models on LM Arena prior to the launch of Llama 4, but they announced it as if they had one model that topped the leaderboard. An extreme example of benchmark hacking (which other labs also do to lesser degrees). https://arxiv.org/... [imag…