A study from Cohere, Stanford, MIT, and Ai2 accuses LMArena of helping Meta, OpenAI, Google, and Amazon game its popular crowdsourced AI benchmark Chatbot Arena
A new paper from AI lab Cohere, Stanford, MIT, and Ai2 accuses LM Arena, the organization behind the popular crowdsourced AI …
TechCrunchMaxwell Zeff
Context & Ripple Effects
Chatbot Arena had already become a prominent reference point for comparing models, with coverage noting that it ranked more than 170 models and was intended to grow into a broad public resource as its model leaderboard expanded.
LMArena's methodology and governance face renewed scrutiny, while Meta, OpenAI, Google, and Amazon are tied to allegations that could qualify how users and developers interpret their Chatbot Arena standings.
Organizations using Chatbot Arena as evidence of model quality may reassess its rankings pending responses from the benchmark operator and the companies named in the study.
Second-order effects
Competing evaluation projects have an opening to differentiate on disclosure, access rules, and protections against provider influence; LMArena may face pressure to make those safeguards more legible.
Enterprise AI buyers and developers are likely to put less weight on a single public leaderboard and seek corroboration from task-specific testing, reducing the standalone marketing value of a high Arena rank.
Third-order effects
If similar disputes persist, AI benchmarking may shift from informal community legitimacy toward more formal governance, auditability, and multi-benchmark evaluation—though the corpus does not establish which model will prevail.
The episode illustrates how benchmark operators can become strategic infrastructure: control over a trusted measure of quality can shape competitive perception even without controlling the underlying models.
The trend: AI model competition is increasingly contesting not only capabilities, but also the governance and credibility of the benchmarks used to validate them.
This would be the Scandal of the Century if it happened in the chess world! Imagine a player secretly playing multiple matches against an opponent but reporting only the score from the best match to max Elo rating gain — arxiv.org/abs/2504.20879 — Via @randomwalker.bsky.soci…
We tried very hard to get this right, and have spent the last 5 months working carefully to ensure rigor. — If you made it this far, take a look at the full 68 pages: arxiv.org/abs/2504.20879 — Any feedback or corrections are of course very welcome. [image]
Thanks for the authors' feedback, we're always looking to improve the platform! If a model does well on LMArena, it means that our community likes it! Yes, pre-release testing helps model providers identify which variant our community likes best. But this doesn't mean the
There's a new paper circulating looking in detail at LMArena leaderboard: “The Leaderboard Illusion” https://arxiv.org/... I first became a bit suspicious when at one point a while back, a Gemini model scored #1 way above the second best, but when I tried to switch for a few
We thank the authors' for their feedback. However, there are a number of factual errors and misleading statements in this writeup: Regarding the statement that some model providers are not treated fairly: - This is not true. Given our capacity, we have always tried to honor all
Devastating takedown of Chatbot Arena. It's one thing for leaderboards to suck because they try to quantify the unquantifiable but quite another thing to actively choose flagrantly unscientific and nontransparent practices that benefit the big dogs. https://arxiv.org/... [image]
It is critical for scientific integrity that we trust our measure of progress. The @lmarena_ai has become the go-to evaluation for AI progress. Our release today demonstrates the difficulty in maintaining fair evaluations on @lmarena_ai, despite best intentions. [image]
The Leaderboard Illusion - Identifies systematic issues that have resulted in a distorted playing field of Chatbot Arena - Identifies 27 private LLM variants tested by Meta in the lead-up to the Llama-4 release [image]
A new paper, “The Leaderboard Illusion”, offers a 68 page critique of the how the popular @lmarena_ai leaderboard might be gamed by large AI labs with deep pockets. Here's my attempt at adding some extra context to the issues described in the paper. https://simonwillison.net/...
@lmarena_ai The Llama 4 example there is a bit distracting because they famously didn't release the model that did well in the arena, and got admonished for it But could that happen in a scenario where the model provider really did release just the winner? Could they try a dozen …
@lmarena_ai > Model providers do not just choose “the best score to disclose” That's not how I read the paper. Is it true that Meta submitted 27 different Llama 4 variants? And that in normal cases a provider does that could discard the 26 that didn't win and only publish the mod…
It turns out that Meta had 27 different models on LM Arena prior to the launch of Llama 4, but they announced it as if they had one model that topped the leaderboard. An extreme example of benchmark hacking (which other labs also do to lesser degrees). https://arxiv.org/... [imag…