Meta VP of Generative AI Ahmad Al-Dahle denies a rumor that the company trained Llama 4 Maverick and Scout on test sets, saying that Meta “would never do that”
but the EU doesn't get everything Pascale Davies / Euronews : From a political shift to a more powerful AI: Everything to know about Meta's Llama 4 models Jay Bonggolto / Android Central : Meta is coming for Google and OpenAI with its fresh Llama 4 models Ben Wodecki / Capacity Media : Meet Meta's Llama 4: Bigger brains, sharper vision, more modalities Katelyn Chedraoui / CNET : Meta Dropped Llama 4: What to Know About the Two New AI Models Tony Peng on Substack : A Chinese whistleblower from Meta's AI team on Llama 4: After repeated training … Matthias Bastian / The Decoder : Meta's Llama 4 models show promise on standard tests, but struggle with long-context tasks Nathan Lambert / Interconnects : Llama 4: Did Meta just push the panic button? Bluesky: Corey Quinn / @quinnypig.com : Yes, when I think “paragon of business ethics,” I think of Facebook. [embedded post] X: Ahmad Al-Dahle / @ahmad_al_dahle : ...That said, we're also hearing some reports of mixed quality across different services. Since we dropped the models as soon as they were ready, we expect it'll take several days for all the public implementations to get dialed in. We'll keep working through our bug fixes and onboarding partners. We've also heard claims that we trained on test sets — that's simply not true and we would never do that. @thexeophon : Llama 4 on LMsys is a totally different style than Llama 4 elsewhere, even if you use the recommended system prompt. Tried various prompts myself META did not do a specific deployment / system prompt just for LMsys, did they? 👀 Paul Gauthier / @paulgauthier : Llama 4 Maverick scored 16% on the aider polyglot coding benchmark. https://aider.chat/... [image] @artificialanlys : Llama 4 independent evals: Maverick (402B total, 17B active) beats Claude 3.7 Sonnet, trails DeepSeek V3 but more efficient; Scout (109B total, 17B active) in-line with GPT-4o mini, ahead of Mistral Small 3.1 We have independently benchmarked Scout and Maverick as scoring 36 and 49 in Artificial Analysis Intelligence Index respectively... Markus Zimmermann / @zimmskal : Preliminary results for Meta's Llama v4 for DevQualityEval v1.0 DO NOT LOOK GOOD 😱😿 It seems that both models TANKED in Java, which is a big part of the eval. Good in Go and Ruby but not TOP10 good. Meta: Llama v4 Scout 109B - 🏁 Overall score 62.53% mid-range - 🐕🦺 With [image] Chase Brower / @chasebrowe32432 : This is the first time for any major LLM that I'm genuinely thinking they just straight up trained on the benchmark answers for the mainline benchmarks Llama 4 is failing spectacularly on like every 3rd party bench i've seen Susan Zhang / @suchenzang : 4D chess move 🧐: use llama4 experimental to hack lmsys, expose the slop preference, and finally discredit the entire ranking system (lather rinse repeat above for academic benchmark maxing too) Yuchen Jin / @yuchenj_uw : If Meta actually did this for Llama 4 training to maximize benchmark scores, it's fucked. [image] Andriy Burkov / @burkov : “We've also heard claims that we trained on test sets — that's simply not true, and we would never do that.” No one said you trained on the test set. What they said is that you seem to have finetuned to benchmarks. It's especially obvious when you look at the Elo rating and then test the model right there and see how GPT-3.5-ish the output is... Min Choi / @minchoi : Yikes. Llama 4 benchmarks looked insane but something feels off. Reddit leaks claim Meta cooked it. Here's what people are saying: [image] Kache / @yacinemtb : BREAKING: pseudonymous internet user believes an anonymous internet user. I mean, who would go on the internet and tell lies? @kimmonismus : Doesnt look all to good for Llama 4. [image] Susan Zhang / @suchenzang : > Company leadership suggested blending test sets from various benchmarks during the post-training process If this is actually true for Llama-4, I hope they remember to cite previous work from FAIR (Llama-1 and https://arxiv.org/...) for this unique approach! 🙏 [image] @abcampbell : tech bros are learning something finance bros have known for millennium quants are going to train on the test they can't help it obvious to anyone actually using these models @kimmonismus : If it is true that Llama, i.e. Meta, cheated in the benchmarks, it would be an unprecedented image loss. Currently, it seems that after the first testings, the mood is rather mediocre anyway. [image] Yifei Hu / @hu_yifei : @Yuchenj_UW Another person said: “I participated in the data mix part for sft and rl. I am not aware of such cases.” 让子弹飞一会儿 [image] @vibagor44145276 : The linked post is not true. There are indeed issues with Llama 4, from both the partner side (inference partners barely had time to prep. We sent out a few transformers wheels/vllm wheels mere days before release) and the model side. But there was no such training on test set. Nathan Lambert / @natolambert : Seems like Llama 4's reputation is maybe irreparably tarnished by having a separate unreleased model that was overfit to LMArena. Actual model is good, but shows again how crucial messaging and details are. Josh Clemm / @joshclemm : Trying to parse Meta's Llama 4 release this weekend? I felt this was a great writeup. In short: Meta's Llama 4 release was very different than past releases, with some odd timing and a different strategy than before. - They added three Mixture-of-Experts models: Scout Nathan Lambert / @natolambert : Llama 4 was a messy release: unreleased finetunes boosting scores, rumors of training on test, released on a weekend, etc As (open) models are commoditized / competition grows, what is the role of Meta's Llama efforts in the future? Should they continue? https://www.interconnects.ai/ ... Chris Paxton / @chris_j_paxton : All this discourse is making me want a survey on who the hell lmsys raters actually are. I thought the lmsys version was unbearable and almost never rated it positively... Lu Liu / @eliza_luth : I have been only trusting the public's choices, my own tests and the judgement of trusted researchers Nathan Lambert / @natolambert : Okay Llama 4 is def a littled cooked lol, what is this yap city [image] Wenting Zhao / @wzhao_nlp : Time to revisit our paper: Open community-driven evaluation platforms could be corrupted from a few sources of bad annotations, making their results not as trustworthy as we'd like. https://arxiv.org/... [image] Susan Zhang / @suchenzang : how did this llama4 score so high on lmsys?? i'm still buckling up to understand qkv through family reunions and weighted values for loving cats... [image] @techdevnotes : for some reason, the Llama 4 model in Arena uses a lot more Emojis on together . ai, it seems better: [image] Ethan Mollick / @emollick : Hopefully the Llama 4 models improve rapidly, as they did in the Llama 3 generation. The initial launch got pretty mixed feedback (including from me) but a good open weights model from Meta would be very useful for many people. Yann LeCun / @ylecun : Some carifications about Llama-4. Elvis / @omarsar0 : Thanks for clarifying this. Maybe some official docs/guide (prompting/usage tips, recommended settings, error expectations, areas/use cases to apply and how to apply, etc) would be helpful here. I am aware of model cards, prompting guides but I think a lot of folks are running into issues with output quality. Nathan Lambert / @natolambert : > be me > be zuck > need llama 4 to land > send a model/prompt to LMSYS to get a top1 score, cringe be damned > release a different model as “open source” > think people won't find out even with weights Forums: Hacker News : Serious issues in Llama 4 training. I Have Submitted My Resignation to GenAI" r/LocalLLaMA : “Serious issues in Llama 4 training. I Have Submitted My Resignation to GenAI” r/LocalLLaMA : “...we're also hearing some reports of mixed quality across different services. Since we dropped the models as soon as they were ready … r/LocalLLaMA : Meta Leaker refutes the training on test set claim
Context & Ripple Effects
Meta has built Llama through successive releases, from Llama 3’s scaled model lineup to a lower-cost Llama 3.3 offering that it said approached the performance of its largest model. Llama 4’s Maverick and Scout are the latest step in that release cadence.
The denial arrives just after Zuckerberg previewed a forthcoming Llama 4 reasoning model. That makes the integrity of published model comparisons consequential for Meta’s ability to sustain confidence in the broader Llama family.
First-order effects
- Meta must defend the credibility of Maverick and Scout’s benchmark results, while Al-Dahle’s categorical denial puts the company’s public position on record.
- Developers and evaluators comparing Llama 4 with alternatives have a reason to scrutinize whether reported test performance is independently reproducible.
Second-order effects
- Benchmark providers and third-party evaluators may face pressure to place more weight on contamination checks and independent testing when assessing newly released models.
- Competitors can use uncertainty around evaluation provenance to differentiate their own releases, shifting attention from model claims toward the methods behind them.
Third-order effects
- If such allegations recur across frontier releases, benchmark scores may become less decisive as a standalone signal; independent evaluation and disclosed testing practices would carry more of the market’s trust burden.
- The episode points to strategic legitimacy becoming part of model competition: labs will be judged not only on capability claims but also on whether their evaluation process is credible.
The trend: Frontier-model competition is increasingly a contest over verifiable evaluation credibility as well as raw reported capability.