HyperWrite's 70B parameter AI model, Reflection, has its performance questioned, after CEO Matt Shumer said something about its upload to Hugging Face was off
something got fucked up during the upload process. Will fix today. Forums: r/LocalLLaMA : Smh: Reflection was too good to be true - reference article
VentureBeatCarl Franzen
Context & Ripple Effects
Reflection was introduced days earlier as a Llama 3.1 70B Instruct-based model whose creator said it outperformed GPT-4o across tested benchmarks, making a working public release central to assessing those claims. The earlier launch claims therefore raised the stakes of a flawed Hugging Face upload.
This episode is an early point in a credibility dispute rather than a standalone distribution glitch: subsequent coverage said a benchmarking-code bug was identified in a postmortem, while Shumer later acknowledged he had “got ahead” of himself.
First-order effects
HyperWrite’s claimed results become difficult for users and evaluators to validate until the hosted model artifact is corrected, putting immediate pressure on the company’s account of its performance.
The upload problem gives community skepticism a concrete technical basis, shifting attention from the announced benchmark lead to whether the released model matches the one evaluated.
Second-order effects
Independent testers and prospective users are more likely to rely on reproducible runs rather than headline benchmark assertions, slowing adoption of Reflection while the discrepancy is resolved.
Competing open-model releases gain an advantage when their weights, evaluation code, and reported results can be checked consistently on the same distribution channel.
Third-order effects
If similar episodes recur, public model launches will increasingly be judged as auditable release packages—weights, prompts, code, and benchmarks—rather than as performance announcements alone.
The episode points toward a higher verification burden for claims of frontier-level open-model performance, though one failed upload by itself does not establish an industry-wide standard.
The trend: Open-model competition is making reproducible distribution and transparent evaluation as important to credibility as benchmark scores.
A story about fraud in the AI research community: On September 5th, Matt Shumer, CEO of OthersideAI, announces to the world that they've made a breakthrough, allowing them to train a mid-size model to top-tier levels of performance. This is huge. If it's real. It isn't. [image]
Matt starts making claims that there's something wrong with the API. There's something wrong with the upload. For *some* reason there's some glitch that's just about to be fixed. [image]
tl;dr Matt Shumer is a liar and a fraud. Presumably he'll eventually throw some poor sap engineer under the bus and pretend he was lied to. Grifters shit in the communal pool, sucking capital, attention, and other resources away from people who could actually make use of them. [i…
The whole Reflection-70B debacle points the the desperate need for a better AI evaluation ecosystem. It needs to be extremely easy to adjudicate: (1) is the model overfit to benchmarks (2) is the model truly unique (i.e. not a wrapper or thin fine-tune)
They get massive news coverage and are the talk of the town, so to speak. *If* this were real, it would represent a substantial advance in tuning LLMs at the *abstract* level, and could perhaps even lead to whole new directions of R&D. But soon, cracks appear in the story. [image…
On September 7th, the first independent attempts to replicate their claimed results fail. Miserably, actually. The performance is awful. Further, it is discovered that Matt isn't being truthful about what the released model actually is based on under the hood. [image]
But the thing about a private API is it's not really clear what it's calling on the backend. They could be calling a more powerful proprietary model under the hood. We should test and see. Trust, but verify. And it turns out that Matt is a liar. [image]
We've figured out the issue. The reflection weights on Hugging Face are actually a mix of a few different models — something got fucked up during the upload process. Will fix today.