HyperWrite CEO unveils Reflection 70B, based on Llama 3.1 70B Instruct and trained using reflection-tuning, and says it beats GPT-4o in all benchmarks tested
There's a new king in town: Matt Shumer, co-founder and CEO of AI writing startup HyperWrite, today unveiled Reflection 70B …
VentureBeat Carl Franzen
Context & Ripple Effects
HyperWrite positioned its model around a claimed benchmark advantage over GPT-4o and a reflection-tuning approach built on Llama 3.1 70B Instruct. That made the announcement a test not just of the model, but of whether its evaluation evidence could withstand outside review.
The claim quickly became contested in subsequent questions about Reflection’s reported performance, and a later postmortem attributed the gap to an initial benchmarking-code bug. The sequence makes verification—not the initial leaderboard position—the durable significance of the story.
First-order effects
- HyperWrite and Matt Shumer gain immediate visibility from a claim of across-the-board benchmark leadership, while putting the credibility of their evaluation process at the center of the release.
- The later inability of evaluators to reproduce some results means the stated comparison with GPT-4o cannot be treated as a settled performance result.
Second-order effects
- Prospective users and model evaluators have reason to seek reproducible test setups and independent runs before using the claim in product or procurement decisions.
- Competing model developers face stronger pressure to publish evaluation details when performance claims depend on custom methods or implementation choices.
Third-order effects
- If releases repeatedly pair bold benchmark claims with later corrections, reproducibility and transparent evaluation harnesses may become a more important differentiator than isolated leaderboard results.
- Reflection-tuning remains a technique claim rather than proof of a durable capability edge unless results can be independently reproduced across relevant tests.
The trend: This is one data point in the shift from benchmark-led model marketing toward independently verifiable evaluation as the basis for AI-model credibility.
Related: HyperWrite · Reflection 70B · Performance questions around Reflection · Reflection 70B benchmarking postmortem · Recursive self-improvement
Related Coverage
- New Open Source AI Model Can Check Itself and Avoid Hallucinations Inc.com · Kit Eaton
- Reflection 70B: A Ground Breaking Open-Source LLM, Trained with a New Technique called Reflection-Tuning that Teaches a LLM to Detect Mistakes in Its Reasoning and Correct Course MarkTechPost · Pragati Jhunjhunwala
- Startup aims to open source the world's most capable AI model The Decoder · Matthias Bastian
Discussion
-
@mattshumer_
Matt Shumer
on x
I'm excited to announce Reflection 70B, the world's top open-source model. Trained using Reflection-Tuning, a technique developed to enable LLMs to fix their own mistakes. 405B coming next week - we expect it to be the best model in the world. Built w/ @GlaiveAI. Read on ⬇️: [ima…
-
@mattshumer_
Matt Shumer
on x
IMPORTANT REFLECTION UPDATE: We have identified and fixed the issue on our Hugging Face repo. If you previously tried to download, run, or host Reflection Llama 70B, please try again now. The outputs should be far better. fp16 version coming in soon as well.
-
@mattshumer_
Matt Shumer
on x
Compute for Reflection 405B secured ✅ We're getting started training now Expect results very soon!
-
@arindam1408
Arindam Mitra
on x
Thread 1/n 1. We know that high quality data => powerful model 2. We used prompt engineering before to show case that we can significantly improve LLAMA2-7B to beat models that 10x of its size (Orca 2: https://arxiv.org/...) @mattshumer_
-
@dev_num0
@dev_num0
on x
i asked Reflection 70B : “give me some tests to see if you have this self-reflection mechanism in yourself or not” and this is the response: [image]
-
@tim_dettmers
Tim Dettmers
on x
Looks like we got project Strawberry/Orion/Q* a bit earlier than expected 😂 Actions speak louder than words. Who is gonna pay $2k/month now?
-
@everydayai_
@everydayai_
on x
Zero-shotting and outbenching: Claude 3.5 (5 shot) Gemini 1.5 (5 shot) and Llama 405 (5 shot) is absolute bananas Hoping there's a paper on this one, cuz we might have to cover it eventually on @EverydayAI_
-
@altryne
Alex Volkov
on x
This from @mattshumer_ and @csahil28 (@GlaiveAI ) is insane! A LLama 70B finetune that has reflection baked into it's weights, does CoT, Reflection and then spits out great answers, beats Sonnet on benchmarks!? This is the #breakingNews I wanted to share today on [image]
-
@msg
Michael S Galpert
on x
amazing that achieving these benchmarks is just using one extra call to reflect on the original answer provided
-
@degeneratoor
@degeneratoor
on x
Holy shit @mattshumer_ , what did you cook? [image]
-
@andthatto
@andthatto
on x
70B beating up sonnet on multiple benchmarks makes me think what openai can do. Gotta try it hands on asap
-
@gospaceport
@gospaceport
on x
Excited @ollama got Reflection model out + cudos @mattschumer_ for making a cool new fine-tune 👍 I prior tested stock Llama3.1-70b so reran prompts on quad 3090s. Did get some differing results. The <reflect> on mine seems broken 🤔but the 8q does seem improved. [video]
-
@shreyshahi
@shreyshahi
on x
This tune of Llama70B smokes 405B on all benchmarks!! A good preview of how synthetic data and more inference time techniques will take the current generation of LLMs so much further.
-
@ikristoph
@ikristoph
on x
I don't want to be negative by any means here about the work @mattshumer_ and his team have . These guys have done some fantastic work and I am super excited to try their models and you should be too! However the benchmark comparisons aren't really fair because the model is [imag…
-
@hyperbolic_labs
@hyperbolic_labs
on x
Reflection 70B at FP16: Now Available on Hyperbolic! ➡️ Chat with it and try our API for FREE: https://app.hyperbolic.xyz/... One trick: adding “Think very carefully” at the end of your message will make the model think more carefully and smarter! [video]
-
@shumochu
@shumochu
on x
The coolest model these days is launching on @huggingface and @hyperbolic_labs at the same time. And this is becoming a norm. I think you are not bullish enough for decentralized AI.
-
@prashant_1722
Prashant
on x
BREAKING NEWS 🔥 World's top open source LLM model is here. Reflection 70B fixes its own mistakes by using a technique caused Reflection tuning, mitigating hallucinations. It's the top LLM in (at least) MMLU, MATH, IFEval, GSM8K. Beats GPT-4o on every benchmark tested. It [image]
-
@yoheinakajima
Yohei
on x
honestly some of the technical details of AI models are beyond me. love the simplicity of this approach! prompt engineering getting baked into model training!
-
@codebreach
Madhav Bhagat
on x
And just like that- an open source 70B model that beats GPT4o level models The reflection tuning and prompting technique seems to have a big impact Especially now that models naturally do chain of thought
-
@p4lsec
@p4lsec
on x
Wow this is super impressive! The reflection tuning is going to be a core component of LLMs going forward.
-
@burkov
Andriy Burkov
on x
This guy just finetuned Llama 3.1 70B and it now beats the top LLMs on benchmarks. The demo page is overloaded right now, but I managed to do a couple of tests, it looks as strong as GPT-4o if not stronger.
-
@kimmonismus
@kimmonismus
on x
I can hardly believe what I'm reading here: an LLM that fixes its own bugs, corrects itself and beats all current models, including GPT-4o in all benchmarks? And the model is still OpenSource? “It's the top LLM in (at least) MMLU, MATH, IFEval, GSM8K. Beats GPT-4o on every [image…
-
@hu_yifei
Yifei Hu
on x
99.2% on gsm8k 😯
-
@deepwhitman
Bilal
on x
Great work! Open Source is our only hope. The labs will keep their fruits so we must cultivate our own garden.
-
@sergatx
Serg
on x
Tech News : A huge win for the open source world in AI. An open source model that matches up and even exceeds the top closed models from OpenAI and Anthropic, @mattshumer_ announces Reflection 70B.
-
@maximelabonne
Maxime Labonne
on x
This is super cool but I have a lot of questions. First, reflection = CoT on steroids. It means you can't compare these scores at all. Remember when people made fun of Gemini for providing CoT results for MMLU? This is a lot worse. Secondly, if you don't parse the output and
-
@gazorp5
@gazorp5
on x
Major grifter energy. - 99.2% performance on GSM8k even though GSM8k has more than >1% error rate. -Doesn't disclose that he's an investor in Glaive. - Gives misleading details about RLHF (the model was finetuned from llama3.1-70b-Instruct).
-
@random_walker
Arvind Narayanan
on x
I want to see how well these results translate from benchmarks to real world tasks, but if they hold up, it's an excellent example of how much low hanging fruit there is in AI development. The idea of doing reasoning using tokens hidden from the user is well known and has been
-
@humblyalex
Alex Gopoian
on x
Replicated @MattShumer_'s Reflection 70b in ChatGPT-4o with a single prompt that not only implemented its strategy, but had immediately used the strategy to refine itself as a strategy before I provided it the questions to tackle. I'll add other answers it gave below. [image]
-
@philip_kiely
Philip Kiely
on x
When a new model like Reflection 70B by @mattshumer_ drops, I want to get it up and running as soon as possible. Here's how I built and deployed a TensorRT-LLM engine for the model with basically zero coding work. Links in the replies if you want to do the same! [video]
-
@martinbowling
Martin Bowling
on x
Since @mattshumer_ dropped Reflection 70b and @hyperbolic_labs was nice enough to provide free inference for a week I decided to have @Replit agent build a chat bot for it. [video]
-
@tommyfalkowski
Tommy Falkowski
on x
Running Reflection Llama 3.1 70b 4bit by @mattshumer_ with @ollama on a 128GB MacBook Pro. What a wild experience! [video]
-
@gabrielmbmb_
Gabriel Martín Blázquez
on x
Yesterday Reflection 70B was released, a model fine-tuned using Reflection-Tuning that achieved impressive scores in several benchmarks such as MMLU. The dataset that was used for the fine-tuning wasn't released, but here is a distilabel recipe you can use to generate a similar […
-
@terryyuezhuo
Terry Yue Zhuo
on x
After verifying the required setup (with system prompt, no prefilling), I can safely say Reflection does not do well on BigCodeBench-Hard, at least. Complete: 20.3 (vs 28.4 from Llama3.1-70B) Instruct: 14.9 (vs 23.6 from Llama3.1-70B) The CoT/thinking/reflection process
-
@_philschmid
Philipp Schmid
on x
Mindblowing! 🤯 A 70B open @AIatMeta Llama 3 better than @AnthropicAI Claude 3.5 Sonnet and @OpenAI GPT-4o using Reflection-Tuning! In Reflection Tuning, the LLM is trained on synthetic, structured data to learn reasoning and self-correction. 👀 In the assistant response, the [imag…
-
@slow_developer
Haider
on x
Huge The first independent benchmark for Reflection 70B shows a 9-point improvement over the Llama 70B model, reaching 50%. So with that: 1. OpenAI 2. Google 3. Matt from the IT department 4. Meta 5. Anthropic [image]
-
@tedx_ai
Ted Werbel
on x
This tool for generating synthetic datasets was used to train the new Reflection 70B open source model which currently seems to outperform every other open source model by a significant margin! Definitely worth exploring... [video]
-
@elder_plinius
@elder_plinius
on x
that's odd... Reflection-70b claims to have been created by Anthropic, not Meta. “Upon careful consideration, I remain confident that Anthropic created me, not Meta.” I wonder if any Anthropic models were used to generate the synthetic data 🤔 [image]
-
@hu_yifei
Yifei Hu
on x
I quickly ran some vibe checks on Reflection-70B (by @mattshumer_ ). The model is served using the latest version of SGLang in BF16. I used the system prompt and sampling parameters recommended in the HF model card. Pros: - It's nice to see the “reasoning” process listed in a [im…
-
@matthewberman
@matthewberman
on x
What is Reflection 70b? This new model uses a novel technique to teach LLMs to SELF CORRECT, greatly reducing hallucinations. Here's everything you need to know: [video]
-
@matthewberman
@matthewberman
on x
Reflection 70b just dropped and is beating every other model, including GPT4o and Claude 3.5. How did this happen? Here's my conversation with Matt Shumer (@mattshumer_) and Sahil Chaudhary (@csahil28), the authors of Reflection 70b. [video]
-
@alphasignalai
Lior
on x
This is huge. A new technique called Reflection-Tuning allows open-source models (Llama 3.1 70B) to outperform Claude 3.5 and GPT-4o. This new technique trains the model on structured, synthetic data to detect reasoning errors and enable LLMs to fix their own mistakes. [image]
-
@mattshumer_
Matt Shumer
on x
Reflection 70B holds its own against even the top closed-source models (Claude 3.5 Sonnet, GPT-4o). It's the top LLM in (at least) MMLU, MATH, IFEval, GSM8K. Beats GPT-4o on every benchmark tested. It clobbers Llama 3.1 405B. It's not even close. [image]
-
@clementdelangue
Clem
on x
Reflection-Llama-3.1-70B by @mattshumer_ @csahil28 @GlaiveAI is #1 trending on HF. I've said it and will say it again: you don't need to be a big tech to fine-tune, optimize and run your own models for your specific constraints and you will benefit massively from it. [image]
-
@mattshumer_
Matt Shumer
on x
The technique that drives Reflection 70B is simple, but very powerful. Current LLMs have a tendency to hallucinate, and can't recognize when they do so. Reflection-Tuning enables LLMs to recognize their mistakes, and then correct them before committing to an answer. [image]
-
r/LocalLLaMA
r
on reddit
Tweet from Matt Shumer: “IMPORTANT REFLECTION UPDATE: We have identified and fixed the issue on our Hugging Face repo. …