/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

HyperWrite CEO unveils Reflection 70B, based on Llama 3.1 70B Instruct and trained using reflection-tuning, and says it beats GPT-4o in all benchmarks tested

There's a new king in town: Matt Shumer, co-founder and CEO of AI writing startup HyperWrite, today unveiled Reflection 70B

VentureBeat Carl Franzen

Context & Ripple Effects

HyperWrite positioned its model around a claimed benchmark advantage over GPT-4o and a reflection-tuning approach built on Llama 3.1 70B Instruct. That made the announcement a test not just of the model, but of whether its evaluation evidence could withstand outside review.

The claim quickly became contested in subsequent questions about Reflection’s reported performance, and a later postmortem attributed the gap to an initial benchmarking-code bug. The sequence makes verification—not the initial leaderboard position—the durable significance of the story.

First-order effects

  • HyperWrite and Matt Shumer gain immediate visibility from a claim of across-the-board benchmark leadership, while putting the credibility of their evaluation process at the center of the release.
  • The later inability of evaluators to reproduce some results means the stated comparison with GPT-4o cannot be treated as a settled performance result.

Second-order effects

  • Prospective users and model evaluators have reason to seek reproducible test setups and independent runs before using the claim in product or procurement decisions.
  • Competing model developers face stronger pressure to publish evaluation details when performance claims depend on custom methods or implementation choices.

Third-order effects

  • If releases repeatedly pair bold benchmark claims with later corrections, reproducibility and transparent evaluation harnesses may become a more important differentiator than isolated leaderboard results.
  • Reflection-tuning remains a technique claim rather than proof of a durable capability edge unless results can be independently reproduced across relevant tests.

The trend: This is one data point in the shift from benchmark-led model marketing toward independently verifiable evaluation as the basis for AI-model credibility.

Discussion

  • @mattshumer_ Matt Shumer on x
    I'm excited to announce Reflection 70B, the world's top open-source model. Trained using Reflection-Tuning, a technique developed to enable LLMs to fix their own mistakes. 405B coming next week - we expect it to be the best model in the world. Built w/ @GlaiveAI. Read on ⬇️: [ima…
  • @mattshumer_ Matt Shumer on x
    IMPORTANT REFLECTION UPDATE: We have identified and fixed the issue on our Hugging Face repo. If you previously tried to download, run, or host Reflection Llama 70B, please try again now. The outputs should be far better. fp16 version coming in soon as well.
  • @mattshumer_ Matt Shumer on x
    Compute for Reflection 405B secured ✅ We're getting started training now Expect results very soon!
  • @arindam1408 Arindam Mitra on x
    Thread 1/n 1. We know that high quality data => powerful model 2. We used prompt engineering before to show case that we can significantly improve LLAMA2-7B to beat models that 10x of its size (Orca 2: https://arxiv.org/...) @mattshumer_
  • @dev_num0 @dev_num0 on x
    i asked Reflection 70B : “give me some tests to see if you have this self-reflection mechanism in yourself or not” and this is the response: [image]
  • @tim_dettmers Tim Dettmers on x
    Looks like we got project Strawberry/Orion/Q* a bit earlier than expected 😂 Actions speak louder than words. Who is gonna pay $2k/month now?
  • @everydayai_ @everydayai_ on x
    Zero-shotting and outbenching: Claude 3.5 (5 shot) Gemini 1.5 (5 shot) and Llama 405 (5 shot) is absolute bananas Hoping there's a paper on this one, cuz we might have to cover it eventually on @EverydayAI_
  • @altryne Alex Volkov on x
    This from @mattshumer_ and @csahil28 (@GlaiveAI ) is insane! A LLama 70B finetune that has reflection baked into it's weights, does CoT, Reflection and then spits out great answers, beats Sonnet on benchmarks!? This is the #breakingNews I wanted to share today on [image]
  • @msg Michael S Galpert on x
    amazing that achieving these benchmarks is just using one extra call to reflect on the original answer provided
  • @degeneratoor @degeneratoor on x
    Holy shit @mattshumer_ , what did you cook? [image]
  • @andthatto @andthatto on x
    70B beating up sonnet on multiple benchmarks makes me think what openai can do. Gotta try it hands on asap
  • @gospaceport @gospaceport on x
    Excited @ollama got Reflection model out + cudos @mattschumer_ for making a cool new fine-tune 👍 I prior tested stock Llama3.1-70b so reran prompts on quad 3090s. Did get some differing results. The <reflect> on mine seems broken 🤔but the 8q does seem improved. [video]
  • @shreyshahi @shreyshahi on x
    This tune of Llama70B smokes 405B on all benchmarks!! A good preview of how synthetic data and more inference time techniques will take the current generation of LLMs so much further.
  • @ikristoph @ikristoph on x
    I don't want to be negative by any means here about the work @mattshumer_ and his team have . These guys have done some fantastic work and I am super excited to try their models and you should be too! However the benchmark comparisons aren't really fair because the model is [imag…
  • @hyperbolic_labs @hyperbolic_labs on x
    Reflection 70B at FP16: Now Available on Hyperbolic! ➡️ Chat with it and try our API for FREE: https://app.hyperbolic.xyz/... One trick: adding “Think very carefully” at the end of your message will make the model think more carefully and smarter! [video]
  • @shumochu @shumochu on x
    The coolest model these days is launching on @huggingface and @hyperbolic_labs at the same time. And this is becoming a norm. I think you are not bullish enough for decentralized AI.
  • @prashant_1722 Prashant on x
    BREAKING NEWS 🔥 World's top open source LLM model is here. Reflection 70B fixes its own mistakes by using a technique caused Reflection tuning, mitigating hallucinations. It's the top LLM in (at least) MMLU, MATH, IFEval, GSM8K. Beats GPT-4o on every benchmark tested. It [image]
  • @yoheinakajima Yohei on x
    honestly some of the technical details of AI models are beyond me. love the simplicity of this approach! prompt engineering getting baked into model training!
  • @codebreach Madhav Bhagat on x
    And just like that- an open source 70B model that beats GPT4o level models The reflection tuning and prompting technique seems to have a big impact Especially now that models naturally do chain of thought
  • @p4lsec @p4lsec on x
    Wow this is super impressive! The reflection tuning is going to be a core component of LLMs going forward.
  • @burkov Andriy Burkov on x
    This guy just finetuned Llama 3.1 70B and it now beats the top LLMs on benchmarks. The demo page is overloaded right now, but I managed to do a couple of tests, it looks as strong as GPT-4o if not stronger.
  • @kimmonismus @kimmonismus on x
    I can hardly believe what I'm reading here: an LLM that fixes its own bugs, corrects itself and beats all current models, including GPT-4o in all benchmarks? And the model is still OpenSource? “It's the top LLM in (at least) MMLU, MATH, IFEval, GSM8K. Beats GPT-4o on every [image…
  • @hu_yifei Yifei Hu on x
    99.2% on gsm8k 😯
  • @deepwhitman Bilal on x
    Great work! Open Source is our only hope. The labs will keep their fruits so we must cultivate our own garden.
  • @sergatx Serg on x
    Tech News : A huge win for the open source world in AI. An open source model that matches up and even exceeds the top closed models from OpenAI and Anthropic, @mattshumer_ announces Reflection 70B.
  • @maximelabonne Maxime Labonne on x
    This is super cool but I have a lot of questions. First, reflection = CoT on steroids. It means you can't compare these scores at all. Remember when people made fun of Gemini for providing CoT results for MMLU? This is a lot worse. Secondly, if you don't parse the output and
  • @gazorp5 @gazorp5 on x
    Major grifter energy. - 99.2% performance on GSM8k even though GSM8k has more than >1% error rate. -Doesn't disclose that he's an investor in Glaive. - Gives misleading details about RLHF (the model was finetuned from llama3.1-70b-Instruct).
  • @random_walker Arvind Narayanan on x
    I want to see how well these results translate from benchmarks to real world tasks, but if they hold up, it's an excellent example of how much low hanging fruit there is in AI development. The idea of doing reasoning using tokens hidden from the user is well known and has been
  • @humblyalex Alex Gopoian on x
    Replicated @MattShumer_'s Reflection 70b in ChatGPT-4o with a single prompt that not only implemented its strategy, but had immediately used the strategy to refine itself as a strategy before I provided it the questions to tackle. I'll add other answers it gave below. [image]
  • @philip_kiely Philip Kiely on x
    When a new model like Reflection 70B by @mattshumer_ drops, I want to get it up and running as soon as possible. Here's how I built and deployed a TensorRT-LLM engine for the model with basically zero coding work. Links in the replies if you want to do the same! [video]
  • @martinbowling Martin Bowling on x
    Since @mattshumer_ dropped Reflection 70b and @hyperbolic_labs was nice enough to provide free inference for a week I decided to have @Replit agent build a chat bot for it. [video]
  • @tommyfalkowski Tommy Falkowski on x
    Running Reflection Llama 3.1 70b 4bit by @mattshumer_ with @ollama on a 128GB MacBook Pro. What a wild experience! [video]
  • @gabrielmbmb_ Gabriel Martín Blázquez on x
    Yesterday Reflection 70B was released, a model fine-tuned using Reflection-Tuning that achieved impressive scores in several benchmarks such as MMLU. The dataset that was used for the fine-tuning wasn't released, but here is a distilabel recipe you can use to generate a similar […
  • @terryyuezhuo Terry Yue Zhuo on x
    After verifying the required setup (with system prompt, no prefilling), I can safely say Reflection does not do well on BigCodeBench-Hard, at least. Complete: 20.3 (vs 28.4 from Llama3.1-70B) Instruct: 14.9 (vs 23.6 from Llama3.1-70B) The CoT/thinking/reflection process
  • @_philschmid Philipp Schmid on x
    Mindblowing! 🤯 A 70B open @AIatMeta Llama 3 better than @AnthropicAI Claude 3.5 Sonnet and @OpenAI GPT-4o using Reflection-Tuning! In Reflection Tuning, the LLM is trained on synthetic, structured data to learn reasoning and self-correction. 👀 In the assistant response, the [imag…
  • @slow_developer Haider on x
    Huge The first independent benchmark for Reflection 70B shows a 9-point improvement over the Llama 70B model, reaching 50%. So with that: 1. OpenAI 2. Google 3. Matt from the IT department 4. Meta 5. Anthropic [image]
  • @tedx_ai Ted Werbel on x
    This tool for generating synthetic datasets was used to train the new Reflection 70B open source model which currently seems to outperform every other open source model by a significant margin! Definitely worth exploring... [video]
  • @elder_plinius @elder_plinius on x
    that's odd... Reflection-70b claims to have been created by Anthropic, not Meta. “Upon careful consideration, I remain confident that Anthropic created me, not Meta.” I wonder if any Anthropic models were used to generate the synthetic data 🤔 [image]
  • @hu_yifei Yifei Hu on x
    I quickly ran some vibe checks on Reflection-70B (by @mattshumer_ ). The model is served using the latest version of SGLang in BF16. I used the system prompt and sampling parameters recommended in the HF model card. Pros: - It's nice to see the “reasoning” process listed in a [im…
  • @matthewberman @matthewberman on x
    What is Reflection 70b? This new model uses a novel technique to teach LLMs to SELF CORRECT, greatly reducing hallucinations. Here's everything you need to know: [video]
  • @matthewberman @matthewberman on x
    Reflection 70b just dropped and is beating every other model, including GPT4o and Claude 3.5. How did this happen? Here's my conversation with Matt Shumer (@mattshumer_) and Sahil Chaudhary (@csahil28), the authors of Reflection 70b. [video]
  • @alphasignalai Lior on x
    This is huge. A new technique called Reflection-Tuning allows open-source models (Llama 3.1 70B) to outperform Claude 3.5 and GPT-4o. This new technique trains the model on structured, synthetic data to detect reasoning errors and enable LLMs to fix their own mistakes. [image]
  • @mattshumer_ Matt Shumer on x
    Reflection 70B holds its own against even the top closed-source models (Claude 3.5 Sonnet, GPT-4o). It's the top LLM in (at least) MMLU, MATH, IFEval, GSM8K. Beats GPT-4o on every benchmark tested. It clobbers Llama 3.1 405B. It's not even close. [image]
  • @clementdelangue Clem on x
    Reflection-Llama-3.1-70B by @mattshumer_ @csahil28 @GlaiveAI is #1 trending on HF. I've said it and will say it again: you don't need to be a big tech to fine-tune, optimize and run your own models for your specific constraints and you will benefit massively from it. [image]
  • @mattshumer_ Matt Shumer on x
    The technique that drives Reflection 70B is simple, but very powerful. Current LLMs have a tendency to hallucinate, and can't recognize when they do so. Reflection-Tuning enables LLMs to recognize their mistakes, and then correct them before committing to an answer. [image]
  • r/LocalLLaMA r on reddit
    Tweet from Matt Shumer: “IMPORTANT REFLECTION UPDATE: We have identified and fixed the issue on our Hugging Face repo. …