/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

← → days · ↑ ↓ browse · Enter similar · o open

DeepSeek says its new DeepSeekMath-V2 model got gold-medal level status on the International Mathematical Olympiad 2025 and Chinese Mathematical Olympiad 2024

Chinese startup Deepseek reports its new DeepseekMath-V2 model has reached gold medal status at the Math Olympiad …

The Decoder Matthias Bastian

Context & Ripple Effects

DeepSeek has been building a math-and-reasoning line: its R1 update was presented as improving mathematics, programming and logic, while Prover-V2 was released as a math-focused model. This result extends that arc from general reasoning claims to olympiad-style evaluation.

The comparison set is tightening. DeepMind had already reported that AlphaGeometry2 exceeded average gold-medalist performance on historical geometry problems, making rigorous mathematical problem solving a visible frontier-model battleground.

First-order effects

  • DeepSeek gains a new, high-prestige performance claim for DeepSeekMath-V2, potentially strengthening its position with users evaluating models for difficult formal reasoning tasks.
  • The reported result raises the bar for how DeepSeek's math models will be scrutinized: benchmark methodology, problem coverage and reproducibility become central to the claim's value.

Second-order effects

  • Rival model developers face greater pressure to publish comparable olympiad, geometry and proof-oriented results rather than rely solely on broad reasoning benchmarks.
  • Teams building coding, scientific and theorem-proving workflows may give more weight to specialist math-model evaluations, provided the reported performance holds up beyond the stated contests.

Third-order effects

  • If repeated across independently assessable evaluations, olympiad-style performance could become a more important differentiator within the broader competition to build reliable reasoning systems.
  • The pattern points toward a split between broad general-purpose models and increasingly capable specialist reasoning models, though contest results alone do not establish real-world reliability.

The trend: AI labs are using demanding mathematical evaluations to demonstrate progress in reasoning, with specialist models becoming a key competitive layer alongside general-purpose systems.

Discussion

  • @clementdelangue Clem on x
    As far as I know, there isn't any chatbot or API that gives you access to an IMO 2025 gold-medalist model.  Not only does this change today, but you get to download the weights with the Apache 2.0 open-source release of @deepseek_ai Math-V2 on @huggingface!
  • @chombabupe @chombabupe on x
    My question is if the generator doesn't get better than the verifier. Isn't the generator then just a distilled version of the verifier and the self improvement is just an illusion? Why not train the proof generator model directly on what the verifier was trained on?
  • @reach_vb @reach_vb on x
    The whale is BACK!!! 👀👀👀 https://huggingface.co/...
  • @scaling01 @scaling01 on x
    DeepSeek is back with DeepSeek-Math V2 a mathematical reasoning model based on DeepSeek-V3.2-Exp-Base it outperforms the Gemini DeepThink model that got IMO gold on ProofBench Basic https://github.com/... [image]
  • @deedydas Deedy on x
    🚨China's DeepSeek just dropped the only open-source model good enough at math to win IMO Gold, and a must-read report!  The key idea draws from things Karpathy and others have spoken about: move beyond “final answer RL” into a generator-verifier-meta-verifier loop in pure languag…
  • @blancheminerva @blancheminerva on x
    DeepSeek is more aligned to my values than any US frontier lab. I wish EleutherAI had the billions it took to do this work, but since we have less than 1/1000th of that I take a lot of joy in seeing someone else do what needs to be done.
  • @mervenoyann Merve on x
    DeepSeek released DeepSeekMathv2 (based on DeepSeek-V3.2-Exp-Base) outperforming Gemini DeepThink on IMO ProofBench and CNML from paper: > they train an LLM-based verifier for reward function > they train this model using verifier, and ask it to resolve issues on its own > [image…
  • @novasarc01 Λux on x
    deepseekmath-v2 basically admits final-answer training is useless for real math so they build a prover-verifier loop that forces the model to generate proofs and call out its own mistakes. the verifier is trained with expert scores + meta-verification so it doesn't hallucinate [i…
  • @bindureddy Bindu Reddy on x
    Deepseek math is an IMO gold medal winning model built on top of their V3.2 base DeepSeek appears to be testing novel scalable post-training methods and will drop an amazing reasoning model Open source is going to close the gap to top SOTA models quickly
  • @presidentlin @presidentlin on x
    Deepseek is such a cool lab. Every time they see the open source community is struggling with something, they give their candidate answer and go back to the bottom of the oceans. Reasoning models almost got locked up in private labs, OAI, and Sam are soooo good at deflecting,
  • @ecommurz Elon Murz on x
    Interesting stuff. Multiple proof attempts runs in parallel, then a pack of verifiers check, fix n improve each other iteratively til nothing breaks. Its a self checking loop on steroids, built to handle natural language theorem proofs.. Open sourced too! [image]
  • @ai_for_success AshutoshShrivastava on x
    Chinese DeepSeek open sourced its IMO gold level model DeepSeek Math V2 on Thanksgiving TLDR - DeepSeek Math V2 is now fully open sourced - It focuses on self verifiable mathematical reasoning - Verifier checks proofs step by step and the generator fixes its own mistakes - [image…
  • @jenzhuscott Jen Zhu on x
    It just dawned on me why Wenfeng named his lab ‘DeepSeek’ w 🐋 as symbol: It rarely surfaces, but each time it does, it makes a huge splash. Then, it just goes deep into the ocean again & seek deeply. It carries heavy weight w quiet persistence. It's intelligent. It's rare.
  • @yuchenj_uw Yuchen Jin on x
    Twitter is such a bubble. Anons can casually declare a frontier lab “dead” and get 5M views. DeepSeek just released the first open-weight IMO 2025 gold-medalist model and is cooking V4. A year ago Google was “dead.” 1 month ago Anthropic. Now Nvidia, thanks to TPUs. [image]
  • @zhs05232838 Zhihong Shao on x
    We just shared some thoughts and results on self-verifiable mathematical reasoning. The released model, DeepSeekMath-V2, is strong on IMO-ProofBench and competitions like IMO 2025 (5/6 problems) and Putnam 2024 (a near-perfect score of 118/120). Github: https://github.com/... [im…
  • @bryancsk Bryan Cheong on x
    DeepSeek keeps doing this. Surfacing to say “here you go” then disappearing into deep waters.
  • @jenzhuscott Jen Zhu on x
    The whale 🐋 saves the world from AI Feudalism. The fact that it's a Chinese lab and team truly challenge people's prejudices. But those who know, know. Thankful for @deepseek_ai and each and everyone who's working and contributing on open-sourced AI 💙
  • @eliebakouch Elie on x
    deepseek math v2 is the first open source model to reach gold on IMO? and we get a tech report, what an amazing release [image]
  • @zjasper Jasper on x
    DeepSeek dropped their latest AI research again on a holiday. DeepSeek-Math-V2 is the first open AI model that can win gold at IMO 2025 and beat Gemini on IMO-ProofBench. They're using a generator-verifier architecture that feels like GAN in the early days. - first train a
  • @dorialexander Alexander Doria on x
    So DeepSeek-Math-V2. It could be subtitled: “how to train better verifiers?” and the bulk of it is simply... better data work and synth pipelines (even if all models are trained with RL). DeepSeek further distances itself from the initial promises of spontaneous
  • @alpindale Alpin on x
    First impressions with DeepSeek-Math-V2: it overthinks everything. I gave it my pre-print, and tasked it with solving an open lemma in it. It thought for 66 minutes (the sheer amount of tokens was too much for Open WebUI so it slowed down to a crawl), but it didn't manage to [ima…
  • @teortaxestex @teortaxestex on x
    This is why I don't think they'll be the first to “AGI”, but they will likely be the first to make it open source. They can replicate anything on a shoestring budget, given some time. Stealing fire from definitely-not-gods will continue until human autonomy improves. [image]
  • @simonw Simon Willison on x
    DeepSeek-Math-V2 means we now have an open weights (Apache 2) model that can achieve gold medal performance on this year's International Mathematical Olympiad - previously proprietary models from OpenAI and Google DeepMind had achieved that 689GB from Hugging Face!
  • @teortaxestex @teortaxestex on x
    There is a uniquely Promethean vibe in Wenfeng's project. Before DS-MoE, only frontier could do efficiency. Before DS-Math/Prover, only frontier could do Real math. Before DS-Prover V2, only frontier could do Putnam level. Before DS-Math V2, only frontier could do IMO Gold... [im…
  • @victormustar Victor M on x
    DeepSeek is silently building all the blocks they need to ship something crazy for R2 instead of rushing it 🐳🤩
  • @jenzhuscott Jen Zhu on x
    The whales 🐋 is back! DeepSeek-Math-V2: 685B-parameter math monster built on V3.2-Exp-Base, fully open under Apache 2.0 - 1st model to use a generator-verifier loop in training: writes proofs → verifier scores them → RL closes the loop for self-verifiable reasoning. - Focuses
  • @askperplexity @askperplexity on x
    🐋 The Whale is back!!  DeepSeek just dropped an IMO gold-medalist model.  On ProofBench-Advanced—where models prove formal mathematical theorems—GPT-5 scores 20%.  Gemini Deep Think IMO Gold hits 65.7%.  DeepSeek Math V2 (Heavy) scores 61.9%.  That's second place—but Gemini isn't…
  • r/math r on reddit
    The first open source model to reach gold on IMO: DeepSeekMath-V2
  • r/LocalLLaMA r on reddit
    deepseek-ai/DeepSeek-Math-V2  · Hugging Face
  • r/singularity r on reddit
    DeepSeek released DeepSeek-Math-V2