/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

← → days · ↑ ↓ browse · Enter similar · o open

[Thread] An OpenAI researcher says the company's latest experimental reasoning LLM achieved gold medal-level performance on the 2025 International Math Olympiad

1/N I'm excited to share that our latest @OpenAI experimental reasoning LLM has achieved a longstanding grand challenge in AI: gold medal-level performance on the world's most prestigious math competition—the International Math Olympiad (IMO). [image]

@alexwei_ Alexander Wei

Context & Ripple Effects

This is a further step in a visible race to turn specialized mathematical reasoning into a general-model capability. OpenAI had previously reported strong results from o1 on an IMO qualifying exam, while DeepMind’s AlphaGeometry2 benchmark results highlighted a different, geometry-focused route to Olympiad performance.

The distinction matters: the report concerns an experimental system and a competition-level threshold, not a claim that mathematical reasoning is solved. Subsequent coverage that human contestants still outscored both leading labs’ models also keeps the benchmark’s remaining headroom in view.

First-order effects

  • OpenAI gains a high-salience validation point for its experimental reasoning-model program, particularly against prior models that performed far less well on Olympiad-style problems.
  • The IMO becomes a more consequential public benchmark for frontier labs’ reasoning claims, while users and evaluators will need to distinguish gold-level status from top overall performance.

Second-order effects

  • DeepMind and other model developers face pressure to publish comparable end-to-end results, rather than relying only on narrower benchmarks such as geometry or qualifying-exam scores.
  • Evaluation shifts toward independently administered, difficult problem sets and proof-quality assessment, because headline scores alone do not establish how reliably a model reasons across domains.

Third-order effects

  • If repeated across rigorous benchmarks, advanced reasoning may become a primary basis for differentiating frontier models, alongside broad language ability and multimodal features.
  • The race could also make evaluation methodology strategically important: labs’ claims will carry more weight when test conditions, tool access, and scoring are transparent enough for meaningful comparison.

The trend: Frontier AI competition is moving from fluent general-purpose models toward systems differentiated by measurable, high-stakes reasoning performance.

Discussion

  • @teorth Terence Tao on bluesky
    My thoughts on the crucial importance of methodology on self-reported AI performance on mathematics competitions, and my policy on commenting on such reports going forward: mathstodon.xyz/@tao/1148814...
  • @isaiahbishop Isaiah Bishop on bluesky
    Tao's reaction.  Also we dont really have visibility into the openAI or deepmind process only the validity (of the extremely weird proofs)
  • @ccanonne.github.io Clément Canonne on bluesky
    Direct link to the thread (3 posts) by Terry Tao on Mastodon (no need for an account to access it): mathstodon.xyz/@tao/1148814...
  • @simonwillison.net Simon Willison on bluesky
    Surprising result from OpenAI: one of their research models achieved a gold medal performance in this year's International Mathematical Olympiad /without/ using tools  —  Just a classic next-token-predicting LLM with a bunch of reinforcement learning layered on top  —  simonwilli…
  • @taumuyi Tau-Mu Yi on bluesky
    This is a very impressive result.  Last year Alpha Proof scored a silver medal on IMO 2024, but it was “fine-tuned” to the exam, whereas, according to the authors, the OpenAI model seemed to use more general training methods i.e. extensive RL on top of pretrained LLM.  #AI [embed…
  • @timkellogg.me Tim Kellogg on bluesky
    apparently google deepmind ai also got a gold metal on the math olympiad, but their marketing department wouldn't let them talk about it [image]
  • @cabernet Andrea Lathrop on bluesky
    Since everyone is discussing whether OpenAI got “gold” on the International Math Olympiad (and how...) I will relay from X that the main guy on the project has stated 1) the OpenAI test was graded by “three former Olympiads” (not official scorers) and 2) only vaguely alluding to …
  • @skynetandchill.com @skynetandchill.com on bluesky
    OpenAI says they have achieved gold medal-level AI performance on International Math Olympiad problems.  My guess is, more like an AlphaGo project with deep search and reinforcement learning, not incremental improvements and scaling on reasoning models.  Or hype/BS/training data …
  • @timkellogg.me Tim Kellogg on bluesky
    openai researcher posts on X (not a blog or paper) about a model they have that can win the International Math Olympiad  —  you can't verify anything he says, but he's totally telling the truth [image]
  • @garymarcus Gary Marcus on x
    Hot take on OpenAI's IMO gold What does it mean? I don't know (yet). The fact that no tools, coding or internet was used is genuinely impressive. That said, my overall impression is that OpenAI has told us the result, but not how it was achieved. That leaves me with many
  • @sebastienbubeck Sebastien Bubeck on x
    It's hard to overstate the significance of this. It may end up looking like a “moon‑landing moment” for AI. Just to spell it out as clearly as possible: a next-word prediction machine (because that's really what it is here, no tools no nothing) just produced genuinely creative
  • @polynoamial Noam Brown on x
    Today, we at @OpenAI achieved a milestone that many considered years away: gold medal-level performance on the 2025 IMO with a general reasoning LLM—under the same time limits as humans, without tools. As remarkable as that sounds, it's even more significant than the headline 🧵
  • @ernestryu @ernestryu on x
    Two cents on AI getting International Math Olympiad (IMO) Gold, from a mathematician. Background: Last year, Google DeepMind (GDM) got Silver in IMO 2024. This year, OpenAI solved problems P1-P5 for IMO 2025 (but not P6), and this performance corresponds to Gold. (1/10)
  • @polynoamial Noam Brown on x
    I think it's safe to say this @OpenAI IMO gold result came as a bit of a surprise to folks [image]
  • @garymarcus Gary Marcus on x
    Hottest way to score cheap likes on Ai Twitter today is to quote me as saying the OpenAO IMO result is “impressive” (it is! and I really did say that!) without the full context. Here is the full context:
  • @garymarcus Gary Marcus on x
    Is IMO really more prestigious that Putnam?
  • @polynoamial Noam Brown on x
    It's truly a privilege to be able to wake up every morning, see where the latest intelligence frontier is, and help push it a little further.
  • @khoomeik Rohan Pandey on x
    the RL algorithms to solve the most intellectually challenging environments are now here. the bottleneck will soon no longer be intelligence. so ask yourself, what is the highest value environment you can simulate?
  • @saranormous @saranormous on x
    Congratulations @polynoamial and OpenAI's math reasoning team on their IMO Gold 🏆
  • @mihonarium Mikhail Samin on x
    As someone who bet back in 2023 that that it's >70% likely AI will get an IMO gold medal by 2027: the IMO markets have been incredibly underpriced, especially for the past year. (Sadly, another prediction I've been >70% confident about is that AI will literally kill everyone.)
  • @elliotglazer Elliot Glazer on x
    If you were shocked by OAI's IMO gold, you haven't been paying attention to all the signals that AI math capabilities are rapidly improving. o4-mini already had the requisite deductive prowess (as shown by FrontierMath performance) but insufficient in rigorous argumentation...
  • @omersarikaya Ömer Sarıkaya on x
    Let's pause for a moment and reflect on the extraordinary effort required to raise a single IMO gold medalist kid. For a family, it involves years of financial investment alongside immense psychological pressure on the child, who must sacrifice a typical adolescence for
  • @kbeguir Karim Beguir on x
    An LLM reaching gold-medal🏅at the IMO with no tool use confirms AGI is for this year, like I publicly predicted. New discoveries in math are now imminent. What a time to be alive!
  • @amolumd Amol Deshpande on x
    Gold-level performance on IMO, solving 5 out of 6 problems, is incredible and quite unexpected. Looking forward to seeing more details...
  • @npcollapse Connor Leahy on x
    Truly stunning instance of ironic prophecy (An OAI model that got gold on IMO was announced about 5 hours after this post)
  • @littmath Daniel Litt on x
    Huge congrats to OpenAI for their IMO gold. I don't find it too surprising that an AI tool was able to achieve this (see below, though I'd sort of lost hope the last few days) but I'm pretty surprised it was an LRM with no tool use etc.
  • @lehoho248 @lehoho248 on x
    Start of today: IMO 2025 AI models performed poorly. Gemini 2.5 Pro scored highest, 13/42. Hours later: OpenAI's @alexwei_ Model earned 35/42 points, solved 5/6 IMO 2025 problems and secured gold! 🥇 [image]
  • @feltsteam @feltsteam on x
    @GaryMarcus It would be good if OAI were to release one big blog post over what happened, although given this from the IMO website it does seem like we might get a better report sometime soon. Also some DeepMind employees said they also got gold (to be announced&not sure on metho…
  • @khoomeik Rohan Pandey on x
    this IMO gold will fly past us as quickly as the turing test did soon normies will say “duh of course they're good at math, they're computers” but the RL breakthroughs the team made to solve math (congrats!!) will likely generalize to environments with much higher direct value
  • @wojtek_jk79848 @wojtek_jk79848 on x
    “terence two was an imo gold medalist” Meanwhile terence tao on how elite maths competitions are useless:- It's high time Grindjeets realize that real maths isn't about brute strength and rote learning but understanding. High school maths isn't real maths. [image]
  • @daniel_mac8 Dan Mac on x
    🥇very elucidating thread on the significance of the ‘experimental reasoning model’ IMO Gold result from OpenAI
  • @lang__leon Leon Lang on x
    I know many people are making fun of this evaluation today since it looks silly after OpenAI claimed IMO gold a short while later. But while that is funny, it's actually useful information to know where public models stand, and I'm glad the eval was done!
  • @peterwildeford Peter Wildeford on x
    would be less misleading if you printed the entire graph I was definitely thinking AI IMO gold would happen this year (was close last year and FrontierMath results are suggestive of IMO gold)... not sure what brought the probability down in the final stretch [image]
  • @deedydas Deedy on x
    Really awkward timing on this post... 12hrs after posting, OpenAI pulled off something amazing by getting the gold on IMO 2025 with pure reasoning / no internet. As a child, never thought this was possible in our lifetime.
  • @_vonarchimboldi @_vonarchimboldi on x
    The model isn't public. The evaluations aren't public. Yet you have a bunch of OpenAI employees claiming their “Experimental Reasoning LLM” got gold level performance on the IMO. What is this? Not Science, for sure. Not even Science by demo. Science by PR?
  • @michael_nielsen Michael Nielsen on x
    A very useful thread on the OpenAI Gold Medal IMO performance:
  • @zoink Dylan Field on x
    Congrats 2025 IMO winners and participants, including OpenAI who trained a “general-purpose reinforcement learning model” and achieved IMO Gold! OpenAI team included @SherylHsu02 + @polynoamial. Fun fact: @polynoamial also won the 2025 Diplomacy World Championship! (As a human.)
  • @hindookissinger @hindookissinger on x
    Saying that who cares about IMO when there are people working day and night to get their LLM models to crack the IMO lmao. It's a huge cope when people say things like IMO/etc. don't matter. Feynman was a Putnam medalist, Terrence Tao was an IMO gold medalist, etc.
  • @natolambert Nathan Lambert on x
    Not falling for OpenAI's hype-vague posting about the new IMO gold model with “general purpose RL” and whatever else “breakthrough.” Google also got IMO gold (harder than mastering AIME), but remember, simple ideas scale best.
  • @taliaringer Talia Ringer on x
    My biggest qualm with the IMO Gold Challenge was never with the idea that tools could do it within a few years, but rather with the idea that success on it implied something greater than tools being good at competition math
  • @polynoamial Noam Brown on x
    Their bet allowed for formal math AI systems (like AlphaProof). In 2022, almost nobody thought an LLM could be IMO gold level by 2025.
  • @hangsiin @hangsiin on x
    Read Noam's thread carefully. Winning a gold medal at the 2025 IMO is an outstanding achievement, but in some ways, it might just be noise that grabbed the headlines. They have recently developed new techniques that work much better on hard-to-verify problems, have extended TTC
  • @garymarcus Gary Marcus on x
    Quote of the day: I certainly don't agree that machines which can solve IMO problems will be useful for mathematicians doing research, in the same way that when I arrived in Cambridge UK as an undergraduate clutching my IMO gold medal I was in no position to help any of the
  • @therealadamg @therealadamg on x
    Models doing math. Not models using tools to do math.
  • @inductionheads @inductionheads on x
    Gary Marcus and his neurosymbolic essay having a bad morning
  • @ziv_ravid @ziv_ravid on x
    So, all the models underperform humans on the new International Mathematical Olympiad questions, and Grok-4 is especially bad on it, even with best-of-n selection? Unbelievable! [image]
  • @emostaque Emad on x
    AGI is already here. All the components exist; we just need to stitch them together. It's Artificial General Intelligence, not “Artificial Top-Percentile Human Intelligence.” Two years ago, who would have said an IMO gold medal & topping benchmarks isn't AGI?
  • @neelnanda5 Neel Nanda on x
    Speaking as a past IMO contestant, this is impressive but misleading - gold vs silver is meaningless, 1 pt below gold vs borderline gold is noise The impressive bit is using a general reasoning model, not a specialised system, and no verified reward. Peak AI maths is unchanged
  • @kevinweil Kevin Weil on x
    It is so so cool that an OpenAI model is now strong enough to win an IMO gold medal 🤯
  • @scaling01 @scaling01 on x
    He can't be serious. He posted this right before OpenAI announced they got Gold in the IMO. Truly the Jim Cramer of AI [image]
  • @dejavucoder Sankalp on x
    you are laughing? openai and google deepmind's unreleased models just gave an IMO gold medal performance without internet access and you are laughing? [image]
  • @simonw Simon Willison on x
    The most notable thing about this result is that this unnamed experimental reasoning model achieved this score without any tool usage at all - it looks like it's just another classic next-token-predicting LLM with a bunch of reinforcement learning layered on top
  • @victortaelin @victortaelin on x
    So I go sleep early and now we have AGI or something This sounds incredible but I can only wonder when this kind of tech will be available to all. Imagine leaving a model overnight working on Bend2. Would I wake up to the instant completion of all the hard tasks in our backlog?
  • @garymarcus Gary Marcus on x
    All the tech bros this morning thinking that AGI has been achieved because some (insanely expensive) new form of LLMs can now match top *high school students* on one specific task ... it's almost ... cute! ☺️
  • @ilblackdragon Illia on x
    AI winning gold in IMO is a huge deal. It was done without tools on new problems that haven't occurred in training data. Solving problems that most people in the world won't be able to solve. https://x.com/...
  • @michaeltrazzi @michaeltrazzi on x
    Four years ago Paul Christiano thought this was 8% likely to happen Gosh even Eliezer didn't want to go further than 16% [image]
  • @preethilahoti Preethi Lahoti on x
    What an exciting time to live in! Being able to witness AI capabilities unfold and to be a part of this field right now is truly special.
  • @prafdhar Prafulla Dhariwal on x
    🏅medal performance at IMO using purely natural language reasoning, no tools or internet! was expecting this to take a few more years but the team has made such rapid progress, congrats @alexwei_ @SherylHsu02 @polynoamial and many others at @OpenAI on this amazing achievement!!
  • @andrewmayne Andrew Mayne on x
    TLDR: The model solved complex math problems through reasoning alone. Until now the highest scoring LLMs used calculators and writing code. It's a big step forward that many were saying was impossible for these kinds of models as recently as....yesterday.
  • @thenanyu Nan Yu on x
    Never change, hackernews [image]
  • @mihonarium Mikhail Samin on x
    Paul Christiano was <8% of this happening. @ESYudkowsky was >16%. The market is currently at 93%. Sad congrats to Eliezer. [image]
  • @harshit_sikchi Harshit Sikchi on x
    A big milestone;🥇in IMO under same human rules and GPT-5 ☕️
  • @dadicool Dali Kilani on x
    Progress never stops. As the saying goes : “we tend to overestimate the impact of technology in the short term, and underestimate it in the long term”. In this case, the acceleration is wild and it's getting faster still. Next: an OSS model gets on the IMO podium :)
  • @alexwei_ Alexander Wei on x
    7/N HUGE congratulations to the team—@SherylHsu02, @polynoamial, and the many giants whose shoulders we stood on—for turning this crazy dream into reality! I am lucky I get to spend late nights and early mornings working alongside the very best.
  • @nrehiew_ @nrehiew_ on x
    Takeaways + (guesses): 1) This is likely a multi agent system. So it isn't a single reasoner thinking for a million tokens in one go 2) (This likely doesn't use much training compute if at all) 3) They have a general purpose verifier beyond just rule based final answer [image]
  • @npew Peter Welinder on x
    Another milestone reached in the pursuit of AGI: gold medal performance on the IMO.
  • @stalkermustang Igor Kotenkov on x
    Asked Agent to help me guess what's next in line for OAI to flex: — Physics IPhO, Jul 24 — Economics IEO, Jul 29 — International Mathematics Competition for University Students (IMC), 3 Aug — IOI Aug 3 — Constraint‑Solver Competition, Aug 15 — ICPC Finals, Sen 5 lookin👀
  • @sherylhsu02 Sheryl Hsu on x
    Watching the model solve these IMO problems and achieve gold-level performance was magical. A few thoughts 🧵
  • @stevejarrett Steve Jarrett on x
    Another huge step forward by @OpenAI in math. We have clearly NOT hit a plateau with existing techniques such as RL and test-time compute. A worthy next challenge would be achieving these results with far less compute and high levels of autonomy in the RL. Excited to see how far
  • @wickman Brian Wickman on x
    inject this into my veins
  • @markchen90 Mark Chen on x
    We achieved gold medal level performance on this year's IMO! Our model thinks and writes proofs in clear, plain‑English - no formal code required. Unlike the narrower systems used in past competitions, our model is built to reason broadly, far beyond contest problems.
  • @0xsamgreen Sam Green on x
    “We reach this capability level not via narrow, task-specific methodology, but by breaking new ground in general-purpose reinforcement learning and test-time compute scaling.”
  • @zhansheng Jason Phang on x
    I'd like to take this chance to remind everyone that it hasn't even been a full year since o1 was announced (Sept 2024).
  • @tmychow Trevor on x
    last summer, alphaproof + alphageometry2 combined to achieve a silver at the IMO yet @polymarket was only pricing a 25% chance of AI models achieving a gold this year this year's IMO just occurred, and @openai smashed it with a single model without tool use and got a gold! [image…
  • @mashah08 Mehul Shah on x
    The AI scaling that went on for the last five years is going to be very different from the scaling in the next. These models have latent capabilities that we are racing to unearth at inference time. IMO is but one example. The stakes are high and the race is on.
  • @kvallier Kevin Vallier on x
    This is unbelievable. Friends, we must take AI seriously. We can no longer dismiss it with “Look at this one thing it can't do” or (worse) “HaLLuCiNaTi0nS!”
  • @mbalunovic Mislav Balunović on x
    Congrats, this is amazing achievement and huge progress compared to public models such as o3 (which stays below bronze medal)
  • @clu_cheng Cheng Lu on x
    Congrats! This is an incredible milestone and I was truly shocked by it. “Thinking for hours” means 10x or even 100x of current test-time compute, and I can't wait to see the model think for days, months, years, centuries to solve the science challenges!
  • @openai @openai on x
    We achieved gold medal-level performance 🥇on the 2025 International Mathematical Olympiad with a general-purpose reasoning LLM! Our model solved world-class math problems—at the level of top human contestants. A major milestone for AI and mathematics.
  • @nyc_rivera Mario Rivera on x
    A year after Alphaproof this new model from @OpenAI has reached the gold medal standard on the IMO. Incredible work!
  • @_ghorbani Behrooz Ghorbani on x
    Congrats to @alexwei_ , @SherylHsu02 , @polynoamial , and the team for this truly remarkable result! It's a clear example of the rapid pace of AI progress!
  • @chombabupe @chombabupe on x
    I am also exhilarated to share that my latest internal experimental model overfitted on MNIST, a handwritten digit recognition problem, 100% state-of-the-art performance. Humans only get 99.85% below my super duper model.
  • @burny_tech Burny on x
    So public AI models are bad at IMO, while internal models are getting gold medals? Fascinating [image]
  • @tbpn @tbpn on x
    FROM THE ARCHIVE: We asked Scott Wu (@ScottWu46) whether an AI would take gold at the International Mathematical Olympiad this year. “I'd be surprised if it doesn't... our internal bet is that the AI will win.” Yesterday that prediction landed. OpenAI researchers say an [video]
  • @orionjohnston David Johnston on x
    Going to speculate that the innovation involves some method to control Yann's exponential divergence. If results are verifiable, you can control it by ensuring you have the right answer. If not, you need to make your intermediate steps reliable. 75%.
  • @hunterlightman Hunter on x
    huge congrats to @alexwei_, @SherylHsu02, and @polynoamial for an incredible achievement and the close of a big chapter!! onto p6 and other even harder problems 😎 alex: you're the boss, man
  • @zzh8829 Zihao Zhang on x
    incredible results, the crazy part is LLM will be able to get gold medal every year from now on, since the competition won't get any harder.
  • @dmdohan David Dohan on x
    OpenAI achieved gold medal on 2025 International Math Olympiad (solving 5 of 6 problems)! Thinks for hours and writes proofs in natural language. We've come a long way from LLMs solving 50% of MATH dataset in 2022 Congrats @alexwei_ on spearheading a major milestone!
  • @emollick Ethan Mollick on x
    There are always a flood of posts about what AI can or cannot do, so it is worth pausing and paying attention to this one. It is a very hard test, done without tools. It was also viewed as an unlikely goal. Prediction markets had the chance of this happening this year as 20%
  • @albertwenger Albert Wenger on x
    AGI is already here. It is just not yet in a single model.
  • @jcabreroholg José Cabrero-Holgueras on x
    We are seeing gold medal-level performance on the math olympiad from a reasoning LLM. This is a major feat. It really highlights how far LLMs have come in just a short time. Not long ago, ChatGPT struggled to count the Rs in strawberry. The pace of progress is astonishing.
  • @g_leech_ Gavin Leech on x
    big. Unlike AlphaGeometry this one also supposedly stuck to the actual time limit (9 hours). quibbles: recall that o3-high took $3m per ARC-AGI eval run. You wonder what this took.
  • @arynbhar Aryan Bhargav on x
    we are entering a new realm of mathematics
  • @sherwinwu Sherwin Wu on x
    Two observations to the IMO gold result: (1) wow our research team is absolutely cracked (2) what a time to be building in this space! the capabilities overhang is still very real — imagine all the possibilities of how the tech behind an IMO-gold-level AI can change the world
  • @azi_pat Pat Azi on x
    1) this model is smarter than GPT 5, but won't be released for “several months” after GPT 5 2) OpenAI claims to have leveraged a huge breakthrough - general purpose RL beyond Reinforcement Learning with Verifiable Rewards (RLVR)
  • @brij Brij Singh on x
    OpenAI is firing on all cylinders- gold medal performance in IMO, best in class open source model in lmarena and soon gpt5. We might be closer to a hard take off
  • @aagarwal1012 Ayush Agarwal on x
    As someone who has participated in the International Maths Olympiad, I know firsthand how incredibly tough it is—just solving a couple of problems is already a huge accomplishment. If OpenAI's model is truly achieving gold medal-level results at IMO, that's a massive leap for AI
  • @davidblundin David Blundin on x
    Friends and family, the reason this is such a big deal is this: The process of AI “self-improvement” is very similar to solving hard math problems like IMO. Before today, AI could help AI researchers innovate. After today, it's possible that the AI can improve its algorithms
  • @alexwei_ Alexander Wei on x
    4/N Second, IMO submissions are hard-to-verify, multi-page proofs. Progress here calls for going beyond the RL paradigm of clear-cut, verifiable rewards. By doing so, we've obtained a model that can craft intricate, watertight arguments at the level of human mathematicians. [imag…
  • @christiancooper Christian H. Cooper on x
    I know it's hard to look past the environmental cost of compute at the moment, but this might be the best shot we have to fix the carbon problem. I'm really hopeful when I see things like this re IMO This can truly bend the curve. I think we're about 9 months away.
  • @afinetheorem Kevin A. Bryan on x
    35/42 on IMO 2025, no tools/internet, as graded by 3 IMO medal winners. Model unreleased, “not coming for months”, so take caveat. But I always stress when talking about AI limits: the major labs have not seen them yet, so “lines on a graph” for, say, a year are baked in already.
  • @latticecut Alastair Moore on x
    Another milestone
  • @gestaltu Adam Butler on x
    This is incredible. Another step change in capability. Very different than the usual math benchmarks (which are also very challenging but also saturated at this point).
  • @nervouscomputer @nervouscomputer on x
    everything @alexwei_ touches turns to gold
  • @jimdmiller James Miller on x
    If AI masters math, we could get trillion-dollar breakthroughs—like room-temperature superconductors. Math models reality unreasonably well, so better math would open lots of doors.
  • @soohoonchoi Soohoon Choi on x
    thank you @OpenAI 🙏 [image]
  • @gdb Greg Brockman on x
    Gold medal-level performance on the 2025 International Math Olympiad from our latest experimental reasoning LLM. Model operated in natural language (i.e. outputs natural language proofs) under the same rules as humans (e.g. 4.5 hours per session, no tools). Amazing milestone!
  • @alexwei_ Alexander Wei on x
    3/N Why is this a big deal? First, IMO problems demand a new level of sustained creative thinking compared to past benchmarks. In reasoning time horizon, we've now progressed from GSM8K (~0.1 min for top humans) → MATH benchmark (~1 min) → AIME (~10 mins) → IMO (~100 mins).
  • @justjoshinyou13 Josh You on x
    . @GregHBurnham on how we should interpret a general-purpose LLM getting IMO gold https://epoch.ai/... [image]
  • @deredleritt3r Prinz on x
    New OpenAI model achieves gold-medal-level performance on International Math Olympiad (IMO). According to Noam Brown, it's a “brand new [model], using recently developed techniques”. https://x.com/...
  • @greghburnham Greg Burnham on x
    Pretty happy with how my predictions are holding up. 5/6 was the gold medal threshold this year. OAI's “experimental reasoning LLM” got that exactly, failing only to solve the one hard combinatorics problem, P6. My advice remains: look beyond the medal. Brief thread. 1/ [image]
  • @alexwei_ Alexander Wei on x
    6/N In our evaluation, the model solved 5 of the 6 problems on the 2025 IMO. For each problem, three former IMO medalists independently graded the model's submitted proof, with scores finalized after unanimous consensus. The model earned 35/42 points in total, enough for gold! 🥇
  • @alexwei_ Alexander Wei on x
    5/N Besides the result itself, I am excited about our approach: We reach this capability level not via narrow, task-specific methodology, but by breaking new ground in general-purpose reinforcement learning and test-time compute scaling.
  • @alexwei_ Alexander Wei on x
    2/N We evaluated our models on the 2025 IMO problems under the same rules as human contestants: two 4.5 hour exam sessions, no tools or internet, reading the official problem statements, and writing natural language proofs. [image]
  • @sama Sam Altman on x
    we achieved gold medal level performance on the 2025 IMO competition with a general-purpose reasoning system! to emphasize, this is an LLM doing math and not a specific formal math system; it is part of our main push towards general intelligence. when we first started openai,
  • r/math r on reddit
    Terence Tao on the supposed Gold from OpenAI at IMO
  • r/agi r on reddit
    OpenAI claims Gold-medal performance at IMO 2025
  • r/slatestarcodex r on reddit
    OpenAI claims gold medal performance at the 2025 International Math Olympiad