[Thread] An OpenAI researcher says the company's latest experimental reasoning LLM achieved gold medal-level performance on the 2025 International Math Olympiad
1/N I'm excited to share that our latest @OpenAI experimental reasoning LLM has achieved a longstanding grand challenge in AI: gold medal-level performance on the world's most prestigious math competition—the International Math Olympiad (IMO). [image]
@alexwei_ Alexander Wei
Context & Ripple Effects
This is a further step in a visible race to turn specialized mathematical reasoning into a general-model capability. OpenAI had previously reported strong results from o1 on an IMO qualifying exam, while DeepMind’s AlphaGeometry2 benchmark results highlighted a different, geometry-focused route to Olympiad performance.
The distinction matters: the report concerns an experimental system and a competition-level threshold, not a claim that mathematical reasoning is solved. Subsequent coverage that human contestants still outscored both leading labs’ models also keeps the benchmark’s remaining headroom in view.
First-order effects
- OpenAI gains a high-salience validation point for its experimental reasoning-model program, particularly against prior models that performed far less well on Olympiad-style problems.
- The IMO becomes a more consequential public benchmark for frontier labs’ reasoning claims, while users and evaluators will need to distinguish gold-level status from top overall performance.
Second-order effects
- DeepMind and other model developers face pressure to publish comparable end-to-end results, rather than relying only on narrower benchmarks such as geometry or qualifying-exam scores.
- Evaluation shifts toward independently administered, difficult problem sets and proof-quality assessment, because headline scores alone do not establish how reliably a model reasons across domains.
Third-order effects
- If repeated across rigorous benchmarks, advanced reasoning may become a primary basis for differentiating frontier models, alongside broad language ability and multimodal features.
- The race could also make evaluation methodology strategically important: labs’ claims will carry more weight when test conditions, tool access, and scoring are transparent enough for meaningful comparison.
The trend: Frontier AI competition is moving from fluent general-purpose models toward systems differentiated by measurable, high-stakes reasoning performance.
Related: Reasoning economics · AI industrialization · LLM · International Math Olympiad · OpenAI’s prior IMO qualifying-exam result · DeepMind’s AlphaGeometry2 benchmark result
Related Coverage
- OpenAI's gold medal performance on the International Math Olympiad. This feels notable to me. OpenAI research scientist Alexander Wei: Simon Willison's Weblog · Simon Willison
- OpenAI's gold medal performance on the International Math Olympiad O · Ben Werdmuller
- OpenAI says its next big model can bring home Math Olympiad gold: A turning point? The Indian Express
- OpenAI's Experimental Model Outscores Human Prodigies at Math Olympiad implicator.ai · Marcus Schuler
- OpenAI's Reasoning Model Wins Gold at 2025 IMO, GPT-5 Coming Soon Analytics India Magazine · Siddharth Jindal
- OpenAI's AI model achieves “gold medal” in Math Olympiad: All the details Moneycontrol
- OpenAI's experimental model achieved gold at the International Math Olympiad Engadget · Jackson Chen
- OpenAI wins gold at the International Math Olympiad. The new LLM surpasses the historic challenge of mathematical reasoning Data Studios ‧Exafin · Graziano Stefanelli
- OpenAI claims a breakthrough in LLM reasoning on complex math problems The Decoder · Matthias Bastian
- It is tempting to view the capability of current AI technology as a singular quantity: either a given task X is within the ability of current tools, or it is not. However, there is in fact a very wide spread in capability (several orders of magnitude) depending on what resources and assistance gives the tool, and how one reports their results. … @tao@mathstodon.xyz · Terence Tao
- I collected notes on OpenAI's announcement that an unreleased and unnamed model of theirs scored at a gold medal level in this year's International … Simon Willison
- OpenAI claims gold-medal performance at IMO 2025 Hacker News
Discussion
-
@teorth
Terence Tao
on bluesky
My thoughts on the crucial importance of methodology on self-reported AI performance on mathematics competitions, and my policy on commenting on such reports going forward: mathstodon.xyz/@tao/1148814...
-
@isaiahbishop
Isaiah Bishop
on bluesky
Tao's reaction. Also we dont really have visibility into the openAI or deepmind process only the validity (of the extremely weird proofs)
-
@ccanonne.github.io
Clément Canonne
on bluesky
Direct link to the thread (3 posts) by Terry Tao on Mastodon (no need for an account to access it): mathstodon.xyz/@tao/1148814...
-
@simonwillison.net
Simon Willison
on bluesky
Surprising result from OpenAI: one of their research models achieved a gold medal performance in this year's International Mathematical Olympiad /without/ using tools — Just a classic next-token-predicting LLM with a bunch of reinforcement learning layered on top — simonwilli…
-
@taumuyi
Tau-Mu Yi
on bluesky
This is a very impressive result. Last year Alpha Proof scored a silver medal on IMO 2024, but it was “fine-tuned” to the exam, whereas, according to the authors, the OpenAI model seemed to use more general training methods i.e. extensive RL on top of pretrained LLM. #AI [embed…
-
@timkellogg.me
Tim Kellogg
on bluesky
apparently google deepmind ai also got a gold metal on the math olympiad, but their marketing department wouldn't let them talk about it [image]
-
@cabernet
Andrea Lathrop
on bluesky
Since everyone is discussing whether OpenAI got “gold” on the International Math Olympiad (and how...) I will relay from X that the main guy on the project has stated 1) the OpenAI test was graded by “three former Olympiads” (not official scorers) and 2) only vaguely alluding to …
-
@skynetandchill.com
@skynetandchill.com
on bluesky
OpenAI says they have achieved gold medal-level AI performance on International Math Olympiad problems. My guess is, more like an AlphaGo project with deep search and reinforcement learning, not incremental improvements and scaling on reasoning models. Or hype/BS/training data …
-
@timkellogg.me
Tim Kellogg
on bluesky
openai researcher posts on X (not a blog or paper) about a model they have that can win the International Math Olympiad — you can't verify anything he says, but he's totally telling the truth [image]
-
@garymarcus
Gary Marcus
on x
Hot take on OpenAI's IMO gold What does it mean? I don't know (yet). The fact that no tools, coding or internet was used is genuinely impressive. That said, my overall impression is that OpenAI has told us the result, but not how it was achieved. That leaves me with many
-
@sebastienbubeck
Sebastien Bubeck
on x
It's hard to overstate the significance of this. It may end up looking like a “moon‑landing moment” for AI. Just to spell it out as clearly as possible: a next-word prediction machine (because that's really what it is here, no tools no nothing) just produced genuinely creative
-
@polynoamial
Noam Brown
on x
Today, we at @OpenAI achieved a milestone that many considered years away: gold medal-level performance on the 2025 IMO with a general reasoning LLM—under the same time limits as humans, without tools. As remarkable as that sounds, it's even more significant than the headline 🧵
-
@ernestryu
@ernestryu
on x
Two cents on AI getting International Math Olympiad (IMO) Gold, from a mathematician. Background: Last year, Google DeepMind (GDM) got Silver in IMO 2024. This year, OpenAI solved problems P1-P5 for IMO 2025 (but not P6), and this performance corresponds to Gold. (1/10)
-
@polynoamial
Noam Brown
on x
I think it's safe to say this @OpenAI IMO gold result came as a bit of a surprise to folks [image]
-
@garymarcus
Gary Marcus
on x
Hottest way to score cheap likes on Ai Twitter today is to quote me as saying the OpenAO IMO result is “impressive” (it is! and I really did say that!) without the full context. Here is the full context:
-
@garymarcus
Gary Marcus
on x
Is IMO really more prestigious that Putnam?
-
@polynoamial
Noam Brown
on x
It's truly a privilege to be able to wake up every morning, see where the latest intelligence frontier is, and help push it a little further.
-
@khoomeik
Rohan Pandey
on x
the RL algorithms to solve the most intellectually challenging environments are now here. the bottleneck will soon no longer be intelligence. so ask yourself, what is the highest value environment you can simulate?
-
@saranormous
@saranormous
on x
Congratulations @polynoamial and OpenAI's math reasoning team on their IMO Gold 🏆
-
@mihonarium
Mikhail Samin
on x
As someone who bet back in 2023 that that it's >70% likely AI will get an IMO gold medal by 2027: the IMO markets have been incredibly underpriced, especially for the past year. (Sadly, another prediction I've been >70% confident about is that AI will literally kill everyone.)
-
@elliotglazer
Elliot Glazer
on x
If you were shocked by OAI's IMO gold, you haven't been paying attention to all the signals that AI math capabilities are rapidly improving. o4-mini already had the requisite deductive prowess (as shown by FrontierMath performance) but insufficient in rigorous argumentation...
-
@omersarikaya
Ömer Sarıkaya
on x
Let's pause for a moment and reflect on the extraordinary effort required to raise a single IMO gold medalist kid. For a family, it involves years of financial investment alongside immense psychological pressure on the child, who must sacrifice a typical adolescence for
-
@kbeguir
Karim Beguir
on x
An LLM reaching gold-medal🏅at the IMO with no tool use confirms AGI is for this year, like I publicly predicted. New discoveries in math are now imminent. What a time to be alive!
-
@amolumd
Amol Deshpande
on x
Gold-level performance on IMO, solving 5 out of 6 problems, is incredible and quite unexpected. Looking forward to seeing more details...
-
@npcollapse
Connor Leahy
on x
Truly stunning instance of ironic prophecy (An OAI model that got gold on IMO was announced about 5 hours after this post)
-
@littmath
Daniel Litt
on x
Huge congrats to OpenAI for their IMO gold. I don't find it too surprising that an AI tool was able to achieve this (see below, though I'd sort of lost hope the last few days) but I'm pretty surprised it was an LRM with no tool use etc.
-
@lehoho248
@lehoho248
on x
Start of today: IMO 2025 AI models performed poorly. Gemini 2.5 Pro scored highest, 13/42. Hours later: OpenAI's @alexwei_ Model earned 35/42 points, solved 5/6 IMO 2025 problems and secured gold! 🥇 [image]
-
@feltsteam
@feltsteam
on x
@GaryMarcus It would be good if OAI were to release one big blog post over what happened, although given this from the IMO website it does seem like we might get a better report sometime soon. Also some DeepMind employees said they also got gold (to be announced¬ sure on metho…
-
@khoomeik
Rohan Pandey
on x
this IMO gold will fly past us as quickly as the turing test did soon normies will say “duh of course they're good at math, they're computers” but the RL breakthroughs the team made to solve math (congrats!!) will likely generalize to environments with much higher direct value
-
@wojtek_jk79848
@wojtek_jk79848
on x
“terence two was an imo gold medalist” Meanwhile terence tao on how elite maths competitions are useless:- It's high time Grindjeets realize that real maths isn't about brute strength and rote learning but understanding. High school maths isn't real maths. [image]
-
@daniel_mac8
Dan Mac
on x
🥇very elucidating thread on the significance of the ‘experimental reasoning model’ IMO Gold result from OpenAI
-
@lang__leon
Leon Lang
on x
I know many people are making fun of this evaluation today since it looks silly after OpenAI claimed IMO gold a short while later. But while that is funny, it's actually useful information to know where public models stand, and I'm glad the eval was done!
-
@peterwildeford
Peter Wildeford
on x
would be less misleading if you printed the entire graph I was definitely thinking AI IMO gold would happen this year (was close last year and FrontierMath results are suggestive of IMO gold)... not sure what brought the probability down in the final stretch [image]
-
@deedydas
Deedy
on x
Really awkward timing on this post... 12hrs after posting, OpenAI pulled off something amazing by getting the gold on IMO 2025 with pure reasoning / no internet. As a child, never thought this was possible in our lifetime.
-
@_vonarchimboldi
@_vonarchimboldi
on x
The model isn't public. The evaluations aren't public. Yet you have a bunch of OpenAI employees claiming their “Experimental Reasoning LLM” got gold level performance on the IMO. What is this? Not Science, for sure. Not even Science by demo. Science by PR?
-
@michael_nielsen
Michael Nielsen
on x
A very useful thread on the OpenAI Gold Medal IMO performance:
-
@zoink
Dylan Field
on x
Congrats 2025 IMO winners and participants, including OpenAI who trained a “general-purpose reinforcement learning model” and achieved IMO Gold! OpenAI team included @SherylHsu02 + @polynoamial. Fun fact: @polynoamial also won the 2025 Diplomacy World Championship! (As a human.)
-
@hindookissinger
@hindookissinger
on x
Saying that who cares about IMO when there are people working day and night to get their LLM models to crack the IMO lmao. It's a huge cope when people say things like IMO/etc. don't matter. Feynman was a Putnam medalist, Terrence Tao was an IMO gold medalist, etc.
-
@natolambert
Nathan Lambert
on x
Not falling for OpenAI's hype-vague posting about the new IMO gold model with “general purpose RL” and whatever else “breakthrough.” Google also got IMO gold (harder than mastering AIME), but remember, simple ideas scale best.
-
@taliaringer
Talia Ringer
on x
My biggest qualm with the IMO Gold Challenge was never with the idea that tools could do it within a few years, but rather with the idea that success on it implied something greater than tools being good at competition math
-
@polynoamial
Noam Brown
on x
Their bet allowed for formal math AI systems (like AlphaProof). In 2022, almost nobody thought an LLM could be IMO gold level by 2025.
-
@hangsiin
@hangsiin
on x
Read Noam's thread carefully. Winning a gold medal at the 2025 IMO is an outstanding achievement, but in some ways, it might just be noise that grabbed the headlines. They have recently developed new techniques that work much better on hard-to-verify problems, have extended TTC
-
@garymarcus
Gary Marcus
on x
Quote of the day: I certainly don't agree that machines which can solve IMO problems will be useful for mathematicians doing research, in the same way that when I arrived in Cambridge UK as an undergraduate clutching my IMO gold medal I was in no position to help any of the
-
@therealadamg
@therealadamg
on x
Models doing math. Not models using tools to do math.
-
@inductionheads
@inductionheads
on x
Gary Marcus and his neurosymbolic essay having a bad morning
-
@ziv_ravid
@ziv_ravid
on x
So, all the models underperform humans on the new International Mathematical Olympiad questions, and Grok-4 is especially bad on it, even with best-of-n selection? Unbelievable! [image]
-
@emostaque
Emad
on x
AGI is already here. All the components exist; we just need to stitch them together. It's Artificial General Intelligence, not “Artificial Top-Percentile Human Intelligence.” Two years ago, who would have said an IMO gold medal & topping benchmarks isn't AGI?
-
@neelnanda5
Neel Nanda
on x
Speaking as a past IMO contestant, this is impressive but misleading - gold vs silver is meaningless, 1 pt below gold vs borderline gold is noise The impressive bit is using a general reasoning model, not a specialised system, and no verified reward. Peak AI maths is unchanged
-
@kevinweil
Kevin Weil
on x
It is so so cool that an OpenAI model is now strong enough to win an IMO gold medal 🤯
-
@scaling01
@scaling01
on x
He can't be serious. He posted this right before OpenAI announced they got Gold in the IMO. Truly the Jim Cramer of AI [image]
-
@dejavucoder
Sankalp
on x
you are laughing? openai and google deepmind's unreleased models just gave an IMO gold medal performance without internet access and you are laughing? [image]
-
@simonw
Simon Willison
on x
The most notable thing about this result is that this unnamed experimental reasoning model achieved this score without any tool usage at all - it looks like it's just another classic next-token-predicting LLM with a bunch of reinforcement learning layered on top
-
@victortaelin
@victortaelin
on x
So I go sleep early and now we have AGI or something This sounds incredible but I can only wonder when this kind of tech will be available to all. Imagine leaving a model overnight working on Bend2. Would I wake up to the instant completion of all the hard tasks in our backlog?
-
@garymarcus
Gary Marcus
on x
All the tech bros this morning thinking that AGI has been achieved because some (insanely expensive) new form of LLMs can now match top *high school students* on one specific task ... it's almost ... cute! ☺️
-
@ilblackdragon
Illia
on x
AI winning gold in IMO is a huge deal. It was done without tools on new problems that haven't occurred in training data. Solving problems that most people in the world won't be able to solve. https://x.com/...
-
@michaeltrazzi
@michaeltrazzi
on x
Four years ago Paul Christiano thought this was 8% likely to happen Gosh even Eliezer didn't want to go further than 16% [image]
-
@preethilahoti
Preethi Lahoti
on x
What an exciting time to live in! Being able to witness AI capabilities unfold and to be a part of this field right now is truly special.
-
@prafdhar
Prafulla Dhariwal
on x
🏅medal performance at IMO using purely natural language reasoning, no tools or internet! was expecting this to take a few more years but the team has made such rapid progress, congrats @alexwei_ @SherylHsu02 @polynoamial and many others at @OpenAI on this amazing achievement!!
-
@andrewmayne
Andrew Mayne
on x
TLDR: The model solved complex math problems through reasoning alone. Until now the highest scoring LLMs used calculators and writing code. It's a big step forward that many were saying was impossible for these kinds of models as recently as....yesterday.
-
@thenanyu
Nan Yu
on x
Never change, hackernews [image]
-
@mihonarium
Mikhail Samin
on x
Paul Christiano was <8% of this happening. @ESYudkowsky was >16%. The market is currently at 93%. Sad congrats to Eliezer. [image]
-
@harshit_sikchi
Harshit Sikchi
on x
A big milestone;🥇in IMO under same human rules and GPT-5 ☕️
-
@dadicool
Dali Kilani
on x
Progress never stops. As the saying goes : “we tend to overestimate the impact of technology in the short term, and underestimate it in the long term”. In this case, the acceleration is wild and it's getting faster still. Next: an OSS model gets on the IMO podium :)
-
@alexwei_
Alexander Wei
on x
7/N HUGE congratulations to the team—@SherylHsu02, @polynoamial, and the many giants whose shoulders we stood on—for turning this crazy dream into reality! I am lucky I get to spend late nights and early mornings working alongside the very best.
-
@nrehiew_
@nrehiew_
on x
Takeaways + (guesses): 1) This is likely a multi agent system. So it isn't a single reasoner thinking for a million tokens in one go 2) (This likely doesn't use much training compute if at all) 3) They have a general purpose verifier beyond just rule based final answer [image]
-
@npew
Peter Welinder
on x
Another milestone reached in the pursuit of AGI: gold medal performance on the IMO.
-
@stalkermustang
Igor Kotenkov
on x
Asked Agent to help me guess what's next in line for OAI to flex: — Physics IPhO, Jul 24 — Economics IEO, Jul 29 — International Mathematics Competition for University Students (IMC), 3 Aug — IOI Aug 3 — Constraint‑Solver Competition, Aug 15 — ICPC Finals, Sen 5 lookin👀
-
@sherylhsu02
Sheryl Hsu
on x
Watching the model solve these IMO problems and achieve gold-level performance was magical. A few thoughts 🧵
-
@stevejarrett
Steve Jarrett
on x
Another huge step forward by @OpenAI in math. We have clearly NOT hit a plateau with existing techniques such as RL and test-time compute. A worthy next challenge would be achieving these results with far less compute and high levels of autonomy in the RL. Excited to see how far
-
@wickman
Brian Wickman
on x
inject this into my veins
-
@markchen90
Mark Chen
on x
We achieved gold medal level performance on this year's IMO! Our model thinks and writes proofs in clear, plain‑English - no formal code required. Unlike the narrower systems used in past competitions, our model is built to reason broadly, far beyond contest problems.
-
@0xsamgreen
Sam Green
on x
“We reach this capability level not via narrow, task-specific methodology, but by breaking new ground in general-purpose reinforcement learning and test-time compute scaling.”
-
@zhansheng
Jason Phang
on x
I'd like to take this chance to remind everyone that it hasn't even been a full year since o1 was announced (Sept 2024).
-
@tmychow
Trevor
on x
last summer, alphaproof + alphageometry2 combined to achieve a silver at the IMO yet @polymarket was only pricing a 25% chance of AI models achieving a gold this year this year's IMO just occurred, and @openai smashed it with a single model without tool use and got a gold! [image…
-
@mashah08
Mehul Shah
on x
The AI scaling that went on for the last five years is going to be very different from the scaling in the next. These models have latent capabilities that we are racing to unearth at inference time. IMO is but one example. The stakes are high and the race is on.
-
@kvallier
Kevin Vallier
on x
This is unbelievable. Friends, we must take AI seriously. We can no longer dismiss it with “Look at this one thing it can't do” or (worse) “HaLLuCiNaTi0nS!”
-
@mbalunovic
Mislav Balunović
on x
Congrats, this is amazing achievement and huge progress compared to public models such as o3 (which stays below bronze medal)
-
@clu_cheng
Cheng Lu
on x
Congrats! This is an incredible milestone and I was truly shocked by it. “Thinking for hours” means 10x or even 100x of current test-time compute, and I can't wait to see the model think for days, months, years, centuries to solve the science challenges!
-
@openai
@openai
on x
We achieved gold medal-level performance 🥇on the 2025 International Mathematical Olympiad with a general-purpose reasoning LLM! Our model solved world-class math problems—at the level of top human contestants. A major milestone for AI and mathematics.
-
@nyc_rivera
Mario Rivera
on x
A year after Alphaproof this new model from @OpenAI has reached the gold medal standard on the IMO. Incredible work!
-
@_ghorbani
Behrooz Ghorbani
on x
Congrats to @alexwei_ , @SherylHsu02 , @polynoamial , and the team for this truly remarkable result! It's a clear example of the rapid pace of AI progress!
-
@chombabupe
@chombabupe
on x
I am also exhilarated to share that my latest internal experimental model overfitted on MNIST, a handwritten digit recognition problem, 100% state-of-the-art performance. Humans only get 99.85% below my super duper model.
-
@burny_tech
Burny
on x
So public AI models are bad at IMO, while internal models are getting gold medals? Fascinating [image]
-
@tbpn
@tbpn
on x
FROM THE ARCHIVE: We asked Scott Wu (@ScottWu46) whether an AI would take gold at the International Mathematical Olympiad this year. “I'd be surprised if it doesn't... our internal bet is that the AI will win.” Yesterday that prediction landed. OpenAI researchers say an [video]
-
@orionjohnston
David Johnston
on x
Going to speculate that the innovation involves some method to control Yann's exponential divergence. If results are verifiable, you can control it by ensuring you have the right answer. If not, you need to make your intermediate steps reliable. 75%.
-
@hunterlightman
Hunter
on x
huge congrats to @alexwei_, @SherylHsu02, and @polynoamial for an incredible achievement and the close of a big chapter!! onto p6 and other even harder problems 😎 alex: you're the boss, man
-
@zzh8829
Zihao Zhang
on x
incredible results, the crazy part is LLM will be able to get gold medal every year from now on, since the competition won't get any harder.
-
@dmdohan
David Dohan
on x
OpenAI achieved gold medal on 2025 International Math Olympiad (solving 5 of 6 problems)! Thinks for hours and writes proofs in natural language. We've come a long way from LLMs solving 50% of MATH dataset in 2022 Congrats @alexwei_ on spearheading a major milestone!
-
@emollick
Ethan Mollick
on x
There are always a flood of posts about what AI can or cannot do, so it is worth pausing and paying attention to this one. It is a very hard test, done without tools. It was also viewed as an unlikely goal. Prediction markets had the chance of this happening this year as 20%
-
@albertwenger
Albert Wenger
on x
AGI is already here. It is just not yet in a single model.
-
@jcabreroholg
José Cabrero-Holgueras
on x
We are seeing gold medal-level performance on the math olympiad from a reasoning LLM. This is a major feat. It really highlights how far LLMs have come in just a short time. Not long ago, ChatGPT struggled to count the Rs in strawberry. The pace of progress is astonishing.
-
@g_leech_
Gavin Leech
on x
big. Unlike AlphaGeometry this one also supposedly stuck to the actual time limit (9 hours). quibbles: recall that o3-high took $3m per ARC-AGI eval run. You wonder what this took.
-
@arynbhar
Aryan Bhargav
on x
we are entering a new realm of mathematics
-
@sherwinwu
Sherwin Wu
on x
Two observations to the IMO gold result: (1) wow our research team is absolutely cracked (2) what a time to be building in this space! the capabilities overhang is still very real — imagine all the possibilities of how the tech behind an IMO-gold-level AI can change the world
-
@azi_pat
Pat Azi
on x
1) this model is smarter than GPT 5, but won't be released for “several months” after GPT 5 2) OpenAI claims to have leveraged a huge breakthrough - general purpose RL beyond Reinforcement Learning with Verifiable Rewards (RLVR)
-
@brij
Brij Singh
on x
OpenAI is firing on all cylinders- gold medal performance in IMO, best in class open source model in lmarena and soon gpt5. We might be closer to a hard take off
-
@aagarwal1012
Ayush Agarwal
on x
As someone who has participated in the International Maths Olympiad, I know firsthand how incredibly tough it is—just solving a couple of problems is already a huge accomplishment. If OpenAI's model is truly achieving gold medal-level results at IMO, that's a massive leap for AI
-
@davidblundin
David Blundin
on x
Friends and family, the reason this is such a big deal is this: The process of AI “self-improvement” is very similar to solving hard math problems like IMO. Before today, AI could help AI researchers innovate. After today, it's possible that the AI can improve its algorithms
-
@alexwei_
Alexander Wei
on x
4/N Second, IMO submissions are hard-to-verify, multi-page proofs. Progress here calls for going beyond the RL paradigm of clear-cut, verifiable rewards. By doing so, we've obtained a model that can craft intricate, watertight arguments at the level of human mathematicians. [imag…
-
@christiancooper
Christian H. Cooper
on x
I know it's hard to look past the environmental cost of compute at the moment, but this might be the best shot we have to fix the carbon problem. I'm really hopeful when I see things like this re IMO This can truly bend the curve. I think we're about 9 months away.
-
@afinetheorem
Kevin A. Bryan
on x
35/42 on IMO 2025, no tools/internet, as graded by 3 IMO medal winners. Model unreleased, “not coming for months”, so take caveat. But I always stress when talking about AI limits: the major labs have not seen them yet, so “lines on a graph” for, say, a year are baked in already.
-
@latticecut
Alastair Moore
on x
Another milestone
-
@gestaltu
Adam Butler
on x
This is incredible. Another step change in capability. Very different than the usual math benchmarks (which are also very challenging but also saturated at this point).
-
@nervouscomputer
@nervouscomputer
on x
everything @alexwei_ touches turns to gold
-
@jimdmiller
James Miller
on x
If AI masters math, we could get trillion-dollar breakthroughs—like room-temperature superconductors. Math models reality unreasonably well, so better math would open lots of doors.
-
@soohoonchoi
Soohoon Choi
on x
thank you @OpenAI 🙏 [image]
-
@gdb
Greg Brockman
on x
Gold medal-level performance on the 2025 International Math Olympiad from our latest experimental reasoning LLM. Model operated in natural language (i.e. outputs natural language proofs) under the same rules as humans (e.g. 4.5 hours per session, no tools). Amazing milestone!
-
@alexwei_
Alexander Wei
on x
3/N Why is this a big deal? First, IMO problems demand a new level of sustained creative thinking compared to past benchmarks. In reasoning time horizon, we've now progressed from GSM8K (~0.1 min for top humans) → MATH benchmark (~1 min) → AIME (~10 mins) → IMO (~100 mins).
-
@justjoshinyou13
Josh You
on x
. @GregHBurnham on how we should interpret a general-purpose LLM getting IMO gold https://epoch.ai/... [image]
-
@deredleritt3r
Prinz
on x
New OpenAI model achieves gold-medal-level performance on International Math Olympiad (IMO). According to Noam Brown, it's a “brand new [model], using recently developed techniques”. https://x.com/...
-
@greghburnham
Greg Burnham
on x
Pretty happy with how my predictions are holding up. 5/6 was the gold medal threshold this year. OAI's “experimental reasoning LLM” got that exactly, failing only to solve the one hard combinatorics problem, P6. My advice remains: look beyond the medal. Brief thread. 1/ [image]
-
@alexwei_
Alexander Wei
on x
6/N In our evaluation, the model solved 5 of the 6 problems on the 2025 IMO. For each problem, three former IMO medalists independently graded the model's submitted proof, with scores finalized after unanimous consensus. The model earned 35/42 points in total, enough for gold! 🥇
-
@alexwei_
Alexander Wei
on x
5/N Besides the result itself, I am excited about our approach: We reach this capability level not via narrow, task-specific methodology, but by breaking new ground in general-purpose reinforcement learning and test-time compute scaling.
-
@alexwei_
Alexander Wei
on x
2/N We evaluated our models on the 2025 IMO problems under the same rules as human contestants: two 4.5 hour exam sessions, no tools or internet, reading the official problem statements, and writing natural language proofs. [image]
-
@sama
Sam Altman
on x
we achieved gold medal level performance on the 2025 IMO competition with a general-purpose reasoning system! to emphasize, this is an LLM doing math and not a specific formal math system; it is part of our main push towards general intelligence. when we first started openai,
-
r/math
r
on reddit
Terence Tao on the supposed Gold from OpenAI at IMO
-
r/agi
r
on reddit
OpenAI claims Gold-medal performance at IMO 2025
-
r/slatestarcodex
r
on reddit
OpenAI claims gold medal performance at the 2025 International Math Olympiad