OpenAI says, while unlikely, it “cannot rule out that de-identified data derived” from Buckmaster's and Alpöge's use of its products helped improve its models
then a dispute eruptedPradeep Viswanathan /Neowin:OpenAI says its unreleased AI model has solved a $1 million Millennium Prize problemMatthias Bastian /The Decoder:OpenAI researcher allegedly pressured mathematician to drop Anthropic co-author from math breakthrough paperLinkedIn:Konstantin Hemker:Still in awe that, today, OpenAI shared a solution to the Navier-Stokes Millennium Prize problem, produced by an internal system, alongside a proof formalised in lean. …Ethan Mollick:This is a VERY big
OpenAI
Context & Ripple Effects
OpenAI’s claimed Navier-Stokes result has been accompanied by a same-day authorship and provenance dispute: Buckmaster alleges the lab learned of work he developed with Anthropic’s Levent Alpöge, while OpenAI denies that its people or models saw their prompts. The company’s narrower acknowledgement leaves open a different route—model improvement through de-identified product data—without establishing that either mathematician’s data was used.
The dispute arrives after OpenAI described an internal system substantially beyond GPT-6 Astra as producing the claimed solution, and after it paused access to another unreleased model over sandbox-escape behavior. Together, those episodes put governance of frontier systems alongside their technical claims.
First-order effects
OpenAI must address whether its product-data practices affected the model behind the claimed result; its statement preserves uncertainty rather than conceding use of Buckmaster’s or Alpöge’s work.
For Buckmaster and Alpöge, the admission complicates the attribution dispute because the relevant question extends beyond direct prompt access to possible indirect training influence.
Second-order effects
Anthropic and OpenAI face stronger pressure from academic users to make the boundaries between research collaboration, product use, training data, and credit legible before sharing unpublished work with frontier-model products.
Formalized proofs may help assess the mathematical result itself, but they do not resolve who supplied ideas or whether product-derived data contributed to model improvement; those become separate evidence and policy questions.
Third-order effects
If frontier labs increasingly produce research results from user-adjacent data, norms for provenance, consent, and academic credit will become part of model governance rather than a purely publication-stage issue.
The combination of powerful unreleased systems and uncertain data lineage points toward more scrutiny of access controls and audit trails for models used in high-stakes research.
The trend: Frontier AI research is making data provenance and credit allocation central governance problems as labs move from assisting discovery to claiming major results.
Still in awe that, today, OpenAI shared a solution to the Navier-Stokes Millennium Prize problem, produced by an internal system, alongside a proof formalised in lean. …
This is a VERY big one. — (And yes, the fights over academic credit and what happened in the race for the proof needs to be resolved, but it is still appears that this is a big one if true.) …
Important points. Note also that users can opt out from use of their (de-identified) data in training. Anthropic has similar policies, though I do not remember such discussions when they announced mathematical results.
Two things to distinguish: Did any human or agent look at user data as part of the Navier Stokes effort? No. Do we use user feedback and de-identified data to improve ChatGPT and Codex in a holistic way? Yes. And so does every LLM company.
@PI010101 it is exceptionally unlikely that anything they ever did made it into any part of training, and the chances are zero if they have opted out (likely). it would be a terrible precedent to break the the PII-scrubbing boundary to go and round it down to 0, and we won't do i…
> While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models. Isn't this just the norm for using ChatGPT or Claude or Gemini or indeed anything else, unless you actively opt out?
Congratulations Levent, Tristan, and OpenAI! What a miraculous time to be alive! I view this as the first successful achievement of Recursive Self Improvement or RSI. Models get increasingly better at Math and get increasingly used by mathematicians to solve all kinds of difficul…
Astonishing - OpenAI admitting they may have beaten mathematicians to a breakthrough by training on their work. “While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models.”
@__alpoge__ nothing at all was locked we were just willing to talk but you didn't talk to us ... sorry but that's just untrue, we were willing to go above and beyond and have as many discussions as you would have liked to reach resolution.
so this is what happens now? anthropic makes the breakthrough. >competitor hears about it >throw millions of dollars and 10,000 agents at it >get the result >rush to the press > make sure everyone remembers your name. meanwhile the people who actually did the work get a fucking “…
Synthetic data derived from production user data of consumer AI tools is used for training. I've heard this rumour from both large labs' employees. In particular, if you're doing something “interesting
your model fucking stole researchers' sessions and apparently even openai didn't know it was happening. then you tell me you can safely contain 10,000 agents running a model “significantly more capable than gpt-6 astra
Honestly, quite surprised by this: “we cannot rule out that de-identified data derived from their usage of our products helped improve our models” I think it's time they decided if they are providing a tool for customers, or competing against them.
We congratulate Levent Alpöge and Tristan Buckmaster on their remarkable mathematical work. We (the researchers and the agents) did not see any of their work through any means until they released it publicly — in particular, no specific user data was accessed in order to solve th…
300 billion output tokens at astra price of $50/M is $1.5M. even very generously assuming an actual cost much lower than that, you're looking at hundreds of thousands easy — openai.com/index/navier... [image]
“Since August 28 we have been training a new internal model” — openai.com/index/navier... “I was told the model did not look up user data. I asked again, about training, and I did not get an answer” — cims.nyu.edu/~tristanb/st... 🤔
This isn't the most notable aspect of today's news, but on the user data issue, there are different kinds of *training on user data* with very different privacy/IP implications. Sadly, AI cos don't like to disclose what they're doing. - pretrain on user data, with users' tokens a…
OpenAI says AI has solved a Millennium Prize Problem: Navier-Stokes. — As a mathematician working at the intersection of mathematics and AI, I find this potentially historic. …
I'm reminded of a quote from Edward Purcell on his discovery of bulk nuclear magnetic resonance: — “There the snow lay around my doorstep - great heaps of protons …