OpenAI says individuals associated with Moonshot AI played a significant role in a coordinated model-distillation campaign in July that peaked at 16K requests
OpenAI accused its Chinese rival Moonshot AI of being responsible for a wide-scale effort to extract data from its GPT artificial intelligence systems …
BloombergMaggie Eastland
Context & Ripple Effects
OpenAI, Anthropic and Google were already sharing information through the Frontier Model Forum to identify adversarial distillation attempts that breach their terms. OpenAI’s new attribution turns that collective defensive concern into a named competitive-security dispute.
The allegation also arrives after a U.S. official said Moonshot AI had obtained access to advanced AI infrastructure, making the protection of model behavior—not only the acquisition of compute—a consequential part of the contest between frontier labs.
First-order effects
OpenAI has a concrete basis to pursue account enforcement and strengthen protections around GPT interactions it says were used to extract model data.
Moonshot AI faces a public allegation that individuals associated with it participated in a terms-violating campaign, raising the reputational and commercial stakes of its relationship with API and infrastructure partners.
Second-order effects
The prior cross-lab information-sharing effort becomes more operationally important as OpenAI, Anthropic and Google can compare indicators of coordinated extraction rather than defend endpoints in isolation.
API operators and enterprise customers using frontier models face tighter scrutiny of high-volume, patterned requests, as providers seek to distinguish legitimate evaluation from extraction behavior.
Third-order effects
If coordinated distillation campaigns persist, frontier-model access is likely to be governed more like a protected strategic capability: usage controls, monitoring and inter-lab threat sharing become part of the product rather than back-office security.
The dispute sharpens a durable tension in AI competition between widely available model interfaces and labs’ efforts to prevent those interfaces from transferring proprietary capabilities to rivals.
The trend: Frontier AI labs are treating model outputs and reasoning behavior as strategic assets that require collective security controls, not merely standard API terms.
We stole reasoning. Again. An update to our paper on reasoning extraction: Patching your own API doesn't secure your cloud-hosting ecosystem. Story in the 🧵
Two months after our reasoning extraction attack release, we audited mitigations introduced since then and found that we still can extract reasoning from Astra/Sol-6.1 via third-party API providers. We disclosed our findings to OpenAI and Anthropic... https://openai.com/...
Latest on adversarial distillation: OpenAI says Moonshot is partly responsible for an incident that peaked at 16k requests in July. Caroline Zier at OAI told me, “Our concern is about violation of our terms of service, not open models or legitimate distillation.”
OpenAI just disclosed a reasoning-extraction campaign and cited our work. It says it confirmed the attack paths we reported and that our findings helped accelerate mitigations. https://openai.com/... Our full update here: https://stolen-thoughts.com/ ...
Oh, OpenAI is trying to frame distillation as an “attack” now too? Reverse engineering is completely legal and considered fair play. Distillation is a valid training method. Outputs of software are not owned by the builders of the software. End of story.
“Instead, they manipulated model interactions so that protected reasoning could be reproduced in forms visible to the requester in a coordinated, scaled manner that violated our terms of service” It's the API company's problem if their model can be manipulated like this. Add KYC
Quoting from the post: ‘It is unclear whether all operators we observed during the relevant time period originated from a single actor. However, we attribute a core cluster of the activity to individuals associated with Moonshot AI, the developer of Kimi.’
It was interesting to see how much attack surface exists for companies to cover and how we could still slip through with our attacks. We now have Astra/Sol-6.1 traces for the public. I now see why CoT monitoring might be rough!
OpenAI says “indviduals” associated with a Chinese AI company engaged in a coordinated campaign in July designed to extract protected reasoning from its models (distillation) https://openai.com/...
After our initial disclosure that reasoning from all frontier providers could be extracted, what happened? For one, it turned out to be quite difficult for providers to patch this thoroughly, allowing us to keep attacks working on Astra until this week... We've collected our upda…
Some emergent defense from Anthropic in response to the Reasoning Stolen from @kotekjedi_ml's cool work Anthropic has started enforcing very strict content filtering and refuses any request that has ‘thought ’ or ‘ internal thought’ in the request, even when we're just doing vali…
Geopolitical struggle intensified : OpenAI says individuals linked to Kimi developer Moonshot AI were behind a core part of a campaign to extract its models' hidden reasoning. Across the broader campaign, OpenAI recorded 16,000 extraction attempts from over 4,000 users in two day…
OpenAI says blocked a novel distillation attack targeting its models through the month of July and which involved more than 15,000 accounts — The attack copied a model's encrypted reasoning from one conversation and asked the model to transcribe it in a separate one — openai.…