David Sacks says there's “substantial evidence” that DeepSeek “distilled knowledge out of OpenAI models and I don't think OpenAI is very happy about this”
White House artificial intelligence czar David Sacks said there's “substantial evidence” …
BloombergJackie Davalos
Context & Ripple Effects
The allegation arrived alongside OpenAI’s own report that it had seen evidence of DeepSeek using model-output distillation, turning a technical training dispute into a public policy issue involving the White House.
It also sits against OpenAI’s acknowledgment that DeepSeek had narrowed its lead and that OpenAI had been on the wrong side of history on open sourcing. The dispute therefore reaches beyond one competitor’s conduct to the terms on which model capabilities can spread.
First-order effects
Sacks’s intervention amplifies OpenAI’s complaint from a company-level allegation into a matter likely to draw greater attention from US AI policymakers.
DeepSeek faces heightened scrutiny over how its models were trained, while OpenAI has a stronger public basis to defend limits on access to its proprietary model outputs.
Second-order effects
Leading model providers may tighten monitoring and usage controls intended to detect automated extraction of outputs for training competing systems.
The dispute gives policymakers a concrete example for voluntary AI standards and release-timeline talks, particularly where proprietary models and open-source competitors intersect.
Third-order effects
If such allegations become a recurring policy focus, access to frontier-model outputs could increasingly be treated as a strategic control point rather than solely a product feature.
The longer-term fault line is whether AI competition is governed mainly through provider terms and technical safeguards or through state-backed rules around model access and cross-border capability transfer.
The trend: AI model access is becoming a geopolitical and regulatory issue as providers and governments contest whether model outputs can be used to build competing systems.
Isn't the entire fucking point of AI to distill what others have created? Sounds like they did it better than any of your companies, David. [embedded post]
they're gonna spend a few weeks throwing excuses at the wall to explain why they were out-innovated by china (again) and ultimately land on xenophobic fear mongering (again)
What I've read so far, the “evidence” is that they can't figure out how else DeepSeek could've done it, which, hrm. — Also, *if* someone stole OpenAI's data, call the police, using copyrighted material without paying its creators is a crime, good point. — techcrunch.com/2025/…
Oh, wow. Apparently using other peoples' IP to train an AI model is theft now, is it? Weird how that changed so very quickly 🤔 — techcrunch.com/2025/01/28/d...
The irony of a US-owned intellectual property stealing technology accusing Chinese-owned intellection property stealing technology of stealing American IP. You could not make this stuff up. — https://www.reuters.com/...
Asked if China's DeepSeek stole American IP, AI Czar David Sacks says it looks like a technique called distillation was used where a student model can “suck the knowledge” out of the parent model and there is evidence that DeepSeek distilled knowledge from OpenAI's models, which …
after seeing finance people attempt to explain evil Chinese concepts like open source, distillation and GPU training, i no longer think the average person can adapt to AI
noooo deepseek is a psyop nooo muh tik tok muh ccp mu- *company trains an AI to refuse your requests when you want smut* *company makes it lock down and say “policy, you are a bad person"* *ceo goes to biden and then lobbies to make it illegal to train ai so he can maintain a
What does DeepSeek R1 & v3 mean for LLM data? Contrary to some lazy takes I've seen, DeepSeek R1 was trained on a shit ton of human-generated data - in fact, the DeepSeek models are setting records for the disclosed amount of post-training data for open-source models...
There's a lot of misconception that China “just cloned” the outputs of openai. This is far from true and reflects incomplete understanding of how these models are trained in the first place. DeepSeek R1 has figured out RL finetuning...