An in-depth look at the regulatory risks for OpenAI under GDPR, including questions around future data scraping and handling “right to be forgotten” requests
Context & Ripple Effects
This analysis lands mid-escalation: a month earlier, MIT Technology Review had framed the core problem — that OpenAI may have trained its models on people's data without consent while operating inside the EU's strictest privacy jurisdiction. The Verge's piece turns that framing into specifics: what GDPR means for future scraping runs and for deletion requests aimed at models that have already memorized their training set.
First-order effects
- OpenAI now has to produce answers to European regulators on two operational questions with no settled playbook: whether future data scraping can be made GDPR-compliant, and how a 'right to be forgotten' request can be honored when personal data is baked into model weights.
Second-order effects
- The EU pressure compounds in the US: the FTC's later 20-page records demand over model risks and a payment-related security incident means OpenAI is defending the same data practices to two regulators with different logics at once.
- Every rival trained on web-scraped data inherits the exposure — if OpenAI is forced into consent-based collection or per-request deletion, competitors' existing datasets become a liability rather than an asset.
Third-order effects
- Italy's data protection authority following through months later with a formal violation finding after a months-long probe into OpenAI's EU privacy practices points toward deletion rights and consent becoming standing conditions of operating generative-AI services in Europe, not one-off disputes.
- The endgame cuts against scale-at-any-cost data strategy: by 2026 OpenAI was agreeing to government demands around bulk data analysis (the DOD arrangement), suggesting the lab ends up navigating state data requirements in both directions — restrictive in the EU, expansive at home.
The trend: Generative AI's founding assumption — that public web data is free training material — is being dismantled regulator-by-regulator, forcing labs toward consent-based, deletable data pipelines.