Former OpenAI policy researcher Miles Brundage criticizes an OpenAI post on safety and alignment, saying it “rewrites the history of GPT-2 in a concerning way”
A high-profile ex-OpenAI policy researcher, Miles Brundage, took to social media on Wednesday to criticize OpenAI for …
Context & Ripple Effects
OpenAI’s GPT-2 release was framed around misuse concerns, making its account of that episode consequential to how the lab’s safety record is evaluated. The dispute also follows earlier criticism over limited disclosure around GPT-4 and reporting that some safety staff felt launch pressure around GPT-4o.
OpenAI had previously brought in the Alignment Research Center to assess GPT-4 risks, including power-seeking behavior, through an external risk assessment. Brundage’s objection puts the credibility of such safety-facing communications under fresh scrutiny.
First-order effects
- Brundage’s public challenge creates an immediate credibility dispute over OpenAI’s characterization of its GPT-2 safety and alignment history.
- The criticism gives researchers, policymakers, and other outside observers a concrete reason to scrutinize how OpenAI documents past safety decisions.
Second-order effects
- Other frontier-model developers face stronger incentives to preserve clear records and evidence for safety claims, since retrospective narratives can become a governance liability.
- External evaluation becomes more salient as a counterweight to labs’ self-descriptions, building on OpenAI’s earlier use of an outside GPT-4 risk review.
Third-order effects
- If former insiders increasingly contest labs’ public safety narratives, assurance may shift from voluntary statements toward more independent, auditable processes.
- The episode illustrates a broader legitimacy test for frontier AI: safety governance must be credible not only at launch, but also in the historical record used to justify it.
The trend: Frontier AI governance is moving from broad safety commitments toward demands for independently verifiable evidence, records, and oversight.