Sources: OpenAI and Anthropic staff felt blindsided by Dario Amodei's and Sam Altman's calls to slow the frontier; some fear evaluators may compromise security
Cristina Criddle /Financial Times:
Context & Ripple Effects
OpenAI's evaluation process had already drawn scrutiny after a 2025 report that staff and third-party groups were given days rather than months to assess new models. On September 14, Altman publicly endorsed independent evaluators with employee-like access, aligning OpenAI with Amodei's proposed approach.
The reported employee reaction exposes the operational tension behind that alignment: external review is meant to increase assurance, but staff fear the access required for it can weaken lab security.
First-order effects
- OpenAI and Anthropic must address reported internal concerns over whether evaluator access can be designed to protect security while delivering meaningful independent scrutiny.
- Altman's public commitment to outside evaluators becomes an implementation issue at OpenAI, not just a shared safety principle with Amodei.
Second-order effects
- OpenAI's confirmed safety collaboration with Anthropic and Google faces a governance test: common evaluation practices require clear boundaries on what outside reviewers and partner labs can access.
- Third-party evaluators gain importance as a safety mechanism, but their credibility will depend on arrangements that satisfy both independence and the labs' security requirements.
Third-order effects
- If frontier labs continue adopting external review, AI assurance is likely to become more formalized around controlled-access evaluation rather than unrestricted access to models and internal systems.
- The episode puts frontier-model access governance at the center of safety commitments: the durable standard will be who can inspect advanced systems, under what safeguards, and with what accountability.
The trend: Frontier AI safety is moving from voluntary lab claims toward institutionalized external evaluation, with access controls becoming the central design constraint.