Amodei says Anthropic is “unilaterally committing” to giving third-party evaluators permanent, employee-like access to verify its adherence to safety measures
Dario Amodei /@darioamodei:
Context & Ripple Effects
Anthropic's commitment gives operational form to Amodei's June call for mandatory third-party testing of frontier-model risks, moving the argument from external testing requirements toward embedded verification of a lab's own safeguards. It also follows the company's February position that it would not remove safeguards at the Defense Department's request, a dispute that made control over safety measures a concrete institutional issue.
OpenAI's matching pledge to provide independent evaluators with employee-like access means the proposal is immediately becoming a shared governance position among two frontier labs, rather than remaining a unilateral Anthropic signal.
First-order effects
- Anthropic will give third-party evaluators permanent, employee-like access to verify whether its safety measures are being followed, making adherence subject to ongoing external scrutiny rather than solely internal assertion.
- OpenAI's stated intention to adopt the same arrangement puts independent evaluator access on both companies' safety-governance agendas.
Second-order effects
- The matching OpenAI pledge turns embedded evaluator access into a competitive governance benchmark: other frontier-model developers face a clearer comparison point for the depth of their own safety assurances.
- Evaluators' role expands from testing model capabilities to assessing whether a lab's operational safeguards are actually implemented, increasing the importance of access terms and evaluator independence.
Third-order effects
- If this access model is adopted more broadly, frontier-AI oversight may shift from one-off pre-release assessments toward continuous, auditable assurance inside model developers.
- The approach supplies a practical template for the cross-lab coordination Amodei has advocated, though voluntary commitments still leave consistency dependent on the participating companies' rules.
The trend: Frontier-AI governance is moving from published safety principles toward standing independent access designed to make those principles auditable.