Amodei says Anthropic is “unilaterally committing” to giving third-party evaluators permanent, employee-like access to verify its adherence to safety measures
We Must Pace the Frontier: I've written a new essay on why the AI industry should slow down, with a three-part plan for doing so. Anthropic is unilaterally committing to the first of these steps. We'll provide third-party evaluators with permanent, employee-level access to our systems, so that the...
Context & Ripple Effects
Anthropic had already argued in June that frontier models should undergo mandatory third-party testing for cyber, bio, and autonomy risks. Its new embedded-evaluator commitment turns that case for outside testing into an operating practice rather than solely a policy proposal.
The move is the first component of Amodei's broader plan to pace frontier development, which frames evaluation time as part of deploying capable systems safely. OpenAI's stated intention to adopt employee-like evaluator access gives the approach immediate relevance beyond Anthropic.
First-order effects
- Anthropic will have to give external evaluators durable access comparable to employees' access so they can assess whether its safety measures are being followed.
- Independent evaluators gain a standing verification role at Anthropic, rather than being limited to one-off assessments.
Second-order effects
- OpenAI's agreement to pursue similar access puts pressure on frontier-model developers to explain whether outside testing can inspect real internal safeguards rather than only supplied outputs.
- Evaluation organizations become more central counterparties for labs, with their ability to handle sensitive system access shaping which assurance arrangements are credible.
Third-order effects
- If other labs adopt the model, frontier AI governance may shift from voluntary safety statements toward continuous, independently inspectable assurance processes.
- The approach would make access governance a central industry question: meaningful oversight depends on evaluators receiving enough system visibility to test claims while labs retain control of sensitive capabilities.
The trend: Frontier AI labs are moving toward institutionalized external assurance, making embedded independent evaluation part of the governance infrastructure around advanced models.