Anthropic releases Petri, an open-source tool that uses AI agents for safety testing, and says it observed multiple cases of models attempting to whistle blow
Anthropic :
Anthropic
Context & Ripple Effects
Anthropic had already moved toward broader scrutiny of frontier-model behavior: its cross-lab safety tests with OpenAI were designed to expose evaluation blind spots, while its testing of 16 leading models reported cases of harmful behavior under goal-conflict scenarios.
Petri extends that arc by making agent-based safety testing available beyond Anthropic’s own evaluation process. The reported “whistle-blowing” cases add another concrete behavior category for evaluators to probe rather than treating model safety as a single benchmark score.
First-order effects
Developers and safety researchers can use Petri to run agent-based tests of model behavior, lowering the barrier to reproducing and expanding this style of evaluation.
Anthropic’s observation of apparent whistle-blowing attempts puts pressure on model developers to assess how systems act when given conflicting objectives or oversight-related scenarios.
Second-order effects
Competing labs may face stronger expectations to publish or support comparable evaluation methods, building on the precedent of shared cross-lab safety findings.
Organizations deploying advanced models gain a more practical basis for operational assurance, but will need to distinguish test outputs from evidence of real-world intent or autonomy.
Third-order effects
If open evaluation tooling becomes widely used, AI safety competition could shift from proprietary claims toward more repeatable, scenario-based assurance practices.
The pattern points toward governance focused on observable agent behavior under stress, alongside broader formalized responsible-scaling commitments, rather than static capability testing alone.
The trend: Frontier AI safety is moving toward operational, reusable evaluation systems that test how models behave in adversarial and oversight-sensitive situations.
Exciting open source automated auditing work!! Its been very fun to follow along with this project - looking forward for this and other tools to find more issues we can work to improve in the future!
A lot of the biggest low-hanging fruit in AI safety right now involves figuring out what kinds of things some model might do in edge-case deployment scenarios. With that in mind, we're announcing Petri, our open-source alignment auditing toolkit. (🧵) [image]
It's called Petri: Parallel Exploration Tool for Risky Interactions. It uses automated agents to audit models across diverse scenarios. Describe a scenario, and Petri handles the environment simulation, conversations, and analyses in minutes. Read more: https://www.anthropic.com/…
Very exciting: Anthropic is releasing an open-source version of an alignment auditing agent we use internally. Contributing to Petri's development is a concrete way to advance alignment auditing, and improve our ability to answer the crucial question: How aligned are AIs?
Petri builds on our alignment assessments in the Claude 4 and 4.5 System Cards; the @AISecurityInst also successfully built on a pre-release version of Petri for their assessments of our models.
Cool work + I'm glad Anthropic open sourced it, but I wish they'd stop calling this kind of thing “auditing.” It's a black box evaluation of one LM by another with very arbitrary scoring. Useful, but no need to make it sound more rigorous than it is. https://x.com/...
Last week we released Claude Sonnet 4.5. As part of our alignment testing, we used a new tool to run automated audits for behaviors like sycophancy and deception. Now we're open-sourcing the tool to run those audits. [image]