A look at Anthropic's Frontier Red Team, which has grown to 11 people overseen by Rhodes scholar Logan Graham to evaluate catastrophic risks in its AI models
Sam Schechner / Wall Street Journal : X: @jackclarksf X: Jack Clark / @jackclarksf : Fun story about the Frontier Red Team at Anthropic. I expect coming up with better and more realistic threat models for frontier risks is going to be one of the more important areas of AI policy to work on in 2025.
Context & Ripple Effects
Anthropic’s safety-first culture was already central to accounts of the company’s decision-making, while the Frontier Model Forum’s creation showed major labs trying to frame safety as a shared frontier-model responsibility. The red team makes that posture an identifiable internal operating function rather than a general principle.
Later coverage of an Anthropic Institute combining red-team and societal-impact work suggests this small unit was part of a broader move to organize technical risk testing, social-impact research, and policy thinking under more durable structures.
First-order effects
- Anthropic has an 11-person team, led by Logan Graham, dedicated to probing catastrophic-risk scenarios in its models, concentrating responsibility for that evaluation inside the lab.
- The team’s findings can give Anthropic a more formal internal input into decisions about frontier-model risks, alongside its broader safety-oriented operating approach.
Second-order effects
- A visible specialist team raises the bar for how peer frontier labs demonstrate that they test severe misuse and capability risks, reinforcing the safety agenda behind the industry’s Frontier Model Forum.
- As internal testing becomes more legible, policymakers and enterprise users gain a clearer organizational counterpart for requests about model-risk evidence and safety processes.
Third-order effects
- If frontier labs continue building distinct evaluation, societal-impact, and policy teams, AI safety is likely to become a standing institutional function rather than an ad hoc research activity—later reflected in Anthropic’s combined institute structure.
- That institutionalization could make internal testing practices more influential in future access-governance and oversight debates, though the corpus does not establish common external standards or independent review requirements.
The trend: Frontier AI labs are institutionalizing catastrophic-risk evaluation as models and the governance expectations around them become more consequential.