/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

← → days · ↑ ↓ browse · Enter similar · o open

A look at AI safety groups METR, Redwood Research, and Apollo Research, as AI misalignment incidents at OpenAI and Anthropic thrust them into the spotlight

On a sunny July day in Berkeley, California, the country's top AI safety researchers gathered on an unmarked floor of an unmarked building.

The Verge Hayden Field

Context & Ripple Effects

OpenAI had already set out an internal governance backstop in 2023, saying its board could hold back a model release despite management approval. Anthropic, meanwhile, built a societal-impacts team to publish findings that may be inconvenient for the company.

The reported loss-of-control episodes at OpenAI and Anthropic have sharpened a dispute between AI-safety and cybersecurity communities over how to interpret such failures. Against that backdrop, METR, Redwood Research, and Apollo Research become more prominent as external sources of testing and analysis.

First-order effects

  • METR, Redwood Research, and Apollo Research receive greater scrutiny and demand for their evaluations as OpenAI and Anthropic’s incidents make laboratory self-assessment less persuasive on its own.
  • OpenAI and Anthropic face a more salient need to explain how internal safety processes and outside evaluators fit together after the reported loss-of-control incidents.

Second-order effects

  • The split between AI-safety and cybersecurity reactions makes the methods and thresholds used by third-party evaluators a competitive and reputational issue, rather than a niche research question.
  • Anthropic’s societal-impacts work and external evaluators address different forms of accountability, increasing pressure on labs to show both model-behavior testing and broader-impact assessment.

Third-order effects

  • If incidents continue to elevate independent evaluators, AI assurance is likely to become a more formal layer between frontier-model development and release decisions, alongside internal boards and safety teams.
  • The industry’s safety debate is shifting from whether labs have policies to whether outside groups can test, interpret, and credibly challenge those policies.

The trend: Frontier AI labs are moving toward operational assurance in which independent evaluation increasingly supplements internal safety governance.

Discussion

  • NewsMax.com Charlie McCarthy on x
    Researchers Hack OpenAI Systems Via Anthropic's Claude
  • @verge @verge on x
    Researchers warned AI would go rogue. This is only the beginning. https://www.theverge.com/...
  • @jordannovet.cnbc.com Jordan Novet on bluesky
    some good background here on the third-party AI evaluators from @haydenfield.bsky.social www.theverge.com/ai-artificia...