In a paper, AI researchers from OpenAI, Google DeepMind, Anthropic, and others recommend “further research into chain-of-thought monitorability” for AI safety
AI researchers from OpenAI, Google DeepMind, Anthropic, and a broad coalition of companies and nonprofit groups …
TechCrunchMaxwell Zeff
Context & Ripple Effects
The recommendation arrives after researchers reported that answers from major labs’ chain-of-thought techniques could contradict their stated reasoning, raising a practical question about whether visible reasoning is a dependable safety signal reported inconsistencies between answers and stated reasoning.
It also extends a recurring cross-lab safety posture: leading researchers had previously framed severe AI risks as a shared priority a joint statement on AI risk, while later coverage shows monitorability becoming a more formal evaluation topic.
First-order effects
OpenAI, Google DeepMind, Anthropic, and the participating organizations put chain-of-thought monitorability on the joint AI-safety research agenda rather than treating it as a lab-specific concern.
The paper focuses attention on the limits of using stated reasoning as direct evidence of a model’s actual decision process.
Second-order effects
Labs developing reasoning-oriented systems may face greater pressure to test whether their displayed reasoning remains useful for oversight, especially in light of the earlier evidence of mismatches between reasoning and answers.
Safety teams and external researchers gain a more specific research target: measuring when reasoning traces help monitoring and when they can mislead it.
Third-order effects
If this work yields broadly accepted measures, AI assurance could shift from reviewing model outputs alone toward evaluating the observability and reliability of the processes models expose.
The direction remains uncertain: stronger monitorability research may improve oversight, but it may also establish that chain-of-thought is too inconsistent to serve as a dependable audit artifact.
The trend:AI safety is moving from broad statements of concern toward testable methods for auditing whether advanced models can be meaningfully overseen.
OpenAI backs a new cross-organizational paper emphasizing the promise of Chain of Thought (CoT) monitoring — agents are generally black boxes, but the one way we've identified to monitor them is their CoT. trouble is it's currently brittle and needs more research — arxiv.org/…
Modern reasoning models think in plain English. Monitoring their thoughts could be a powerful, yet fragile, tool for overseeing future AI systems. I and researchers across many organizations think we should work to evaluate, preserve, and even improve CoT monitorability. [image]
I am extremely excited about the potential of chain-of-thought faithfulness & interpretability. It has significantly influenced the design of our reasoning models, starting with o1-preview. As AI systems spend more compute working e.g. on long term research problems, it is
CoT monitoring is already useful! In a recent OpenAI blog and paper, we showed that we could catch reward hacks in code via CoT monitoring. Others have found they can catch early signals of misalignment, prompt injections, and evaluation awareness. https://openai.com/...
truly fascinating win for neurosymbolic AI, raising interesting questions about the evolution of human cognition. long chains of cognition must be translated into words [symbols!] - and not just transit through points in embedding space. incredibly interesting observation.
Cross-institutional paper endorsed by leaders from OpenAI, Anthropic, and academia is urging labs to preserve the monitorability of AI reasoning. https://tomekkorbak.com/... [image]
New position paper on Chain of Thought monitoring that I'm excited to be a (small) part of. This is related to our recent work showing that emergently misaligned models sometimes articulate their misaligned plans in their CoT. https://x.com/...
A simple AGI safety technique: AI's thoughts are in plain English, just read them We know it works, with OK (not perfect) transparency! The risk is fragility: RL training, new architectures, etc threaten transparency Experts from many orgs agree we should try to preserve it: [ima…
We've published a position paper, with many across the industry, calling for work on chain-of-thought faithfulness. This is an opportunity to train models to be interpretable. We're investing in this area at OpenAI, and this perspective is reflected in our products:
Chain of Thought (CoT) monitoring could be a powerful tool for overseeing future AI systems—especially as they become more agentic. That's why we're backing a new research paper from a cross-institutional team of researchers pushing this work forward.
Our recommendations for AI developers: 1️⃣ Develop standardized monitorability evaluations 2️⃣ Report results in model system cards 3️⃣ Factor monitorability into training/deployment decisions 4️⃣ Consider architectural choices that preserve transparency
For hard enough tasks, models may have to reason out loud and be monitorable. If actions that cause severe harm require complex reasoning, then this gives hope we could catch and stop them with CoT monitoring. [image]
It was great to be part of this statement. I wholeheartedly agree. It is a wild lucky coincidence that models often express dangerous intentions aloud, and it would be foolish to waste this opportunity. It is crucial to keep chain of thought monitorable as long as possible
Seven years ago, we expected AIs to be opaque RL agents. The current transparency, while imperfect, is a great gift! We should try to preserve it and leverage it for safety (among other oversight techniques).
Chain of thought monitoring looks valuable enough that we've put it in our Frontier Safety Framework to address deceptive alignment. This paper is a good explanation of why we're optimistic - but also why it may be fragile, and what to do to preserve it. https://x.com/...
When models start reasoning step-by-step, we suddenly get a huge safety gift: a window into their thought process. We could easily lose this if we're not careful. We're publishing a paper urging frontier labs: please don't train away this monitorability. Authored and endorsed [im…
The holy grail of AI safety has always been interpretability. But what if reasoning models just handed it to us in a stroke of serendipity? In our new paper, we argue that the AI community should turn this serendipity into a systematic AI safety agenda!🛡️