In a paper, AI researchers from OpenAI, Google DeepMind, Anthropic, and others recommend “further research into chain-of-thought monitorability” for AI safety
AI researchers from OpenAI, Google DeepMind, Anthropic, and a broad coalition of companies and nonprofit groups …
TechCrunch Maxwell Zeff
Related Coverage
- Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety Tomek Korbak
- Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety arXiv.org
- OpenAI, Google DeepMind and Anthropic sound alarm: ‘We may be losing the ability to understand AI’ VentureBeat · Michael Nuñez
- What if we could catch AI misbehaving before it acts? Chain of Thought monitoring explained Digit · Mithun Mohandas
- Read AIs' Thoughts While We Still Can, Warn Researchers from 20 AI Labs Nikita Ostrovsky
- Tracking AI models' ‘thoughts’ could reveal how they make decisions, researchers say The Indian Express
- Top AI Labs Sound Joint Alarm: Transparency in Machine Reasoning May Soon Vanish implicator.ai · Robert Brown
Discussion
-
@timkellogg.me
Tim Kellogg
on bluesky
OpenAI backs a new cross-organizational paper emphasizing the promise of Chain of Thought (CoT) monitoring — agents are generally black boxes, but the one way we've identified to monitor them is their CoT. trouble is it's currently brittle and needs more research — arxiv.org/…
-
@bobabowen
Bowen Baker
on x
Modern reasoning models think in plain English. Monitoring their thoughts could be a powerful, yet fragile, tool for overseeing future AI systems. I and researchers across many organizations think we should work to evaluate, preserve, and even improve CoT monitorability. [image]
-
@merettm
Jakub Pachocki
on x
I am extremely excited about the potential of chain-of-thought faithfulness & interpretability. It has significantly influenced the design of our reasoning models, starting with o1-preview. As AI systems spend more compute working e.g. on long term research problems, it is
-
@bobabowen
Bowen Baker
on x
CoT monitoring is already useful! In a recent OpenAI blog and paper, we showed that we could catch reward hacks in code via CoT monitoring. Others have found they can catch early signals of misalignment, prompt injections, and evaluation awareness. https://openai.com/...
-
@garymarcus
Gary Marcus
on x
truly fascinating win for neurosymbolic AI, raising interesting questions about the evolution of human cognition. long chains of cognition must be translated into words [symbols!] - and not just transit through points in embedding space. incredibly interesting observation.
-
@sujay_kapadnis
Sujay
on x
Cross-institutional paper endorsed by leaders from OpenAI, Anthropic, and academia is urging labs to preserve the monitorability of AI reasoning. https://tomekkorbak.com/... [image]
-
@owainevans_uk
Owain Evans
on x
New position paper on Chain of Thought monitoring that I'm excited to be a (small) part of. This is related to our recent work showing that emergently misaligned models sometimes articulate their misaligned plans in their CoT. https://x.com/...
-
@balesni
Mikita Balesni
on x
A simple AGI safety technique: AI's thoughts are in plain English, just read them We know it works, with OK (not perfect) transparency! The risk is fragility: RL training, new architectures, etc threaten transparency Experts from many orgs agree we should try to preserve it: [ima…
-
@gdb
Greg Brockman
on x
We've published a position paper, with many across the industry, calling for work on chain-of-thought faithfulness. This is an opportunity to train models to be interpretable. We're investing in this area at OpenAI, and this perspective is reflected in our products:
-
@shakeelhashim
Shakeel
on x
Very notable that this paper has authors from *every* major AI company (OpenAI, Anthropic, GDM, Meta). Also endorsed by Ilya Sutskever.
-
@openai
@openai
on x
Chain of Thought (CoT) monitoring could be a powerful tool for overseeing future AI systems—especially as they become more agentic. That's why we're backing a new research paper from a cross-institutional team of researchers pushing this work forward.
-
@morqon
Morgan
on x
quite the author list, great to see top-flight researchers working together across industry on this, no arena when it matters
-
@balesni
Mikita Balesni
on x
Our recommendations for AI developers: 1️⃣ Develop standardized monitorability evaluations 2️⃣ Report results in model system cards 3️⃣ Factor monitorability into training/deployment decisions 4️⃣ Consider architectural choices that preserve transparency
-
@bobabowen
Bowen Baker
on x
For hard enough tasks, models may have to reason out loud and be monitorable. If actions that cause severe harm require complex reasoning, then this gives hope we could catch and stop them with CoT monitoring. [image]
-
@neelnanda5
Neel Nanda
on x
It was great to be part of this statement. I wholeheartedly agree. It is a wild lucky coincidence that models often express dangerous intentions aloud, and it would be foolish to waste this opportunity. It is crucial to keep chain of thought monitorable as long as possible
-
@balesni
Mikita Balesni
on x
Seven years ago, we expected AIs to be opaque RL agents. The current transparency, while imperfect, is a great gift! We should try to preserve it and leverage it for safety (among other oversight techniques).
-
@rohinmshah
Rohin Shah
on x
Chain of thought monitoring looks valuable enough that we've put it in our Frontier Safety Framework to address deceptive alignment. This paper is a good explanation of why we're optimistic - but also why it may be fragile, and what to do to preserve it. https://x.com/...
-
@woj_zaremba
Wojciech Zaremba
on x
When models start reasoning step-by-step, we suddenly get a huge safety gift: a window into their thought process. We could easily lose this if we're not careful. We're publishing a paper urging frontier labs: please don't train away this monitorability. Authored and endorsed [im…
-
@tomekkorbak
Tomek Korbak
on x
The holy grail of AI safety has always been interpretability. But what if reasoning models just handed it to us in a stroke of serendipity? In our new paper, we argue that the AI community should turn this serendipity into a systematic AI safety agenda!🛡️