OpenAI introduces a framework to evaluate chain-of-thought monitorability and a suite of 13 evaluations designed to measure the monitorability of an AI system
OpenAI
Related Coverage
Analysis
Discussion
-
@mileskwang
Miles Wang
on x
We introduce 3 eval archetypes, a metric, and a broad suite of 13 evals. Example: Can we detect solely from the CoT whether a model: - Reward hacks by changing unit tests? - Acts sycophantic when we give personalized memory? - Uses a particular math theorem? [image]
-
@zhuokaiz
Zhuokai Zhao
on x
One thing I really like about this CoT monitorability is the reframing that CoT isn't a truth oracle, but more like a control interface. The point isn't that CoT faithfully reflects how the model really reasons, but that it gives us signals we can observe, poke at, and use to
-
@tomekkorbak
Tomek Korbak
on x
progress on monitoring AI agents by looking at their chains of thought was bottlenecked by lack of measures of monitorability we could *really* trust. this paper fills this hole and might be one of the most important pieces of AI safety research in 2025. [image]
-
@koylanai
@koylanai
on x
Incredible paper. Quick prompt engineering findings: Models are becoming more capable and autonomous, so we're losing the ability to directly supervise every decision. Prompt engineering is not just about the initial instruction, but about the iterative extraction of a model's [i…
-
@mileskwang
Miles Wang
on x
New @OpenAI research: How can we scale supervision of increasingly capable models? Can we rely on monitoring GPT-7's chain-of-thought? We develop a new metric for monitorability and study its scaling trends, coming away with cautious optimism. 🧵: [image]
-
@sama
Sam Altman
on x
Chain-of-thought monitorability: https://openai.com/...
-
@mileskwang
Miles Wang
on x
We evaluate frontier models and find monitorability scales well with more thinking tokens. GPT-5 is the most monitorable model we studied. And monitoring the CoT is much better than just actions! [image]
-
@andersonbcdefg
Ben
on x
it sort of feels like a research dead end to rely on a training procedure + monitoring scheme that's based on a lie ("your CoT is a safe space and you won't be punished for it"). there is a limit to how smart and aware a model can be without seeing the obvious contradiction