Source: OpenAI's Astra model uses “recurrent depth”, a technique that improves cost and performance but obscures the AI's reasoning, making it harder to monitor
The Information
Context & Ripple Effects
OpenAI had already shown the Astra family to US policymakers and regulators, emphasizing its ability to handle long-running tasks. The company also said Astra had reached its “Critical” cyber-capability threshold, placing the model’s safeguards and evaluability at the center of its rollout case.
The Information’s source-based report adds a trade-off to that arc: the same reported architecture that improves cost and performance may leave less reasoning exposed in forms monitors can inspect. Public reaction was divided, with some safety-focused commenters treating the report as a major monitoring concern and others disputing its characterization.
First-order effects
- If the reported recurrent-depth design is used in Astra, OpenAI’s safety teams face a harder task of detecting problematic reasoning through natural-language traces while evaluating a model already subject to stronger cyber safeguards.
- Regulators briefed on Astra’s long-running-task capabilities have a more difficult assurance question: whether observed outputs and safeguards are sufficient when internal reasoning is less legible.
Second-order effects
- OpenAI’s deployment controls would need to carry more of the safety burden if reasoning traces provide less useful evidence, increasing the importance of behavioral evaluations and access restrictions around advanced capabilities.
- The report raises the bar for customers and government counterparts seeking auditability from frontier models: performance gains and lower costs become less separable from the evidence available for oversight.
Third-order effects
- If recurrent architectures deliver competitive gains while reducing inspectable reasoning, frontier-model governance will shift further from reviewing chains of thought toward testing behavior, controlling access, and validating safeguards.
- That would sharpen the private observability paradox: the most commercially valuable model designs may be the ones least compatible with straightforward external monitoring.
The trend: Frontier AI development is increasingly forcing a trade-off between cheaper, stronger models and the transparency that safety and regulatory oversight rely on.
Related: Private observability paradox · The state-compatible AI lab · OpenAI · OpenAI says Astra reached its Critical cyber threshold · OpenAI demoed Astra to policymakers and regulators
Related Coverage
- OpenAI's new reasoning technique alarms AI safety experts TechCrunch · Russell Brandom
- Path to Astra: critical capabilities and frontier safeguards OpenAI
- This Is the Worst Possible Time for OpenAI to BfЖ7!م#2.$9&क Gizmodo · Webb Wright
- Researchers fear safety disaster ahead of OpenAI's Astra release The Verge · Robert Hart
- Anthropic Has Some Alignment Problems Don't Worry About the Vase · Zvi Mowshowitz
- OpenAI says upcoming model is so capable it requires stronger guardrails Reuters · Deepa Seetharaman
- OpenAI Astra: All about the quantum math-solving model with ‘critical’ hacking skills Mashable · Timothy Beck Werth
- OpenAI says it plans to publicly release a version of Astra “soon” but will give access to “its most advanced cyber capabilities” only to testers and partners Wired
- OpenAI says Astra is its first model to reach its Critical cybersecurity threshold, and its safeguards may mistakenly flag legitimate activity as cyber misuse Axios · Ina Fried
- Path to Astra: critical capabilities and frontier safeguards Hacker News
- AI & Tech Brief: Fable 5.1 and data privacy Washington Post · Benjamin Guggenheim
- OpenAI's Astra model is on the way — and very good at breaking into computer systems TechCrunch · Tim Fernholz
- ICYMI: OpenAI's Astra crosses Critical cybersecurity threshold TestingCatalog AI News
Discussion
-
@merettm
Jakub Pachocki
on x
I want to prevent a race into unmonitorability kicked off by confused reporting. The depth of the computation graph for our present frontier models, including Astra, is within a factor of two of GPT-4. OpenAI has worked to preserve and utilize chain-of-thought monitoring since ou…
-
@amir
Amir Efrati
on x
new: OpenAI & others quietly using loop transformers that don't show their ‘thinking’ when scaled up a leap forward on performance, but sparking concerns inside & outside OpenAI re: security as this takes off
-
@_nathancalvin
Nathan Calvin
on x
Really huge and extremely concerning story from the Information tonight. Looks like OpenAI utilized a breakthrough in neuralese for Astra that could destroy chain of thought monitorability - though the Informations source told them that OpenAI is currently
-
@deanwball
Dean W. Ball
on x
this latest panic over the false claim that OpenAI is “doing neuralese
-
@ryangreenblatt
Ryan Greenblatt
on x
OpenAI's newest AI, Astra, is reported to use an ‘opaque reasoning’ architecture where more of the reasoning occurs in activations instead of natural language. This may be the single worst development for AI security/safety to date. The details of Astra aren't publicly known, but…
-
@tszzl
Roon
on x
i agree with this and think that CoT is at best an epiphenomenon of current training methods. it will break (in the future, not now). safety community too often centers thought terminating taboos like “neuralese” and “training on interp”
-
@thlarsen
Thomas Larsen
on x
Very bad if true. I previously thought that in a short timelines world, the most likely case was that (1) the AIs would be misaligned, but (2) we would get a lot of evidence about it from reading the COTs. This evidence would increase the chance of a reasonable response from labs…
-
@mobav0
Mo Bavarian
on x
It's remarkable that The Information gets to mix 10% truth with 90% untruth and sell it to the public as “Information
-
@scaling01
@scaling01
on x
just ~240 layers of effective depth what could possibly go wrong? this is genuinely the single worst thing that OpenAI ever did. it's like they committed a war crime. they opened Pandora's box. remember what happened when they announced strawberry and a new reasoning approach? ev…
-
@jeremiecharris
Jeremie Harris
on x
I don't understand how you can look at the Hugging Face incident and go “we should make this happen again except next time lets allow the AI swarm bots to communicate in a secret language humans can't understand” It would make for less shocking METR post-mortems at least.
-
@dkokotajlo
Daniel Kokotajlo
on x
Holy shit fuck
-
@mattyglesias
Matthew Yglesias
on x
More powerful but also more opaque, what could go wrong?
-
@austinc3301
Agus
on x
This is extremely bad. I would have thought that the almost universal consensus of this being a terrible idea would dissuade OAI from it, but alas, you can never quite trust them with anything. This is horrible news for safety and for society as a whole.
-
@amir
Amir Efrati
on x
There seems to be a misunderstanding about the piece and what OpenAI is up to. There are new techniques at frontier labs that involve loop transformers and similar. As we say in the piece, OpenAI is putting limits on the loops and trying to make sure CoT is visible with Astra. Th…
-
@industriaalist
@industriaalist
on x
this is such an unwinnable fight and i hope people are not naive enough to bank on labs agreeing to not make their models deeper (either with more layers or recurrence), esp when computational depth will become the bottleneck to scaling soon. for some context, sparsity and comput…
-
@eliebakouch
Elie
on x
a few thoughts on “recurrent depth” transformers the main question: why recurrent depth instead of just scaling depth? it's not faster at inference or training* since you still go through the full “effective depth
-
@kimmonismus
@kimmonismus
on x
Holy, Astra is build different: OpenAI's Astra model reportedly uses “recurrent depth
-
@tenobrus
@tenobrus
on x
glad to see this post, and glad to see the situational awareness it represents from OAI. still quite concerned about the direction.
-
@andrewginns
Andrew Ginns
on x
We can still monitor the CoT. Reporting on it being opaque is misinformed
-
@ryangreenblatt
Ryan Greenblatt
on x
Transparency about the opaque serial depth is great, but this statement is consistent with Astra having a configurable “dial
-
@teortaxestex
@teortaxestex
on x
> The depth of the computation graph for our present frontier models, including Astra, is within a factor of two of GPT-4. so <240 layers? Makes sense, hardware punishes anything deeper, and training becomes harder too... But... been a while since we really Stacked More Layers.
-
@zeffmax
Max Zeff
on x
OpenAI's chief scientist responds to The Information's report that the company is using a looped transformer, an architecture that could degrade COT monitorability—one of the few windows into how AI models work. Broadly tho, he says CoT monitorability is fragile and
-
@tomekkorbak
Tomek Korbak
on x
i think the day when a frontier lab trains a frontier-scale recurrent (or otherwise unmonitorable) language model would be one of the darkest in the current AI era. this day is not today and i would love frontier labs to coordinate on a commitment that it never comes.
-
@eliebakouch
Elie
on x
note that the latest big model by oai with architecture details was already very deep by today's standards, here is the number of layers for a few models Llama-3.1 405B: 126 GPT3: 96 Kimi K3: 93 Qwen 3.8 max: 92 GLM-5.3: 78 DeepSeek-V4-Pro: 61 gpt-oss 120b: 36 gpt-oss 20b: 24
-
@jachiam0
Joshua Achiam
on x
In the interest of people understanding OpenAI's position, everyone should read this message from the Chief Scientist:
-
@miles_brundage
Miles Brundage
on x
I'm not in the weeds enough to have a view on how much the 2x GPT-4 thing clarifies/reassures, but glad OAI quickly issued a statement. Would love to move towards continuous embedded auditing of things like this rather than reacting to incidents + leaks https://x.com/...
-
@alextmallen
Alex Mallen
on x
Important clarification on the opaque(?) serial depth. Though I still have some important uncertainty about what specific architecture is being used.
-
@balesni
Mikita Balesni
on x
switching to fully recurrent LLM architectures would be the biggest blow to safety, probably in history of AI all AI labs should commit to limit the opaque serial depth of their models, for the foreseeable future. this will not ensure monitorable CoTs but will protect us from the…
-
@j_asminewang
Jasmine Wang
on x
we should avoid a race into unmonitorability & have a multilab commitment/standard to avoid neuralese
-
@peterwildeford
Peter Wildeford
on x
I hope to see OpenAI continue their commitment to monitorability!
-
@micahcarroll
Micah Carroll
on x
A race to the bottom in monitorability due to a false belief that OpenAI is using neuralese models would be incredibly stupid
-
@davidmanheim
David Manheim
on x
How are things going now? OpenAI:
-
@mkinniment
Megan Kinniment
on x
Broadly agree with Ryan's takes here. FWIW, seems like OpenAI has been experimenting with monitorability definitions that permit recurrence for a while. This is from March 2025:
-
@ryangreenblatt
Ryan Greenblatt
on x
It's possible OpenAI and others will use these architectures with only a few recurrent iterations (because that's most performant or due to safety concerns). If so, this would only make things moderately worse. But I expect much more opaque reasoning than this in the future...
-
@jachiam0
Joshua Achiam
on x
A very hot take: chain of thought interpretability was always going to be so fragile as to be an unacceptable backstop for long-term AI safety, and while I admire the optimism and effort involved in protecting its fidelity (and consider such effort to have been worthwhile), I do …
-
@ryangreenblatt
Ryan Greenblatt
on x
> This may be the single worst development for AI security/safety to date. By this “development
-
@amir
Amir Efrati
on x
Astra's chain of thought can be monitored, y'all. OpenAI said it today, as we wrote in our piece. This is about the future and what happens as new techniques proliferate and get supercharged.
-
@0xdoug
Doug Colkitt
on x
AI safety hysteria is where level headed reasoning goes to die. “Recurrent depth” is just a relatively minor architectural change, where the transformer spends more time on harder tokens. No, it's not going to replace chain of thought with “neuralese
-
@emostaque
Emad
on x
I see a lot of my TL freaking out about this Here is the reality: Frontier models will one shot just about anything at 10,000 tokens per second in a few years Do you really think we can monitor that? Would need speed limits on the model that it would code around anyway
-
@aaronscher
Aaron Scher
on x
Big shame on OpenAI for seemingly breaking the norms they agreed to 10 months ago and doing one of the most dangerous AI methods we know of.
-
@imjustnewatai
@imjustnewatai
on x
looped transformers are real, public, and already here. the information says astra puts them inside openai's first critical-cyber model. a standard transformer sends each token through a fixed stack. recurrent depth reuses some blocks, feeding the hidden state through them again …
-
@alexiglad
Alexi Gladstone
on x
feed-forward transformers -> recurrent transformers -> energy-based transformers is the natural progression
-
@daniel_mac8
Dan McAteer
on x
Astra uses “recurrent depth” reasoning, a more powerful but less legible form of reasoning. Fable 5.1 explains:
-
@garrisonlovely
Garrison Lovely
on x
OpenAI: At long last, we have created The Neuralese Model from the classic OpenAI blogpost from 10 months ago, Don't Create The Neuralese Model.
-
@zephyr_z9
@zephyr_z9
on x
So those looped transformer rumors that used to float every 6 months turned out to be true????
-
@teortaxestex
@teortaxestex
on x
Did OAI finally make the looped meme work?
-
@jessesingal
Jesse Singal
on x
we're begging to get our asses skynetted
-
@steph_palazzolo
Stephanie Palazzolo
on x
OpenAI's Astra AI uses a new reasoning approach called “recurrent depth.
-
@so8res
Nate Soares
on x
They sure are rushing to build the “it” from the New York Times Bestseller “If Anyone Builds It, Everyone Dies”
-
@pigeon__s
@pigeon__s
on x
wait Astra is a latent space thinking model? i dont really buy that tbh but if thats true thats fucking awesome ive been wanting a big model to use this for years
-
@scaling01
@scaling01
on x
holy shit they did it Astra uses “recurrent depth” start the countdown for open models to use that too
-
@dimitrispapail
Dimitris Papailiopoulos
on x
LoOpEd TrANsfOrMeRs :D
-
@scaling01
@scaling01
on x
there goes CoT monitoring
-
@_nathancalvin
Nathan Calvin
on x
This may seem niche but it is a huge deal and genuinely scary. A remarkable list of top AI researchers, including at OpenAI, previously described monitorable chain of thought as “A New and Fragile Opportunity for AI safety.
-
@chrisgpt
Chris
on x
WAIT PLEASE TELL ME OPENAI PARALLELIZED RECURRENT DEPTH. Microsoft Research already showed us in June that you can run K latent blocks in parallel for R recurrent iterations. LoopCoder has also scaled recurrent depth to 40B parameters in July + a few others. But NOBODY has public…
-
@robertskmiles
Rob Miles
on x
Can you believe these dipshits? Here's OpenAI's blog (being correct) less than a year ago:
-
@andrewcurran_
Andrew Curran
on x
So this is how they did it.
-
@sjgadler
Steven Adler
on x
If this is true, OpenAI seems to be violating one of the few redlines that exist in the AI industry. Absolutely do not train your models like this - what is going on??
-
Kanean S A
Kanean S A
on linkedin
we previewed how we evaluated Astra (https://lnkd.in/g7xiBYap) — 2 things stand out to me: — 1. Astra is significantly more capable, and …
-
r/technology
r
on reddit
OpenAI Is About to Release Its First AI Model With ‘Critical’ Cyber Abilities | The company will give select partners early access to its Astra AI model …
-
r/codex
r
on reddit
Breaking: OpenAI's new Astra model will reportedly think in neuralese
-
r/OpenAI
r
on reddit
Fable 5.1 outperforms Fable 5, Opus 5, and GPT-5.6 Sol. Isn't time to reveal GPT-6 Astra?
-
r/accelerate
r
on reddit
OpenAI-Path to Astra: critical capabilities and frontier safeguards
-
r/singularity
r
on reddit
Path to Astra: critical capabilities and frontier safeguards
-
r/OpenAI
r
on reddit
Soon as in tomorrow (Astra)?