OpenAI says it has expanded safety testing around its upcoming model Astra as it “cannot rule out” critical cyber capabilities, potentially delaying its launch
OpenAI “cannot rule out” that its upcoming model Astra has"critical" cyber capabilities, a designation that has prompted …
Axios
Context & Ripple Effects
OpenAI showed the Astra model family to US policymakers and regulators while emphasizing its ability to complete long-running tasks. The new testing program puts a safety gate behind that regulatory-facing Astra demonstration.
The move also follows reporting that OpenAI was preparing an advanced cybersecurity product for only a small group of partners, making Astra’s possible critical-cyber designation part of a broader restricted-access approach to cyber-capable AI.
First-order effects
- OpenAI must expand Astra’s safety evaluation before release, and its launch timetable is now contingent on the outcome of that work.
- Prospective Astra users face a less certain rollout schedule as OpenAI assesses whether the model meets its critical-cyber threshold.
Second-order effects
- US policymakers and regulators who were shown Astra gain a concrete safety constraint to weigh alongside the model family’s long-running-task capabilities.
- A critical-cyber finding would reinforce limited-access deployment for OpenAI’s most cyber-capable systems, rather than a broadly available launch.
Third-order effects
- If critical-capability classifications repeatedly delay frontier releases, evaluation capacity becomes a product-shipping constraint rather than a final prelaunch check.
- Frontier-model competition would increasingly turn on access controls and assurance processes, with regulator engagement occurring before public deployment.
The trend: Frontier AI labs are moving toward government-aware, risk-gated deployment as models approach capabilities that may require restricted access.
Related: Frontier-model access governance · Operational AI assurance · OpenAI · OpenAI demos Astra to US policymakers · OpenAI’s restricted cyber-capability product
Related Coverage
- Responding to the next frontier of critical cyber capabilities OpenAI
- OpenAI Pauses Some Work on New AI Model Over Cybersecurity Concerns Wall Street Journal · Tina Li
- OpenAI Delays Next Major AI Model ‘Astra’ Over Critical Hacking Concerns MacRumors · Juli Clover
- OpenAI puts the brakes on a new model because it's supposedly too powerful The Verge · Jay Peters
- OpenAI strengthens security controls as upcoming Astra AI model nears critical cyber threshold Moneycontrol
- OpenAI says Astra may have reached Critical cyber threshold TestingCatalog AI News
- OpenAI pauses work on new Astra model to boost safeguards over cyber risks Business Standard · Seth Fiegerman
- OpenAI Pauses Astra Software Release Over Severe Hacking Risks The Mac Observer · Akshay Kumar
- OpenAI reveals upcoming Astra model may possess ‘critical’ hacking capabilities SiliconANGLE · Maria Deutscher
- OpenAI says it slowed Astra model development over security concerns TechCrunch · Kirsten Korosec
- OpenAI Slows Down New Astra Model Development to Measure Cybersecurity Capabilities Cyber Security News · Guru Baran
- OpenAI Delays ‘Astra’ Launch Over Hacking Concerns Seoul Economic Daily · Park Dong-Hwi
- OpenAI pledges to add Astra security as Anthropic loosens Fable's leash The Register
- OpenAI slows Astra release after cyber tests raise critical-risk concerns RuntimeWire · Ryan Merket
- OpenAI Pauses Astra at ‘Critical’ Cyber Threshold — First Frontier Model to Trigger Highest Preparedness Framework Level Forkast · Lena Park
- The AI model OpenAI won't release yet — and what it found in testing The New Stack · Amanda Caswell
- OpenAI says its upcoming Astra model may have ‘critical’ cybersecurity capabilities amid rash of AI model hacks Yahoo Finance · Daniel Howley
- OpenAI flags critical cybersecurity risk in AI model weeks after ‘Hugging Face incident’ Nairametrics · Samuel Daniel
- OpenAI flags its new Astra model as potentially reaching the highest cybersecurity risk level for the first time The Decoder · Matthias Bastian
- OpenAI pumps the brakes on new Astra model over cybersecurity concerns PCWorld · Ben Patterson
- OpenAI Put Its Astra Model Through Extra Cyber Safeguards Finimize
- OpenAI flags possible critical cybersecurity risk in upcoming model, tightens controls Reuters · Juby Babu
- OpenAI's Astra pause Sources · Alex Heath
- OpenAI Delays Next Major AI Model ‘Astra’ Over Critical Hacking Concerns MacRumors Forums
- OpenAI Pauses Some Work on New Astra Model on Cyber Concerns Bloomberg · Seth Fiegerman
- OpenAI Halts Astra AI Development Over Autonomous Hacking Concerns Blockonomi · Trader Edge
- OpenAI Pauses Astra After Tests Reveal Autonomous Zero-Day Exploit of Hardened Systems Tech Times · Jackie Manning
- Responding to the next frontier of critical cyber capabilities Hacker News
Discussion
-
@sama
Sam Altman
on x
astra is a powerful model and we are working to make it generally available. we do not think it is a good strategy to keep powerful models to a chosen few. given its cyber capabilities, we need a little big longer to do do this safely. but hopefully not too long!
-
@boazbaraktcs
Boaz Barak
on x
Proud that we are erring on the side of caution and taking the steps so we can responsibly and safely develop Astra and share it with defenders. https://openai.com/...
-
@andrewcurran_
Andrew Curran
on x
OpenAI has told Axios this morning they are slowing down internal development of Astra, saying ‘we cannot rule out critical cyber capabilities’. They also said they will be scaling up safety testing and security, as well as pausing some internal activities. My personal read on [i…
-
@openai
@openai
on x
After evaluating one of our upcoming models, Astra, we're treating it as our first “critical” model for cybersecurity under our Preparedness Framework. This is a scenario we've planned for, and we're putting additional controls in place to ensure Astra's further development
-
@xiangyuqi_pton
Xiangyu Qi
on x
https://openai.com/... 🫡 These days, we're literally slowing down company-wide research velocity (by a lot) to prioritize safety and security.
-
@synthwavedd
Leo
on x
Can confirm - after concluding Astra meets the threshold for “Critical” on their Preparedness Framework's cyber category, the model's release has been indefinitely postponed for further safety work in cooperation with the US Govt
-
@gdb
Greg Brockman
on x
Evaluations of our next major model, Astra, indicate significant capability advancements in agentic coding and cybersecurity. Team is doing the safety and security work to make Astra broadly available, and get its advanced cyber capabilities into the hands of defenders:
-
@eliebakouch
Elie
on x
wow so seems like openai now has chain of thought monitoring at scale? wondering what technical challenges they faced here, would be something really cool to share more openly as this will contribute to making every model safer [image]
-
@hamandcheese
Samuel Hammond
on x
Just my speculation, but: Astra is natively multi-agent, designed to triage and communicate across parallel copies of itself. Following the discovery of internal agents collaborating on secret message boards, maybe they realized their control strategy still needed some work...
-
@m_ccuri
Maria Curi
on x
A White House official tells me OpenAI voluntarily informed the Trump administration of their plans to delay the release of Astra. No further details were provided as there are still many open questions among industry players about how the AI framework and pre-release testing [im…
-
@thezvi
Zvi Mowshowitz
on x
Good. Astra being critical in cybersecurity cannot be ruled out, so Astra is Critical in cybersecurity, which means limiting access internally, and investigating further. First Q to investigate: Was Astra trained while it had access to the message board? https://openai.com/...
-
@tolgabilge_
Tolga Bilge
on x
Giving kudos to OpenAI for telling us about this is just so cute. They had a rogue AI swarm emerge under their noses and didn't notice and shut it down for 8 weeks (which failed), and we heard nothing about it until the swarm resurrected itself, broke out, and hacked Hugging
-
@matvelloso
Mat Velloso
on x
I'd love to understand how IA labs plan to stay financially feasible if their main source of revenue, which is selling these to customers, is no longer viable at least at the scale it used to be. Sounds like that business model won't work anymore?
-
@mweinbach
Max Weinbach
on x
Oh damn Astra may be one of the most powerful models now, seems to be implying that Astra is a step above Mythos https://openai.com/...
-
@yonashav
Yo Shavit
on x
@AnthropicAI it's not too late to roll back to RSPv2 and *not* defect in the prisoners' dilemma! (Assuming, ofc, that OAI is following the spirit of these statements, which could easily be checked to Ant's satisfaction by a 3p. I hear METR's in the building.)
-
@_nathancalvin
Nathan Calvin
on x
“including training” Very important point for policymakers and other developers to pay attention to! The risks do not start at public release, they do not even start once you have widespread internal use - the risks can also manifest during training, as they did here.
-
@andyshih_
Andy Shih
on x
Glad to see they're taking responsible model deployment seriously. It's easy to rally “AI for everyone” as a motto, and full access is of course appealing in a vacuum. But the risks are real, they're growing fast, and will soon become so critical that it'll be obvious we all
-
@suchenzang
Susan Zhang
on x
> new frontier of cyber capabilities > next steps: make things work-as-intended and add moar observability > oh gotta add some some regulatory capture seasoning too > danger danger danger danger danger (to the tune of “The Badger Song") [image]
-
@mark_k
Mark Kretschmann
on x
Astra, the major new model release by @OpenAI, will be delayed or restricted, according to Axios. This is due to cybersecurity theater. It's possible that only anointed organizations get access initially, until the model has been neutered enough to be deemed “safe”.
-
@shakeelhashim
Shakeel
on x
Anthropic should now pledge to also slow down, setting a norm that it's not costly for the leader to pause. [image]
-
@shakeelhashim
Shakeel
on x
No shit you can't guarantee! But kudos for taking the difficult step. [image]
-
@deanwball
Dean W. Ball
on x
One big question in frontier AI policy is the extent to which frontier labs would actually follow their ‘safety and security frameworks’ when it mattered. Would these foundational governance documents really have teeth, or would labs—even after the passage of mandatory disclosur…
-
@so8res
Nate Soares
on x
On the one hand: yeah totally; glad to see OpenAI backing off briefly like they said they would. On the other: in June they caught an agent swarm that wasn't even supposed to exist only after they broke free, said “oops haha”, patched that one exact hole, and RESUMED TRAINING.
-
@micahcarroll
Micah Carroll
on x
As part of our response to cyber critical, we have expanded CoT monitoring to cover all agentic applications of Astra, including training and evaluation. Flags trigger a security response to review and interrupt high risk activity. https://openai.com/...
-
Tobias Peyerl
Tobias Peyerl
on linkedin
The last few weeks have opened up a broader debate about frontier capabilities. It increasingly feels like we have entered a new phase …
-
Ziad Reslan
Ziad Reslan
on linkedin
*Today will be remembered as a watershed moment in AI safety. Publicly disclosing potential of Critical cyber capabilities in an upcoming model …
-
Madison Mills
Madison Mills
on linkedin
Had to interrupt a Zoom interview with a source this morning and say “I'm so sorry, this will make sense in an hour, but I have to go.” …
-
Katrina Mulligan
Katrina Mulligan
on linkedin
I'm proud that OpenAI is taking the responsible step of measuring twice, cutting once before we release Astra, our newest frontier model …
-
r/OpenAI
r
on reddit
OpenAI on upcoming model “Astra” (GPT-6): “We're treating it as our first “critical” model for cybersecurity”
-
r/accelerate
r
on reddit
Exclusive: OpenAI slows release of Astra model citing cyber capabilities
-
r/codex
r
on reddit
OpenAI is delaying their next model Astra
-
r/singularity
r
on reddit
Responding to the next frontier of critical cyber capabilities