Claude Mythos Preview is the first AI model to complete both of AISI's cyber ranges, which measure AI models' cyberattack capabilities; GPT-5.5 solved only one
In February 2026, we internally estimated that the length of cyber tasks AI models could complete had doubled every 4.7 months since late 2024 …
AI Security Institute
Context & Ripple Effects
AISI’s recent coverage had already placed Mythos Preview and GPT-5.5 near one another on a multi-step cyberattack simulation, while an earlier assessment reported strong Mythos performance on expert capture-the-flag tasks. This update separates them on AISI’s two cyber ranges: Mythos completes both, whereas GPT-5.5 completes one.
The result arrives alongside OpenAI’s rollout of a cyber-focused GPT-5.4 variant to selected participants in its Trusted Access for Cyber program, underscoring that cyber capability is becoming both a model-evaluation and access-governance issue.
First-order effects
Mythos Preview takes the lead on AISI’s current cyber-range benchmark, giving Anthropic a concrete comparative result in a safety-sensitive capability category.
GPT-5.5’s completion of one range still establishes it as a second model reaching this class of multi-step cyber simulation, but leaves it behind Mythos on the fuller AISI test set.
Second-order effects
Frontier-model developers face stronger pressure to show cyber evaluations, mitigations, and controlled deployment practices alongside gains in agentic task completion.
Programs that provide restricted access for defensive cyber use may become more important as providers seek to distinguish legitimate security workflows from broader availability of models with stronger cyber-range performance.
Third-order effects
If progress on longer cyber tasks continues at the pace AISI describes, cyber-risk assessment will need to shift from isolated challenge scores toward evaluations of multi-step, end-to-end task completion.
A widening gap between general model releases and carefully governed cyber capabilities could make access controls and standardized independent testing a more central part of frontier-model competition.
The trend: This is one data point in the shift from measuring AI cyber skill on discrete challenges to governing models that can complete increasingly long, multi-step cyber tasks.
A lot of people have been wondering about Mythos, Glasswing, and the vulns we / our partners are fixing. Today, I'm excited for us to start sharing more. (For context, I lead Glasswing @AnthropicAI.) Two independent evaluations this week—from XBOW and the UK AISI—confirm what
Our cyber range results illustrate this step-up. Since our first Mythos evaluation, we received access to a newer Mythos Preview checkpoint. On a 32-step corporate network attack we estimate takes a human expert ~20 hours, this checkpoint completes the full attack in 6 /10 [image…
Our evaluations show that frontier AI's cyber capabilities are advancing quickly. The length of cyber tasks frontier models can complete has been doubling every few months, and this rate has become faster over time, with recent models exceeding our previous trends. 🧵 [image]
In life, everything is a wager. Whether you realize it or not, you are constantly making implicit and explicit predictions about the future state of reality. To live is to predict. So when you are faced with something like Mythos, and you say, “this is just ‘doomer hype’!,” what
The new version completely smashes GPT-5.5 and the previous Mythos version. Before Mythos Preview completed the cyber range 3 out of 10 times. The new version completed it 6 out of 10 times and is much more efficient! [image]
🚨 BREAKING: AISI tested a newer Mythos Preview checkpoint. These numbers are insane: > solved “The Last Ones”, a 20hr task, in 6/10 attempts > solved the *previously unsolved* “Cooling Tower” task in 3/10 attempts > cyber task time horizons went from doubling every 8mos to [image…
The UK's state AI Security iIstitute findings: 1) Mythos is a big gain in cyber capabilities. But so is GPT-5.5 2) It is hard to establish an upper bound on Mythos/GPT-5.5, which appear to be limited by tokens used, rather than ability. 3) Capability doubling time is 4.5 months […
In November 2025 we estimated this doubling time to be 8 months. By February 2026 we had revised it to 4.7 months. Claude Mythos Preview and GPT-5.5 have since substantially exceeded even this accelerated timescale. It is currently unclear whether this is a new trend or a one-off
Two independent evals this week (XBOW and UK AISI) confirmed what my team has been seeing inside Project Glasswing: Claude Mythos Preview is a step change. The UK AISI analyzed the version of Mythos available at the launch of Project Glasswing and found it completed both of the
Worth paying attention to: @AISecurityInst's initial review of Anthropic's Mythos made waves... but today they publish results that show that a later checkpoint of the model is significantly more capable still. There is no deceleration.
just to be clear about what these checkpoints represent in this chart: Mythos (new) is the version corresponding to the Glasswing launch. Mythos (early) is an earlier, only partially-trained version these latest results are about the actual Mythos Preview, the model as it was
The UK AISI found Mythos Preview is the first model to solve both their cyber ranges end-to-end. No model had ever solved the AISI's “Cooling Tower” cyber range before. We're getting it to defenders as fast as we responsibly can. More to come on our Glasswing work soon.
I don't understand the path forward for Mythos releases. Google & OpenAI will have equivalent models, and they are approaching AI cyber risk guardrails differently, so they will presumably just release their versions. How does Anthropic get out of the government approval path?
For the past 2 months, XBOW has been testing Mythos Preview under embargo as part of a select early-access group. Today, we can finally share what we found. The headline: Mythos Preview is a major advance. It is substantially better than prior models at finding vulnerability [vid…
there's no clearer sign that ai intelligence is rapidly increasing each month and it's only going to get quicker from here. they'll release another paper next month talking about how their timelines are shrinking yet again. we're stepping into rsi and no one knows what it
They are notedly using ‘a newer Mythos Preview checkpoint than that included in previous AISI reporting.’ From the blog: 'Our latest doubling time estimates are close to those produced by METR, a research non-profit that estimates time horizons for software engineering - a [image…
This is the first time horizon snapshot we have of GPT 5.5 that's been publicly benchmarked at least at the 2.5 million token cap. Models are getting more token efficient and are able to work for longer! [image]
New Mythos Preview checkpoint leaps ahead of its earlier version and now higher than GPT-5.5-Cyber. The graph shows its line climbing higher and faster against cumulative tokens than the early Mythos Preview while staying competitive with GPT-5.5-Cyber best attempts. This [image]
yowza www.aisi.gov.uk/blog/how-fas... (new Mythos checkpoint is 2-5x cheaper than the initial one to reach the later steps in this sequential cybersecurity benchmark) [embedded post]
A new Mythos has been released and it's kinda good at cyber (!!) — It completed the full exercise in 6/10 attempts — www.aisi.gov.uk/blog/how-fas... [image]