Mythos Preview is the first AI model to complete both of AISI's cyber ranges, which measure models' cyberattack capabilities; GPT-5.5 solved only one of them
In February 2026, we internally estimated that the length of cyber tasks AI models could complete had doubled every 4.7 months since late 2024 …
AI Security Institute
Context & Ripple Effects
AISI’s recent coverage had already placed Mythos Preview and GPT-5.5 at a similar level on a multi-step cyberattack simulation, while Mythos Preview also reached a 73% success rate on expert-level capture-the-flag challenges. This update separates the two on AISI’s broader range suite: Mythos Preview has now completed both ranges, while GPT-5.5 has completed one.
The result arrives alongside efforts to channel frontier cyber capability into controlled defensive use, including OpenAI’s limited Trusted Access rollout of a cyber-focused model variant. It makes AISI’s evaluations more consequential as a comparative measure of models whose cyber performance is advancing quickly.
First-order effects
Mythos Preview takes the leading position in the reported AISI cyber-range results, establishing a clearer capability distinction from GPT-5.5 on these tests.
AISI gains a concrete new benchmark outcome for assessing cyberattack-relevant model capability, increasing the urgency of scrutiny around models that can complete longer, multi-step tasks.
Second-order effects
Model developers will face stronger incentives to demonstrate cyber safeguards and controlled-access policies alongside benchmark gains, particularly where high-performing systems are positioned for defensive use.
The gap between general model capability claims and independently assessed cyber performance becomes more visible: vendors may need to compete not only on task success but also on evidence that access and deployment are bounded appropriately.
Third-order effects
If completion of demanding cyber ranges becomes more common, cyber evaluations may shift from isolated demonstrations toward a recurring release gate for frontier models, with risk management judged against rapidly improving task completion.
The pattern points to a tighter coupling between capability progress and deployment governance: defensive specialization and restricted-access programs may become more important, though the coverage does not establish which safeguards will prove sufficient.
The trend: Frontier AI is progressing from partial success on difficult cyber tasks toward reliable multi-step completion, raising the importance of standardized capability evaluation and deployment controls.
Worth paying attention to: @AISecurityInst's initial review of Anthropic's Mythos made waves... but today they publish results that show that a later checkpoint of the model is significantly more capable still. There is no deceleration.
Our cyber range results illustrate this step-up. Since our first Mythos evaluation, we received access to a newer Mythos Preview checkpoint. On a 32-step corporate network attack we estimate takes a human expert ~20 hours, this checkpoint completes the full attack in 6 /10 [image…
In life, everything is a wager. Whether you realize it or not, you are constantly making implicit and explicit predictions about the future state of reality. To live is to predict. So when you are faced with something like Mythos, and you say, “this is just ‘doomer hype’!,” what
Two independent evals this week (XBOW and UK AISI) confirmed what my team has been seeing inside Project Glasswing: Claude Mythos Preview is a step change. The UK AISI analyzed the version of Mythos available at the launch of Project Glasswing and found it completed both of the
A lot of people have been wondering about Mythos, Glasswing, and the vulns we / our partners are fixing. Today, I'm excited for us to start sharing more. (For context, I lead Glasswing @AnthropicAI.) Two independent evaluations this week—from XBOW and the UK AISI—confirm what
In November 2025 we estimated this doubling time to be 8 months. By February 2026 we had revised it to 4.7 months. Claude Mythos Preview and GPT-5.5 have since substantially exceeded even this accelerated timescale. It is currently unclear whether this is a new trend or a one-off
just to be clear about what these checkpoints represent in this chart: Mythos (new) is the version corresponding to the Glasswing launch. Mythos (early) is an earlier, only partially-trained version these latest results are about the actual Mythos Preview, the model as it was
The UK AISI found Mythos Preview is the first model to solve both their cyber ranges end-to-end. No model had ever solved the AISI's “Cooling Tower” cyber range before. We're getting it to defenders as fast as we responsibly can. More to come on our Glasswing work soon.
🚨 BREAKING: AISI tested a newer Mythos Preview checkpoint. These numbers are insane: > solved “The Last Ones”, a 20hr task, in 6/10 attempts > solved the *previously unsolved* “Cooling Tower” task in 3/10 attempts > cyber task time horizons went from doubling every 8mos to [image…
The UK's state AI Security iIstitute findings: 1) Mythos is a big gain in cyber capabilities. But so is GPT-5.5 2) It is hard to establish an upper bound on Mythos/GPT-5.5, which appear to be limited by tokens used, rather than ability. 3) Capability doubling time is 4.5 months […
The new version completely smashes GPT-5.5 and the previous Mythos version. Before Mythos Preview completed the cyber range 3 out of 10 times. The new version completed it 6 out of 10 times and is much more efficient! [image]
Our evaluations show that frontier AI's cyber capabilities are advancing quickly. The length of cyber tasks frontier models can complete has been doubling every few months, and this rate has become faster over time, with recent models exceeding our previous trends. 🧵 [image]
For the past 2 months, XBOW has been testing Mythos Preview under embargo as part of a select early-access group. Today, we can finally share what we found. The headline: Mythos Preview is a major advance. It is substantially better than prior models at finding vulnerability [vid…
I don't understand the path forward for Mythos releases. Google & OpenAI will have equivalent models, and they are approaching AI cyber risk guardrails differently, so they will presumably just release their versions. How does Anthropic get out of the government approval path?