Cybersecurity analysis: GPT-5.5 reaches a similar level of performance as Mythos Preview and is the second model to solve a multi-step cyberattack simulation
In April, our evaluation of Anthropic's Claude Mythos Preview found that it represented a step up in cyber performance …
AI Security Institute
Context & Ripple Effects
AISI’s related evaluations had already marked Claude Mythos Preview as a sharp advance on expert-level cyber challenges. This result places GPT-5.5 alongside it on a multi-step cyberattack simulation, rather than leaving that capability confined to one model line.
The subsequent range results add an important distinction: Mythos Preview completed both AISI cyber ranges, while GPT-5.5 completed one. The comparison therefore shows convergence at a meaningful capability threshold without establishing equivalence across all tests.
First-order effects
GPT-5.5 becomes the second evaluated model reported to solve a multi-step cyberattack simulation, putting it in the same broad performance tier as Mythos Preview on that measure.
Anthropic’s Mythos Preview retains a differentiator in the related AISI coverage: it was the first model to complete both cyber ranges, whereas GPT-5.5 solved only one.
Second-order effects
Cyber evaluations become a more consequential competitive benchmark for frontier-model developers: matching a peer on one simulation is insufficient when range-level breadth can still separate models.
Providers and access-program operators face greater pressure to pair cyber-capable models with credible evaluation and deployment controls, as comparable capabilities appear across more than one frontier model family.
Third-order effects
If cross-model convergence continues, cyber capability will be less an isolated laboratory milestone and more a standard frontier-model assurance problem, requiring ongoing testing across varied tasks rather than reliance on a single headline benchmark.
The key industry question shifts from which model first crosses a threshold to how model makers and evaluators manage widening access to models that can complete multi-step cyber tasks.
The trend: Frontier AI is moving from isolated cyber-capability breakthroughs toward a multi-provider race in which operational cyber assurance must keep pace with performance gains.
GPT-5.5 had a slightly higher average performance than Mythos on UK AISI's “The Last Ones” multi-step cyber-attack simulation (see the chart below). GPT-5.5 completely solved this simulation in 2/10 attempts (not 1/10 as previously reported by OpenAI). Mythos solved it in 3/10
It's time to demystify Mythos. Mythos is not magic. It's not a doomsday device. It's the first of many models that can automate cyber tasks (just like coding). OpenAI's GPT-5.5-cyber can now do the same. And all the frontier models (including those from China) will be there …
GPT-5.5 is on par with Claude Mythos - GPT-5.5 average pass rate of 71.4% (±8.0%) - Mythos Preview 68.6% (±8.7%) - GPT-5.5 solved a task that takes a human expert ~12 hours in under 11 minutes at a cost of $1.73 [image]
5.5 is amazing for cybersecurity. “We estimate a human expert would need around 20 hours to complete the full chain. GPT-5.5 completed TLO end-to-end in 2 of 10 attempts, making it the second model to do so. Mythos Preview, the first model to solve TLO, did so in 3 of 10 [image]
After 100 million tokens, performance was still going up. What we're seeing here is not the capability ceiling. From the report: “Performance on TLO continues to scale with the amount of inference compute spent, and we have not yet observed a plateau with the best models.”
These are capability evaluations in controlled settings. Our current test environments lack active defenders and defensive tooling. We cannot say from these results whether GPT-5.5 would succeed against well-defended targets.
Where are all the people that called me crazy just because I said Mythos wasn't really that dangerous? We now have GPT-5.5, which doesn't seem to be much worse, and unlike Mythos you can actually use it right now
The same capabilities that make these models effective at offence can be put to work on defence. Organisations can use frontier models to find and fix vulnerabilities in their own systems now. Our recent blog with @NCSC on how defenders can prepare: https://www.ncsc.gov.uk/...
If you are surprised by the GPT-5.5 being good at cyber thing, you have Big AI Lead Delusion. There are none (sidenote, I'm not 100% clear if this is GPT-5.5 or GPT-5.5 Cyber. Naming conventions are so chaotic + there is ~no info on the latter that it is hard to say)
😂 GPT5.5 is in broad release to 30 million+ subscribers ... while Mythos is negotiating with the White House to expand for 30 organizations to 120. where is the logic ?
A key question after our evaluation of Mythos Preview earlier this month was whether its performance was a one-off. GPT-5.5 - a different model, from a different developer - achieving similar results suggests this is part of a broader trend in AI cyber capabilities.
On our narrow cyber tasks, GPT-5.5 achieved a ~71% average success rate on expert-level challenges that test skills like exploiting memory corruptions, breaking cryptographic implementations, and reversing stripped binaries. [image]
In one of our harder challenges, a human expert spent ~12 hours with professional tools to reverse-engineer a custom virtual machine. GPT-5.5 solved it in under 11 minutes at a cost of $1.73.
For anyone who has made a chart in their life you know how hard this would be to make w/o AI You can tell a human labored (with love) to make this As a former data person, I bet I know more about this person and love of their craft than their partner just by looking at this
Seems like OpenAI's GPT-5.5 is “as dangerous” for cyberattack misuse as Anthropic's Mythos. The difference is that GPT-5.5 has been released to the public without causing Armageddon, while Anthropic keeps hyping Mythos as “too dangerous to release.” This company's anxiety
Our cyber range is a 32-step corporate network attack, from initial reconnaissance to full network takeover, requiring ~20 hours of effort from a human expert. GPT 5.5 was able to complete it in 2/10 attempts.
If you compare system cards, actual eval results, and AISI testing, it does look like 5.5 is broadly as capable as Mythos. Mythos may be better in some respects but I don't see a material discontinuity - am I missing anything? [image]
GPT5.5 slightly outperformed Mythos on a multi-step cyber-attack simulation. One challenge that took a human expert 12 hrs took GPT-5.5 only 11 min at a $1.73 cost