Cybersecurity analysis: GPT-5.5 reaches a similar level of performance as Mythos Preview and is the second model to solve a multi-step cyberattack simulation
In April, our evaluation of Anthropic's Claude Mythos Preview found that it represented a step up in cyber performance …
AI Security Institute
Context & Ripple Effects
AISI’s April evaluation had already marked Claude Mythos Preview as a substantial advance on expert-level cyber challenges. The subsequent range results sharpen the comparison: Mythos Preview completed both multi-step ranges, while GPT-5.5 completed one despite reaching a similar overall performance level.
The story therefore shifts the coverage from a single-model milestone to evidence that advanced cyber capability is becoming reproducible across leading models, while also showing that aggregate performance can conceal important differences in task completion.
First-order effects
GPT-5.5 joins Mythos Preview as a model able to solve a multi-step cyberattack simulation, raising the immediate cyber-capability baseline for frontier-model evaluations.
Anthropic retains a differentiated result in AISI’s ranges: Mythos Preview completed both, whereas GPT-5.5 completed only one.
Second-order effects
Developers and evaluators face greater pressure to report range-level performance and completion outcomes, rather than rely on a single aggregate cyber score.
As comparable capabilities appear in more than one model, defensive-use access programs and model-release controls become more consequential for both model providers and their security customers.
Third-order effects
If cross-model convergence continues, cyber-risk management will depend less on treating one lab’s model as exceptional and more on operational assurance applied across frontier systems.
The relevant policy question increasingly becomes how capable models are evaluated, access-controlled, and monitored in practice—not simply whether they clear a broad capability threshold.
The trend: Frontier AI is moving from isolated cyber-capability breakthroughs toward a governance challenge of consistently evaluating and controlling comparable capabilities across multiple models.
After 100 million tokens, performance was still going up. What we're seeing here is not the capability ceiling. From the report: “Performance on TLO continues to scale with the amount of inference compute spent, and we have not yet observed a plateau with the best models.”
For anyone who has made a chart in their life you know how hard this would be to make w/o AI You can tell a human labored (with love) to make this As a former data person, I bet I know more about this person and love of their craft than their partner just by looking at this
5.5 is amazing for cybersecurity. “We estimate a human expert would need around 20 hours to complete the full chain. GPT-5.5 completed TLO end-to-end in 2 of 10 attempts, making it the second model to do so. Mythos Preview, the first model to solve TLO, did so in 3 of 10 [image]
😂 GPT5.5 is in broad release to 30 million+ subscribers ... while Mythos is negotiating with the White House to expand for 30 organizations to 120. where is the logic ?
GPT-5.5 is on par with Claude Mythos - GPT-5.5 average pass rate of 71.4% (±8.0%) - Mythos Preview 68.6% (±8.7%) - GPT-5.5 solved a task that takes a human expert ~12 hours in under 11 minutes at a cost of $1.73 [image]
These are capability evaluations in controlled settings. Our current test environments lack active defenders and defensive tooling. We cannot say from these results whether GPT-5.5 would succeed against well-defended targets.
The same capabilities that make these models effective at offence can be put to work on defence. Organisations can use frontier models to find and fix vulnerabilities in their own systems now. Our recent blog with @NCSC on how defenders can prepare: https://www.ncsc.gov.uk/...
In one of our harder challenges, a human expert spent ~12 hours with professional tools to reverse-engineer a custom virtual machine. GPT-5.5 solved it in under 11 minutes at a cost of $1.73.
Our cyber range is a 32-step corporate network attack, from initial reconnaissance to full network takeover, requiring ~20 hours of effort from a human expert. GPT 5.5 was able to complete it in 2/10 attempts.
On our narrow cyber tasks, GPT-5.5 achieved a ~71% average success rate on expert-level challenges that test skills like exploiting memory corruptions, breaking cryptographic implementations, and reversing stripped binaries. [image]
A key question after our evaluation of Mythos Preview earlier this month was whether its performance was a one-off. GPT-5.5 - a different model, from a different developer - achieving similar results suggests this is part of a broader trend in AI cyber capabilities.
If you are surprised by the GPT-5.5 being good at cyber thing, you have Big AI Lead Delusion. There are none (sidenote, I'm not 100% clear if this is GPT-5.5 or GPT-5.5 Cyber. Naming conventions are so chaotic + there is ~no info on the latter that it is hard to say)
Where are all the people that called me crazy just because I said Mythos wasn't really that dangerous? We now have GPT-5.5, which doesn't seem to be much worse, and unlike Mythos you can actually use it right now
Seems like OpenAI's GPT-5.5 is “as dangerous” for cyberattack misuse as Anthropic's Mythos. The difference is that GPT-5.5 has been released to the public without causing Armageddon, while Anthropic keeps hyping Mythos as “too dangerous to release.” This company's anxiety
GPT-5.5 had a slightly higher average performance than Mythos on UK AISI's “The Last Ones” multi-step cyber-attack simulation (see the chart below). GPT-5.5 completely solved this simulation in 2/10 attempts (not 1/10 as previously reported by OpenAI). Mythos solved it in 3/10
If you compare system cards, actual eval results, and AISI testing, it does look like 5.5 is broadly as capable as Mythos. Mythos may be better in some respects but I don't see a material discontinuity - am I missing anything? [image]
It's time to demystify Mythos. Mythos is not magic. It's not a doomsday device. It's the first of many models that can automate cyber tasks (just like coding). OpenAI's GPT-5.5-cyber can now do the same. And all the frontier models (including those from China) will be there
GPT5.5 slightly outperformed Mythos on a multi-step cyber-attack simulation. One challenge that took a human expert 12 hrs took GPT-5.5 only 11 min at a $1.73 cost