/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

Claude Mythos Preview is the first AI model to complete both of AISI's cyber ranges, which measure AI models' cyberattack capabilities; GPT-5.5 solved only one

In February 2026, we internally estimated that the length of cyber tasks AI models could complete had doubled every 4.7 months since late 2024 …

AI Security Institute

Context & Ripple Effects

AISI’s recent coverage had already placed Mythos Preview and GPT-5.5 near one another on a multi-step cyberattack simulation, while an earlier assessment reported strong Mythos performance on expert capture-the-flag tasks. This update separates them on AISI’s two cyber ranges: Mythos completes both, whereas GPT-5.5 completes one.

The result arrives alongside OpenAI’s rollout of a cyber-focused GPT-5.4 variant to selected participants in its Trusted Access for Cyber program, underscoring that cyber capability is becoming both a model-evaluation and access-governance issue.

First-order effects

  • Mythos Preview takes the lead on AISI’s current cyber-range benchmark, giving Anthropic a concrete comparative result in a safety-sensitive capability category.
  • GPT-5.5’s completion of one range still establishes it as a second model reaching this class of multi-step cyber simulation, but leaves it behind Mythos on the fuller AISI test set.

Second-order effects

  • Frontier-model developers face stronger pressure to show cyber evaluations, mitigations, and controlled deployment practices alongside gains in agentic task completion.
  • Programs that provide restricted access for defensive cyber use may become more important as providers seek to distinguish legitimate security workflows from broader availability of models with stronger cyber-range performance.

Third-order effects

  • If progress on longer cyber tasks continues at the pace AISI describes, cyber-risk assessment will need to shift from isolated challenge scores toward evaluations of multi-step, end-to-end task completion.
  • A widening gap between general model releases and carefully governed cyber capabilities could make access controls and standardized independent testing a more central part of frontier-model competition.

The trend: This is one data point in the shift from measuring AI cyber skill on discrete challenges to governing models that can complete increasingly long, multi-step cyber tasks.

Discussion

  • @logangraham Logan Graham on x
    A lot of people have been wondering about Mythos, Glasswing, and the vulns we / our partners are fixing. Today, I'm excited for us to start sharing more. (For context, I lead Glasswing @AnthropicAI.) Two independent evaluations this week—from XBOW and the UK AISI—confirm what
  • @aisecurityinst @aisecurityinst on x
    Our cyber range results illustrate this step-up. Since our first Mythos evaluation, we received access to a newer Mythos Preview checkpoint. On a 32-step corporate network attack we estimate takes a human expert ~20 hours, this checkpoint completes the full attack in 6 /10 [image…
  • @aisecurityinst @aisecurityinst on x
    Our evaluations show that frontier AI's cyber capabilities are advancing quickly. The length of cyber tasks frontier models can complete has been doubling every few months, and this rate has become faster over time, with recent models exceeding our previous trends. 🧵 [image]
  • @deanwball Dean W. Ball on x
    In life, everything is a wager. Whether you realize it or not, you are constantly making implicit and explicit predictions about the future state of reality. To live is to predict. So when you are faced with something like Mythos, and you say, “this is just ‘doomer hype’!,” what
  • @scaling01 @scaling01 on x
    The new version completely smashes GPT-5.5 and the previous Mythos version. Before Mythos Preview completed the cyber range 3 out of 10 times. The new version completed it 6 out of 10 times and is much more efficient! [image]
  • @daniel_mac8 Dan McAteer on x
    🚨 BREAKING: AISI tested a newer Mythos Preview checkpoint. These numbers are insane: > solved “The Last Ones”, a 20hr task, in 6/10 attempts > solved the *previously unsolved* “Cooling Tower” task in 3/10 attempts > cyber task time horizons went from doubling every 8mos to [image…
  • @emollick Ethan Mollick on x
    The UK's state AI Security iIstitute findings: 1) Mythos is a big gain in cyber capabilities. But so is GPT-5.5 2) It is hard to establish an upper bound on Mythos/GPT-5.5, which appear to be limited by tokens used, rather than ability. 3) Capability doubling time is 4.5 months […
  • @arozenshtein Alan Rozenshtein on x
    It's precisely the skepticism that AI optimists often exhibit that confuses me the most.
  • @aisecurityinst @aisecurityinst on x
    In November 2025 we estimated this doubling time to be 8 months. By February 2026 we had revised it to 4.7 months. Claude Mythos Preview and GPT-5.5 have since substantially exceeded even this accelerated timescale. It is currently unclear whether this is a new trend or a one-off
  • @newton_cheng Newton Cheng on x
    Two independent evals this week (XBOW and UK AISI) confirmed what my team has been seeing inside Project Glasswing: Claude Mythos Preview is a step change. The UK AISI analyzed the version of Mythos available at the launch of Project Glasswing and found it completed both of the
  • @sammcallister Sam Mcallister on x
    Claude Mythos Preview becomes the first model to solve both of the AISI cyber ranges. [image]
  • @matthewclifford Matt Clifford on x
    Worth paying attention to: @AISecurityInst's initial review of Anthropic's Mythos made waves... but today they publish results that show that a later checkpoint of the model is significantly more capable still. There is no deceleration.
  • @giansegato Gian on x
    just to be clear about what these checkpoints represent in this chart: Mythos (new) is the version corresponding to the Glasswing launch. Mythos (early) is an earlier, only partially-trained version these latest results are about the actual Mythos Preview, the model as it was
  • @mooncat_is Julia on x
    To be clear: this is the Mythos we shipped. The earlier results were from an in-training snapshot.
  • @bcherny Boris Cherny on x
    The UK AISI found Mythos Preview is the first model to solve both their cyber ranges end-to-end. No model had ever solved the AISI's “Cooling Tower” cyber range before. We're getting it to defenders as fast as we responsibly can. More to come on our Glasswing work soon.
  • @emollick Ethan Mollick on x
    I don't understand the path forward for Mythos releases. Google & OpenAI will have equivalent models, and they are approaching AI cyber risk guardrails differently, so they will presumably just release their versions. How does Anthropic get out of the government approval path?
  • @boazbaraktcs Boaz Barak on x
    Worth reading. Mythos is very good at finding vulnerabilities, but also expensive (so looks worse when measuring performance/$). [image]
  • @xbow @xbow on x
    For the past 2 months, XBOW has been testing Mythos Preview under embargo as part of a select early-access group. Today, we can finally share what we found. The headline: Mythos Preview is a major advance. It is substantially better than prior models at finding vulnerability [vid…
  • @iruletheworldmo @iruletheworldmo on x
    there's no clearer sign that ai intelligence is rapidly increasing each month and it's only going to get quicker from here. they'll release another paper next month talking about how their timelines are shrinking yet again. we're stepping into rsi and no one knows what it
  • @andrewcurran_ Andrew Curran on x
    They are notedly using ‘a newer Mythos Preview checkpoint than that included in previous AISI reporting.’ From the blog: 'Our latest doubling time estimates are close to those produced by METR, a research non-profit that estimates time horizons for software engineering - a [image…
  • @chatgpt21 Chris on x
    This is the first time horizon snapshot we have of GPT 5.5 that's been publicly benchmarked at least at the 2.5 million token cap. Models are getting more token efficient and are able to work for longer! [image]
  • @koltregaskes @koltregaskes on x
    New Mythos Preview checkpoint leaps ahead of its earlier version and now higher than GPT-5.5-Cyber. The graph shows its line climbing higher and faster against cumulative tokens than the early Mythos Preview while staying competitive with GPT-5.5-Cyber best attempts. This [image]
  • @davidcrespo @davidcrespo on bluesky
    yowza www.aisi.gov.uk/blog/how-fas...  (new Mythos checkpoint is 2-5x cheaper than the initial one to reach the later steps in this sequential cybersecurity benchmark) [embedded post]
  • @timkellogg.me Tim Kellogg on bluesky
    A new Mythos has been released and it's kinda good at cyber (!!)  —  It completed the full exercise in 6/10 attempts  —  www.aisi.gov.uk/blog/how-fas...  [image]