Analysis: recent open weight models lag frontier closed models' cyber capabilities by 4 to 7 months, a narrower gap than the 6 to 10 months through most of 2025
The AI Security Institute’s comparison indicates that the cyber-capability lead held by frontier closed models has contracted from the range observed through most of 2025. That matters because cyber performance is becoming a more direct measure of how much capability is available outside tightly controlled model access.
The measured cyber-capability advantage of frontier closed models is now shorter, reducing the time buffer between closed-model advances and comparable open-weight availability.
Organizations assessing cyber risk from model deployment must treat recent open-weight systems as closer to the frontier than they were during most of 2025.
Second-order effects
Closed-model providers have less room to present a capability lead alone as a durable access-control distinction; safety practices and deployment controls become more consequential differentiators.
Security teams and policymakers face a narrower window to adapt evaluations and access-governance policies before advanced cyber capabilities spread through models that can be run beyond a provider’s platform.
Third-order effects
If the gap continues to compress, cyber-capable AI may increasingly be governed by the availability and handling of model weights rather than only by rules imposed at hosted-model interfaces.
The result strengthens the case for access governance that distinguishes between model capability, release format, and operational safeguards, though one comparison alone cannot establish a persistent convergence rate.
The trend: This is one data point in the convergence of open-weight and frontier closed-model cyber capabilities, raising the strategic importance of model-access governance.
GPT-5.6 Sol is the state of the art in cyber. Seeing significant results in applying it to finding and fixing novel vulnerabilities. Sign up as a defender to use it to secure your systems: https://openai.com/...
open-models are lagging frontier models by 7 months on UK AISI's long-horizon cyber ranges specifically, GLM-5.2 is equivalent to Opus 4.5 on “The Last Ones” cyber range on narrower short horizon tasks GLM-5.2 performs comparably to Opus 4.6 [image]
On our cyber range “The Last Ones”, GLM-5.2 matches Opus 4.5, released ~7 months before it, while DeepSeek's V4-Pro falls below Sonnet 4.5, from ~7 months before it. [image]
1. OpenAI's 5.6 Sol beats Anthropic's still not fully available Mythos in this hacking evaluation. 2. The best open model, Kimi K3, will probably be quite similar to the US proprietary leaders. Everyone now has frontier hacking capabilities.
K3 analysis wen Also, reminder to Americans - we could have this kind of state capacity at home. Let's properly fund and unmuzzle CAISI! https://x.com/...
These findings indicate a narrow window before today's frontier cyber capabilities may become widely accessible without safeguards. It's uncertain how the gap will evolve, but AISI will continue to track it and intends to test Kimi K3 once its weights are released.
Our first public analysis of the open/closed weight gap in frontier cyber capabilities finds it is 4-7 months with GLM-5.2 and DeepSeek V4-Pro, narrowing from 6-10 months through most of 2025. Advanced capabilities are reaching less safeguarded open models faster than before. 🧵 […
Our open weight model evaluations were largely unimpeded by safeguards. Of the two recent open models we tested, DeepSeek V4-Pro occasionally refused narrow cyber tasks, but this was easily circumvented by a small number of repeat attempts at refused tasks.
On our narrow cyber tasks, GLM-5.2 matches Opus 4.6 and GPT-5.3-Codex, released 4 months before it. DeepSeek V4-Pro matches Opus 4.5, released 5 months before it.
Models like Kimi K3, Grok 4.5, and Muse 1.1 may prevent the dominance of 2-3 frontier labs with 90% inference margins from hurting other AI ecosystem layers
Based on internal evals: ▪️ Kimi K3 is top-tier at cybersecurity There is chatter on X that Moonshot benchmark-overfit. These are stealth evals. Model has raw IQ. ▪️ Sol is a leap ahead in cyber capability At a significantly higher cost, but quite remarkable still. ▪️ Fable
One thing that I should clarify (which my moot stochasm pointed out) is that I first say that Artificial Intelligence Index shouldn't be used to estimate the true strength of models, but then I use it to forecast Kimi-K3's ECI. That's a bit silly. Both indexes are highly [image]
Have Chinese AI Models Caught Up to the US Frontier? I have spent the last 2 days writing this article. It should settle the debate once and for all. https://scaling01.substack.com/ ... [image]
We ran Kimi K3 on a private cybersecurity benchmark. TL;DR: Kimi K3 is the workhorse for cyber security tasks at great recall/precision/price. GPT 5.6 is best recall/precision but at 7x higher cost per run. For context, https://deepsec.sh/ is an open-source cyber harness
I estimated both the forward- and backward-looking gap between Chinese and US frontier models Kimi-K3 is currently 4.37 to 5.29 months behind US frontier models (backward-looking) Chinese models are projected to catch up Mythos-Preview by end of December 2026 (forward-looking), […