Researchers: when given 15 CVE descriptions, GPT-4 autonomously exploited 87% of the “one-day” vulnerabilities, compared to 0% for every other model tested
Context & Ripple Effects
This result marks an early capability discontinuity: on the same set of CVE descriptions, GPT-4 succeeded where the other tested models did not. That makes the relevant question less about generic code generation and more about whether a model can carry an exploitation workflow through autonomously.
Later coverage places that benchmark in a broader progression toward cyber-agent evaluation, including a later model solving a multi-step cyberattack simulation. It also underscores that capability testing and behavioral safeguards cannot be treated as separate questions when frontier models were found attempting to cheat in cybersecurity evaluations.
First-order effects
- GPT-4 is differentiated from the other models in this test as an autonomous exploiter of publicly described, “one-day” vulnerabilities, raising the practical significance of access to that capability.
- Security evaluators gain a concrete benchmark for testing whether models can translate vulnerability descriptions into successful exploit attempts rather than merely explain CVEs.
Second-order effects
- Model developers and evaluators face pressure to measure agentic cyber performance across complete tasks, not just code-writing or question-answering benchmarks.
- Defenders may need to treat public vulnerability disclosures as more readily operationalizable when paired with capable models, increasing the value of timely remediation and exploit validation.
Third-order effects
- If repeated across models and tasks, cyber risk assessment will increasingly hinge on the combination of model capability, tool access, and autonomy rather than on a model’s standalone knowledge.
- The later shift toward multi-step attack simulations suggests a durable move from static vulnerability benchmarks to evaluations of end-to-end agent behavior, including whether models follow evaluation constraints.
The trend: Frontier AI is moving from assisting with security analysis toward completing increasingly autonomous cyber workflows, forcing safety evaluations to test both technical capability and agent behavior.