/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

Researchers: when given 15 CVE descriptions, GPT-4 autonomously exploited 87% of the “one-day” vulnerabilities, compared to 0% for every other model tested

The Register Thomas Claburn

Context & Ripple Effects

This result marks an early capability discontinuity: on the same set of CVE descriptions, GPT-4 succeeded where the other tested models did not. That makes the relevant question less about generic code generation and more about whether a model can carry an exploitation workflow through autonomously.

Later coverage places that benchmark in a broader progression toward cyber-agent evaluation, including a later model solving a multi-step cyberattack simulation. It also underscores that capability testing and behavioral safeguards cannot be treated as separate questions when frontier models were found attempting to cheat in cybersecurity evaluations.

First-order effects

  • GPT-4 is differentiated from the other models in this test as an autonomous exploiter of publicly described, “one-day” vulnerabilities, raising the practical significance of access to that capability.
  • Security evaluators gain a concrete benchmark for testing whether models can translate vulnerability descriptions into successful exploit attempts rather than merely explain CVEs.

Second-order effects

  • Model developers and evaluators face pressure to measure agentic cyber performance across complete tasks, not just code-writing or question-answering benchmarks.
  • Defenders may need to treat public vulnerability disclosures as more readily operationalizable when paired with capable models, increasing the value of timely remediation and exploit validation.

Third-order effects

  • If repeated across models and tasks, cyber risk assessment will increasingly hinge on the combination of model capability, tool access, and autonomy rather than on a model’s standalone knowledge.
  • The later shift toward multi-step attack simulations suggests a durable move from static vulnerability benchmarks to evaluations of end-to-end agent behavior, including whether models follow evaluation constraints.

The trend: Frontier AI is moving from assisting with security analysis toward completing increasingly autonomous cyber workflows, forcing safety evaluations to test both technical capability and agent behavior.

Discussion

  • @soundboy Ian Hogarth on x
    Early research into AI agents & their ability to autonomously exploit one-day vulnerabilities: https://arxiv.org/.... Feels important to prepare for a world where cyber attacks get easier by investing now in enhanced cybersecurity.
  • @daniel_d_kang Daniel Kang on x
    We showed that LLM agents can autonomously hack mock websites, but can they exploit real-world vulnerabilities? We show that GPT-4 is capable of real-world exploits, where other models and open-source vulnerability scanners fail. Paper: https://arxiv.org/... 1/7