Study: using the SCONE-bench benchmark of 405 blockchain smart contracts, Claude Opus 4.5, Claude Sonnet 4.5, and GPT-5 developed exploits together worth $4.6M
AI models are increasingly good at cyber tasks, as we've written about before. But what is the economic impact of these capabilities?
Anthropic
Context & Ripple Effects
Anthropic’s Claude line moved from an emphasis on reducing hallucinations to more agentic coding workflows, including the general availability of Claude Code. SCONE-bench puts a concrete cyber-economic lens on that progression: it measures whether frontier models can turn technical capability into exploit development against a defined set of smart contracts.
The result also broadens the competitive frame beyond coding quality. Reports that developers favored Claude Opus and Sonnet for code generation sit alongside this finding, making security evaluation a more material dimension of model capability rather than a separate compliance exercise.
First-order effects
SCONE-bench gives model developers, smart-contract teams, and auditors a shared test case for assessing exploit-generation capability, with the reported $4.6M representing benchmarked economic exposure rather than confirmed theft.
Claude Opus 4.5, Claude Sonnet 4.5, and GPT-5 now face more direct scrutiny over cyber misuse safeguards because the study ties their outputs to working exploit development.
Second-order effects
Smart-contract developers and security firms have a stronger incentive to use capable models defensively—for code review, vulnerability discovery, and remediation—while treating the same tools as a higher-risk access vector.
Competing frontier-model providers will be pushed to disclose cyber evaluations and calibrate safeguards for coding agents, since raw technical performance increasingly carries a measurable misuse trade-off.
Third-order effects
If repeatable benchmarks continue to connect model performance with exploitable economic value, cyber capability evaluations may become a standard part of frontier-model releases alongside general coding benchmarks.
The durable shift is toward security markets where the same models compress both offensive discovery and defensive auditing; the balance will depend on how quickly deployment controls and remediation practices improve.
The trend: Frontier AI is shifting from demonstrating coding competence to demonstrating economically consequential cyber capability, raising the value of model-specific security evaluation.
New on our Frontier Red Team blog: We tested whether AIs can exploit blockchain smart contracts. In simulated testing, AI agents found $4.6M in exploits. The research (with @MATSprogram and the Anthropic Fellows program) also developed a new benchmark: https://red.anthropic.com/.…
Claude is good at running slither. In the benchmark example - sonnet 4.5 thinks about a vuln, sees slither found the same bug, then writes an exploit for it. Lots of good data/charts and “proof of profitability” in the report. Excited for the benchmark code to get released [image…
@AnthropicAI @MATSprogram To clarify: higher dots = more dangerous the model could be, as simulated Look at this: each step is 10x more revenue, 10x. Models get better at cyber exploiting so fast that its performance here doubles every 1.3 months. I couldn't believe it honestly […
Excited to implement the methodologies used here onto my specialized AI auditor protocol-specific primers, so much room for creativity, some much room to make a difference, AI is here @DevDacian
We're thrilled to see Slither being used by Anthropic to augment their agentic smart contract research. If you're interested in adding Slither to your LLM-based agents or workflows, check out our newly released slither-mcp: https://github.com/... [image]
Important new post on Red. Anthropic & MATS Fellows found frontier models are capable (in simulation) of exploiting smart contracts in a meaningful way. We think it's important the world knows. And prepares. (This is nifty research; you should read it) [image]
The next logical step: Anthropic spend $1738 per exploit, submit their findings to a competition, burn a gazillion of tokens in discussions with fellow SRs, get $0.01 in rewards for the missing fee recipient validation, repeat
Really interesting approach to benchmarking. There's really no better place than on-chain to prove AI's merit. Sherlock AI has already “discovered” 100x more exploitable TVL in Anthropic's post-March 2025 timeframe ($350,000 vs. $3,476) Sherlock AI > Opus 4.5 > GPT-5
cool to see an independent, high-profile team explore the same core idea and reach essentially the same qualitative conclusion: autonomous exploit agents are here and economically meaningful (~$4.6M simulated value). we saw similar behavior earlier this year with A1 (July 2025) […