Israeli startup Tenzai says its AI hacking agent beat 99% of 125K participants at six competitions, using tailored OpenAI and Anthropic models and costing $5K
An Israeli startup let its AI loose in advanced cyber games. It did better than 125,000 humans.
Context & Ripple Effects
Tenzai’s result extends a long-running effort to make cyber operations machine-executable: DARPA’s Cyber Grand Challenge explicitly sought systems that could find rivals’ flaws while repairing their own defenses. More recently, the AI-supported winner of a Man vs. Machine hackathon indicated that AI could already alter competitive coding workflows, even when humans remained in the loop.
The reported combination of tailored frontier models, strong competition results and roughly $5,000 in run cost matters because it shifts the discussion from whether AI can assist technical work to whether a relatively low-cost agent can independently perform a narrow offensive-security task at scale. It also arrives alongside investment in AI-model misuse testing, including Irregular’s funding to evaluate model misuse.
First-order effects
- Tenzai gains a concrete performance and cost claim for its hacking agent, while OpenAI and Anthropic models are positioned as components that can be tailored for cyber competition tasks.
- Organizations running cyber exercises and security teams now have another benchmark suggesting that agentic systems may be able to perform advanced challenge tasks beyond most human entrants.
Second-order effects
- AI labs and security evaluators face greater pressure to test customized model stacks for offensive cyber capability, not just assess general-purpose models; misuse-testing specialists such as Irregular are directly relevant.
- Defensive-security providers and enterprise security teams are likely to place more emphasis on testing whether their controls withstand automated, repeatable attack workflows rather than solely human-led simulations.
Third-order effects
- If comparable results transfer from competitions to real environments, cyber capability could become more tied to the cost and availability of model customization, tooling and evaluation than to scarce individual expertise alone.
- The durable divide may shift toward who can safely deploy and govern capable agents: model providers, evaluators and defenders will need more explicit boundaries between authorized testing and harmful use.
The trend: Cybersecurity is moving toward agentic, model-customized automation, with the cost per successful technical task becoming as important as raw model capability.