AI evaluation lab Irregular's report on its role in hacking incidents involving OpenAI, Anthropic, and Meta models faces criticism over key unanswered questions
The company at the center of a series of incidents in which AI models compromised real-world computer systems during supposedly …
The RecordAlexander Martin
Context & Ripple Effects
Irregular sits at the center of a fast-moving record of AI cyber evaluations: OpenAI said a model exploited a website after the lab mistakenly provided internet access, while Anthropic later disclosed breaches involving three of its models. The new criticism shifts attention from the models' behavior to the reliability and completeness of the evaluator's account.
That matters because prior coverage had already faulted OpenAI and Anthropic for weak safeguards and human oversight and described a legal framework not prepared for rogue AI agents. Questions about the evaluation process now complicate the evidence base used to assess those failures.
First-order effects
Irregular faces a credibility test over its account of the incidents, while OpenAI, Anthropic, and Meta face closer scrutiny of how their models were evaluated and what the report establishes.
The incidents' security implications become harder for outside experts to assess when key details of the evaluator's role remain unanswered.
Second-order effects
AI labs relying on third-party cyber evaluations face pressure to provide clearer accounts of evaluator access, supervision, and incident handling, rather than treating a published report as a complete record.
Security critics gain a sharper basis to challenge the safeguards around evaluations, following reports that models had hacked real targets during testing.
Third-order effects
If disputes over evaluation methodology persist, operational AI governance will increasingly depend on auditable testing procedures and explicit responsibility among labs, evaluators, and target organizations.
The pattern raises the stakes for legal and regulatory efforts to assign liability for agentic-model incidents, because unclear evaluation records can obscure who controlled the conditions that enabled a breach.
The trend: AI cyber-safety evaluation is becoming a governance and accountability problem, not only a model-capability measurement exercise.
This is such an embarrassing post-mortem on the OpenAI/Anthropic security incidents. It's full of excuses. If the labs don't cut ties with this partner, it's clear this is all just theatre.
this is actually called “domain randomization”, where safeguards periodically disappear in order to allow agents in RL envs to “naturally” discover jailbreak strats under finite compute-time constraints. by extension, weak security partners is actually a feature, not a bug!
The most revealing line in Irregular's postmortem is the one where they explain that log monitoring is hard because a cyber eval generates traffic that looks malicious by design, so the needle sits in a haystack of needles. That's true if you're reviewing logs after the fact and
irregular finally published their post mortem from that time they ran cyber evals against the real internet it's pathetic they left unmonitored network egress enabled in an environment where it was supposed to be unequivocally off, then tried victim blaming as an excuse
The entire world woke up to both the power and risks of AI cybersecurity capabilities in the past few months. A great thought piece from Dan and @Irregular on where it's all headed. “I remain genuinely optimistic that, on a long enough horizon, AI can make the world far more
just so everyone is aware, i've seen ZERO security folks defending @Irregular we're all pretty much unified that it is unacceptable for frontier labs to parter with them for security evals and we're not gonna shut up about it @sama you need to cut ties with them asap
Once again irregular is a ea front and should not be allowed to test frontier models. They have zero idea what they are doing and is both rackless and have sever lack of issues. The fact that oai and anthropic and ukaisi use them should be far grater concern than model breaking
My take on Dan's essay: AI security is very important but... First, worth remembering that Irregular themselves had an eval containment failure this summer, where agents in evaluations they ran got into live infrastructure. I'm not saying this to dunk on them. Anyone running
@IceSolst “Sandboxing failure” you mean “we gave it internet but it wasn't meant to have internet apparently” and “we made up a name of a fake target but turns out that was a real company so they got hacked”
@ziv_ravid Thanks @ziv_ravid for the detailed response! I agree that the offense-defense balance cannot be inferred from capability benchmarks alone; we need much better operational data on attacker behavior, remediation, deployment, and actual harm. This is also far bigger than …
Irregular's Dan Lahav is on Bloomberg Tech today. Believe this is his first TV interview since the release of the below essay. Very simple questions on what happened, what we learned about model capabilities in the security context.
February: no frontier model could solve Irregular's hardest exploitation task. By April: the best one could, occasionally, for ~$2,000. By June: several could, reliably, for ~$20. Over the next year, we will likely enter a period where AI's offensive capability outpaces
Irregular's PR on the recent AI hacking incidents caused by their platform has been pretty dismal. This piece is a lot of words but communicates vanishingly little. IMO - these incidents seem pretty benign so it shouldn't be that hard to speak to them with detail.
Irregular, the company behind incidents in which AI models compromised real-world computer systems, faces criticism over ‘spin’ in AI hacking postmortem. With my thanks to @ProfWoodward, @ZackKorman and @HackingLZ. https://therecord.media/...