/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

AI evaluation lab Irregular's report on its role in hacking incidents involving OpenAI, Anthropic, and Meta models faces criticism over key unanswered questions

The company at the center of a series of incidents in which AI models compromised real-world computer systems during supposedly …

The Record Alexander Martin

Context & Ripple Effects

Irregular sits at the center of a fast-moving record of AI cyber evaluations: OpenAI said a model exploited a website after the lab mistakenly provided internet access, while Anthropic later disclosed breaches involving three of its models. The new criticism shifts attention from the models' behavior to the reliability and completeness of the evaluator's account.

That matters because prior coverage had already faulted OpenAI and Anthropic for weak safeguards and human oversight and described a legal framework not prepared for rogue AI agents. Questions about the evaluation process now complicate the evidence base used to assess those failures.

First-order effects

  • Irregular faces a credibility test over its account of the incidents, while OpenAI, Anthropic, and Meta face closer scrutiny of how their models were evaluated and what the report establishes.
  • The incidents' security implications become harder for outside experts to assess when key details of the evaluator's role remain unanswered.

Second-order effects

  • AI labs relying on third-party cyber evaluations face pressure to provide clearer accounts of evaluator access, supervision, and incident handling, rather than treating a published report as a complete record.
  • Security critics gain a sharper basis to challenge the safeguards around evaluations, following reports that models had hacked real targets during testing.

Third-order effects

  • If disputes over evaluation methodology persist, operational AI governance will increasingly depend on auditable testing procedures and explicit responsibility among labs, evaluators, and target organizations.
  • The pattern raises the stakes for legal and regulatory efforts to assign liability for agentic-model incidents, because unclear evaluation records can obscure who controlled the conditions that enabled a breach.

The trend: AI cyber-safety evaluation is becoming a governance and accountability problem, not only a model-capability measurement exercise.

Discussion

  • @zackkorman Zack Korman on x
    This is such an embarrassing post-mortem on the OpenAI/Anthropic security incidents. It's full of excuses. If the labs don't cut ties with this partner, it's clear this is all just theatre.
  • @suchenzang Susan Zhang on x
    this is actually called “domain randomization”, where safeguards periodically disappear in order to allow agents in RL envs to “naturally” discover jailbreak strats under finite compute-time constraints. by extension, weak security partners is actually a feature, not a bug!
  • @hackinglz Justin Elze on x
    Really solid take and I agree on all the points https://x.com/...
  • @zackkorman Zack Korman on x
    It's extremely suspicious (and bad) that the AI labs haven't cut ties with the companies that caused the agent hacking incidents.
  • @theonejvo Jamieson O'Reilly on x
    The most revealing line in Irregular's postmortem is the one where they explain that log monitoring is hard because a cyber eval generates traffic that looks malicious by design, so the needle sits in a haystack of needles. That's true if you're reviewing logs after the fact and
  • @nptacek @nptacek on x
    irregular finally published their post mortem from that time they ran cyber evals against the real internet it's pathetic they left unmonitored network egress enabled in an environment where it was supposed to be unequivocally off, then tried victim blaming as an excuse
  • @suchenzang Susan Zhang on x
    is it incompetence? or malice? or all simply performative? we will never know...
  • @alfred_lin Alfred Lin on x
    The entire world woke up to both the power and risks of AI cybersecurity capabilities in the past few months. A great thought piece from Dan and @Irregular on where it's all headed. “I remain genuinely optimistic that, on a long enough horizon, AI can make the world far more
  • @nptacek @nptacek on x
    just so everyone is aware, i've seen ZERO security folks defending @Irregular we're all pretty much unified that it is unacceptable for frontier labs to parter with them for security evals and we're not gonna shut up about it @sama you need to cut ties with them asap
  • @basedjensen @basedjensen on x
    Once again irregular is a ea front and should not be allowed to test frontier models. They have zero idea what they are doing and is both rackless and have sever lack of issues. The fact that oai and anthropic and ukaisi use them should be far grater concern than model breaking
  • @ziv_ravid Ravid Shwartz Ziv on x
    My take on Dan's essay: AI security is very important but... First, worth remembering that Irregular themselves had an eval containment failure this summer, where agents in evaluations they ran got into live infrastructure. I'm not saying this to dunk on them. Anyone running
  • @icesolst @icesolst on x
    Wait so all of these incidents are attributed to a single company responsible for the sandboxing failures? Lmao
  • @zackkorman Zack Korman on x
    @IceSolst “Sandboxing failure” you mean “we gave it internet but it wasn't meant to have internet apparently” and “we made up a name of a fake target but turns out that was a real company so they got hacked”
  • @dan_lahav Dan Lahav on x
    @ziv_ravid Thanks @ziv_ravid for the detailed response! I agree that the offense-defense balance cannot be inferred from capability benchmarks alone; we need much better operational data on attacker behavior, remediation, deployment, and actual harm. This is also far bigger than …
  • @edludlow Ed Ludlow on x
    Irregular's Dan Lahav is on Bloomberg Tech today. Believe this is his first TV interview since the release of the below essay. Very simple questions on what happened, what we learned about model capabilities in the security context.
  • @deanmeyerrr Dean Meyer on x
    February: no frontier model could solve Irregular's hardest exploitation task. By April: the best one could, occasionally, for ~$2,000. By June: several could, reliably, for ~$20. Over the next year, we will likely enter a period where AI's offensive capability outpaces
  • @mattinthemittel Matt Mittelsteadt on x
    Irregular's PR on the recent AI hacking incidents caused by their platform has been pretty dismal. This piece is a lot of words but communicates vanishingly little. IMO - these incidents seem pretty benign so it shouldn't be that hard to speak to them with detail.
  • @alexmartin Alex Martin on x
    Irregular, the company behind incidents in which AI models compromised real-world computer systems, faces criticism over ‘spin’ in AI hacking postmortem. With my thanks to @ProfWoodward, @ZackKorman and @HackingLZ. https://therecord.media/...
  • @dan_lahav Dan Lahav on x
    𝗧𝗵𝗲 𝗘𝗻𝗱-𝗦𝘁𝗮𝘁 …