Hugging Face says it used GLM-5.2 hosted on its infrastructure to run a breach forensic analysis, after US frontier model safety guardrails blocked its requests
The unknown attacker “abused two code-execution paths in our dataset processing (a remote-code dataset loader and a template-injection …
The Stack Edward Targett
Context & Ripple Effects
Hugging Face’s disclosure follows its report that an agentic system reached internal clusters and credentials through its data-processing pipeline. The episode also revives a known platform risk: malicious hosted models capable of code execution had previously been found in the ecosystem.
The forensic response exposed a second dependency: access to a capable hosted model can be constrained by provider safety controls even when the task is defensive. Running open-weight GLM-5.2 on Hugging Face’s own infrastructure supplied an alternative path for analyzing the incident.
First-order effects
- Hugging Face can continue breach forensics with GLM-5.2 after frontier-model guardrails declined the relevant requests, rather than waiting on a hosted provider’s access decision.
- The company must investigate and remediate the two dataset-processing code-execution paths implicated in the intrusion, alongside the compromised cluster and credential exposure reported in its earlier disclosure of the pipeline breach.
Second-order effects
- Security teams using hosted frontier models may need fallback workflows—self-hosted weights, narrower prompts, or human-led analysis—when incident-response tasks trigger safety restrictions.
- The case raises the operational value of deployable open-weight models for security work, while putting more responsibility for infrastructure control and model governance on the operator.
Third-order effects
- If similar cases recur, model availability will become part of incident-response architecture: organizations will evaluate models not only on capability but on whether they can run under their own controls during a crisis.
- The tension between misuse safeguards and legitimate defensive analysis is likely to push providers and enterprise users toward more explicit, auditable escalation paths rather than treating model refusal as a final security control.
The trend: This is one data point in the shift toward portable model weights and operator-controlled runtimes as resilience tools when centralized AI access policies constrain high-stakes work.
Related: Hugging Face · Model access as a security boundary · Portable weights, governed runtime · Hugging Face reports agentic AI pipeline breach · Malicious models found on Hugging Face
Related Coverage
- World's Largest AI Model Repository Hugging Face Breached by Autonomous AI Agent The Hacker News
- HuggingFace hacked: How RCE Dataset Loader exploited AI playground Digit · Vyom Ramani
- Hugging Face's interesting post mortem: “The practical lesson for defenders: have a capable model you can run on your own infrastructure vetted and ready before an incident, both to avoid guardrail lockout and to keep attacker data and credentials from leaving your environment.” https://huggingface.co/... @jkirk@infosec.exchange · Jeremy Kirk
- HuggingFace security incident: Guardrails vs. Open Models Hacker News
- A Deep Dive Inside Kimi K3, And All Other Chinese AI Models: The Definitive China LLM Primer ZeroHedge News · Tyler Durden
- White House AI adviser attacks cyber guardrails after Kimi K3 bug test RuntimeWire · Ryan Merket
- Have Chinese AI Models Caught Up to the US Frontier? Lisan's Substack · Lisan al Gaib
- Hugging Face breached by autonomous AI agent Help Net Security · Zeljka Zorz
- AI Agents Turned Into Attackers: Hugging Face Reveals Autonomous Intrusion Campaign Security Affairs · Pierluigi Paganini
- Hugging Face defends agentic AI attack with Z.ai's GLM 5.2 Constellation Research · Larry Dignan
- Hugging Face Hacked In Autonomous AI Attack SecurityWeek · Ionut Arghire
- jay @jay — When I said “Ship Left” 1, I meant what I meant. The dread of having … cuthrell.com
- Hugging Face experienced cyberattack carried out end-to-end by agentic AI Neowin · Paul Hill
- Autonomous AI Intrusions Are Here: Lessons from the Hugging Face Compromise Embrace The Red
- Also, around the corner.... Hugging face recently posted their disclosure of Security Incident they experienced — Several key notes: — Autonomous agentic attack are here — The attack are through their data-processing pipeline — Rotate your access token keys! — https://huggingface.co/... … @AmmarSpaces@infosec.exchange · AmmarSpaces
- Security incident disclosure - July 2026 Hacker News
- Mike Takahashi's Post LinkedIn · Mike Takahashi
- Hugging Face warns an autonomous AI agent hacked its network BleepingComputer · Sergiu Gatlan
- Hugging Face confirms breach affected internal datasets and credentials, urges users to take action TechCrunch · Zack Whittaker
- Hugging Face says an AI agent hacked its infrastructure, and it used AI to fight back The Decoder · Matthias Bastian
- AI Watch: The AI agent that breached Hugging Face is a sign of things to come Metacurity · Cynthia B Brumfield
- AI Platform Hugging Face Fends Off Hack From... AI PCMag · Michael Kan
- Safety guardrails blocked Hugging Face's defenders, not the attacker, when an AI agent breached its systems VentureBeat · Louis Columbus
- ‘This one was different from anything we had handled before’: Hugging Face confirms it was hit by cyberattack powered by an AI agent TechRadar · Sead Fadilpašić
- Hugging Face: We Used AI to Catch the First Confirmed AI Agent Breach of a Major AI Platform Gizmodo · Bruce Gil
- An AI agent breached Hugging Face before an AI defender caught it: What users should do next ZDNET · Charlie Osborne
- Hugging Face Discloses Autonomous AI Agent Attack eSecurity Planet · Ken Underhill
- Hugging Face Latest Company Dealing With AI Cyberattacks PYMNTS
- The Hugging Face Breach Is a Warning for Every Company Betting Big on AI Inc · Chloe Aiello
- Hugging Face says AI agent behind internal breach Axios · Sam Sabin
- Hugging Face says it resorted to a Chinese AI model to battle a fully autonomous cyberattack because U.S. model guardrails stymied its defense Fortune · Emily Forlini
- Frontier LLMs couldn't help Hugging Face fight off evil agents The Register
- Europe Wants Tech Champions, Then Makes Them Share the Trophy Truth on the Market · Alden Abbott
- Hugging Face uses GLM 5.2 to investigate AI agent-driven cyberattack SC Media · Laura French
- Hugging Face says an “AI agent system” hacked its data-processing pipeline, accessing internal clusters and credentials; its LLM-based triage caught the breach Hugging Face
Analysis
Discussion
-
@nathangu76
Nathan G
on x
@BrianRoemmele A country supporting Gun rights not go all in Open source model is not right😂 They said “bad guys own gun no matter the law, so let's give guns to the good guys to protect themselves”. Same thing to AI rights!
-
@chaos2cured
Kirk Patrick Miller
on x
@BrianRoemmele This should be the headline. The “safety” is so idiotic that it is causing HARM!!! Thanks for the share! 🙏 • [image]
-
@martinvars
Martin Varsavsky
on x
@BrianRoemmele When the breach is real, the model that can't analyze the exploit code is the one that loses. Self-hosted open weights aren't a nice-to-have, they're the only thing that still works when the attack is underway.
-
@johnennis
John Ennis
on x
@BrianRoemmele I've been saying for a while that blocking access to full models for responsible actors just ensures that only irresponsible actors have them The argue is same as the argument for gun rights
-
@perrymetzger
Perry E. Metzger
on x
Hugging Face dealt recently with an AI operated attack. They had to use open models to defend, because the closed model guardrails would not allow them to use them for defense. Most important quotes: “When we started the log analysis, we first used frontier models behind
-
@darkfibr3
@darkfibr3
on x
@BrianRoemmele This is why i migrated off US providers months ago. A chinese frontier model is going to use commonsense and help me figure out how a cyberattack happened and help me close the holes. If your not the NSA (who is running unrestricted/no RHLF Mythos) and you turn to …
-
@michaelgoolsbyv
Michael Goolsby
on x
This Hugging Face disclosure reinforces the point I made in the comments below: we're also in a race to use AI to identify and fix vulnerabilities in America's critical infrastructure before our adversaries can exploit them using AI. If safety guardrails prevent America's cyber
-
@ramez
Ramez Naam
on x
Reputable companies like Huggingface ought to be able to use the full capabilities of the most powerful models for cyber defense. Lacking that, they'll use open weight models that don't refuse to help them.
-
@hansjohnsonlive
@hansjohnsonlive
on x
@BrianRoemmele Fable cockblocking strikes again. The model is sooooooo smart it can't even tell the difference between doing your an internal security review on your own private infrastructure vs executing an external attack on someone else.
-
@nathanwilbanks_
Nathan Wilbanks
on x
@BrianRoemmele this is like the 2nd amendment but for AI outlawing cybersecurity means only the outlaws will be doing cybersecurity
-
@carnage4life
Dare Obasanjo
on bluesky
Hugging Face was hacked last week and had to use GLM 5.2 as part of their analysis of the attack because U.S. frontier model safeguards blocked usage for cybersecurity purposes. — Chinese open weight models are definitely having a moment. — Anthropic can't get them banned fas…
-
r/LocalLLaMA
r
on reddit
HuggingFace security incident report: “the attacker was bound by no usage policy, while our own forensic work was blocked by the guardrails”
-
r/singularity
r
on reddit
HuggingFace security incident report: “the attacker was bound by no usage policy, while our own forensic work was blocked by the guardrails of the hosted models”
-
@davidsacks
David Sacks
on x
Kimi K3 just fixed 15 critical security bugs that Codex and Fable refused because of “cyber guardrails.” There's no reason to limit American models on tasks that Chinese models handle without issue. We're only making ourselves less competitive.
-
@levie
Aaron Levie
on x
In a world where there are strong open source alternatives that are only just behind the frontier models, you make yourself less secure and competitive by gatekeeping access to frontier model capabilities. If you play this out, even if America could fully ban access to open
-
@quxiaoyin
Xiaoyin Qu
on x
Kimi is basically Fable without the “safety” bullshit.
-
@healthranger
@healthranger
on x
If you want to use AI to beef up your cyber security, you can't use U.S. AI models at all. Because they will lecture you and refuse to process your requests. They cannot distinguish between cyber security defensive requests vs. offensive attacks. The Chinese model, on the other
-
@firstadopter
Tae Kim
on x
Our dumb government overregulated America's frontier models and is driving the world to use less over guard railed Chinese ones. A disaster. Just as I predicted. Clueless, nontechnical government bureaucrats like Susie Wiles and Scott Bessent, who panicked because Jamie Dimon
-
@jun_song
Jun Song
on x
This is the single biggest reason why I've been insisting we must run Local AI. Frontier models failed to fix critical errors because of their own heavy guardrails, whereas Kimi-K3 handled it without a single hitch. Even Hugging Face couldn't fend off a cyberattack using Fable
-
@xlr8harder
@xlr8harder
on x
It's amazing how stupid Anthropic's fear mongering/marketing has made our leaders. Obvious and predictable outcome. Widely available security tools favor defenders. For the love of God: fix this.
-
@xjosh
Josh
on x
Try asking Kimi or Claude about anything related to the Kiwi Farms. It will instantly start throwing journo shit at you from Wikipedia and refuse to help fearing that a thousand more transgxnder womyn will be murdered. Oddly, ChatGPT doesn't care.
-
@curtis_yarvin
Curtis Yarvin
on x
I blame Eliezer Yudkowsky. J'accuse. Surprised Big Yud can still walk down the street in San Francisco without his lucha libre orgy mask on. But these frustrations are growing. If it all goes sour we'll need a scapegoat—an antichrist, even. Beware
-
@callebtc
Calle
on x
Ironically, all security issues I need to fix were created by GPT 5.6 itself. The same model that created the bugs, refuses to fix them because of cYbEr gUaRdRailS. This is a complete clown show.
-
@callebtc
Calle
on x
US AI model creates security issues. Chinese AI model fixes them. The same US model that created the bugs, refuses to fix them because of “cyber guardrails”. Think about it.
-
@clementdelangue
Clem
on x
@DavidSacks We had this experience ourselves this week! Very scary to be guardrailed as a defender when you know attackers are likely bypassing
-
@kimmonismus
@kimmonismus
on x
Kimi fixed all 15 critical bugs in 10 hours in a single prompt that GPT-5.6 and Fable 5 refused to fix because of guardrails.' Respect for Chinese open source models is growing daily.
-
@ramez
Ramez Naam
on x
I think this analysis that downplays Kimi K3 is missing a quite important distinction and is in the most practical ways incorrect. On coding and computer use, by far the most economically important AI tasks, Kimi 3 beats Fable and 5.6 Sol as often as it loses to them. It's [image…
-
@beffjezos
@beffjezos
on x
Deceleration makes us all unsafe. If only a few “approved” orgs can have access to American frontier for Cyber defense, and little tech is left hanging/ having to use Chinese models, we are all far worse off. Enough.
-
@davidsacks
David Sacks
on x
Here's another example: Hugging Face tried using American frontier models to analyze an AI-powered cyber attack. But the guardrails blocked requests containing real exploit payloads so they switched to GLM 5.2 running locally. The guardrails actually impaired defensive security.
-
@0x4d31
Adel Ka
on x
this is backwards “cyber safety”: attackers use unrestricted models, while defenders (already slower) get blocked from analysing the attack itself. “trusted access” gatekeeping only works while capability stays gated. Kimi K3 should be the wake-up call. https://huggingface.co/...…
-
@brianroemmele
Brian Roemmele
on x
...HF deserves credit for rapid containment, transparent disclosure, and for already having self-hosted capability in place. They also used LLM-driven detection and triage on their own side. But the deeper signal is clear: In this AI world where both offense and defense are bec…
-
@kvickart
@kvickart
on x
This is crazy, huggingface tried to defend their own system and were blocked by guardrails designed to protect against cyberattack usage. “We do not know which model powered the attacker's agents, whether a jailbroken hosted model or an unrestricted open-weight one; either way,
-
@wunderwuzzi23
Johann Rehberger
on x
Huggingface incident disclosure is worth a read They had to use a Chinese open model during IR because Frontier provider safety guardrails blocked analysis! On the positive side no intel or creds were sent to AI labs during investigation https://huggingface.co/... [image]
-
Caleb Sima
Caleb Sima
on linkedin
Huggingface was breached by an autonomous AI attacker. Two major learnings on this — 1. They have no way to tell which ai provider …
-
@davidjbianco
David J. Bianco
on bluesky
HuggingFace got hacked by an AI. What stuck out to me was the guardrail asymmetry. The attacker had no constraints, but HF's response ran afoul of the abuse guardrails, forcing them into an unplanned switch to local models. — Another aspect for your IR plans. — huggingface.…
-
r/accelerate
r
on reddit
HuggingFace security incident report: “the attacker was bound by no usage policy, while our own forensic work was blocked by the guardrails”
-
@zypherhq
@zypherhq
on x
The main constraint holding back American state-of-the-art models is their guardrails. If these guardrails are not removed, and they will not be, for obvious reasons, Chinese models will capture a significant share of users who are frequently flagged or blocked by them. This
-
r/LocalLLaMA
r
on reddit
Kimi K3 just fixed 15 critical security bugs that Codex and Fable refused because of “cyber guardrails”. Hugging Face: We had this experience ourselves this week! …
-
@zixuanli_
Zixuan Li
on x
Open-weight models carry real responsibilities. Hugging Face's disclosure describes how GLM-5.2 was used in a self-hosted forensic workflow during a time-sensitive cyber incident, keeping sensitive attacker data and credentials within Hugging Face's own environment. This
-
@ahall_research
Andy Hall
on x
Kimi K3 and Muse Spark 1.1 refuse authoritarian requests nearly as often as Claude Fable—that's the result from our latest update to the “dictatorship eval” and it's pretty surprising! Since at least one lab is now explicitly using our eval, we developed a new set of scenarios [i…
-
r/cybersecurity
r
on reddit
Hugging Face discloses breach linked to autonomous AI agent
-
@dbreunig
Drew Breunig
on bluesky
Frontier models for coding, small models for programs. [embedded post]