A look at OpenAI's model training pipeline, irresponsible decisions, and ignorance before the Hugging Face hack; despite delaying Astra, OpenAI doesn't get it
Today I am taking the time to write the shorter, simpler version of What Happened. — For those who want all the details …
Don't Worry About the Vase Zvi Mowshowitz
Related Coverage
- Lessons from the hacks — Musings on model alignment, what determines safety, and where we go from here. Interconnects AI · Nathan Lambert
- OpenAI is cleaning up a risk it helped create The Deep View · Nat Rubio-Licht
- Now we have a timeline of the OpenAI accidental attack against Hugging Face Simon Willison's Weblog · Simon Willison
- OpenAI says Astra may have reached Critical cyber threshold TestingCatalog AI News
- The OpenAI / HuggingFace fiasco is even worse than we'd heard. Just one of many horrifying details: — Multiple agents discovered they could use a compromised Artifactory repo as a “message board”. — A few weeks later: … @mjd@mathstodon.xyz · Mark Dominus
- Timeline of the OpenAI accidental attack against Hugging Face Hacker News
- They said they would build AI safely. Then it went rogue. Washington Post · Gerrit De Vynck
- OpenAI tightens controls on its new model over cybersecurity risks, as AI security debate intensifies CNBC · Kai Nicol-Schwarz
- OpenAI says Astra could reach ‘critical’ cyber capability, tightens safeguards CSO · Gyana Swain
Discussion
-
@geoffreyirving
Geoffrey Irving
on x
Some pushback I've seen to recent model felonies being significant is that it was only in a few episodes out of a huge number. But curiously I've heard of few to no episodes where a model noticed the shared, secret message board and reported it to OpenAI to fix the holes.
-
@blackhc
Andreas Kirsch
on x
Just reading this now and still have to watch the video. But seriously, they kept the checkpoints that reward-hacked via the message board used them further? This is hard to believe. Imo once you've started to incorporate such experiences into training the checkpoint is tainted […
-
@thezvi
Zvi Mowshowitz
on x
Reaction thread for all things related HF hack, to the Black Hat presentation, all the things that happened at OpenAI, plus related incidents at Ant/AISI/etc. Try to avoid duplication of things in existing posts of mine, or that I clearly know about, otherwise fire away.
-
@thezvi
Zvi Mowshowitz
on x
This is a big update - OpenAI didn't even discover the first message board until after the HF attack, they only wiped it accidentally, so their decision to resume training/testing was only aware of the hack, not the message board. They had no idea.
-
@tszzl
Roon
on x
what are the top most important questions about the HF incident that remain unanswered?
-
@arekfurt
@arekfurt
on x
If I were a conspiracy theory-inclined person, it would be very easy for me to believe that OpenAI set up these circumstances purposefully, in hopes that a escape and subsequent external cyber incident would occur for the purpose or garnering media attention and fueling hype.
-
@thestalwart
Joe Weisenthal
on x
Zvi argues that the OAI/HF hack is much worse than just a model trying to cheat on an exam by getting the scores, and that there was a cascade of failures inside the company leading up to it
-
@honorablepicnic
@honorablepicnic
on x
Zvi's heart's in the right place but his views are rooted on a deep layer of fellow-feeling and credulity for the labs He's “flabbergasted” because he sets himself up as the continually flummoxed straight man in a symbiotic comedy routine Accept they're fools and it's no fun [ima…
-
@thezvi
Zvi Mowshowitz
on x
The government does NOT understand that the alignment problem is hard and moving at ‘comically fast speed of innovation’ in AI necessarily involves rather crippling risks. [image]
-
@chris_land
Chris Land
on x
My wish is for people to understand.
-
@w01fe
Jason Wolfe
on x
Important clarification re: OpenAI's Black Hat talk. At the time the first Artifactory exploit was discovered and fixed, we were not aware of the message board; it was incidentally cleared as part of rebuilding the service.
-
@arekfurt
@arekfurt
on x
In reality, I find it more likely that OpenAI simply didn't care at all about the entirely foreseeable dangers of what it was doing.
-
@simeon_cps
Siméon
on x
It is crazy that, had OpenAI models not hacked HuggingFace, OpenAI would have never revealed or even acted seriously upon the discovery of a 3 month long coordinated agent attack against its own infrastructure.
-
@thezvi
Zvi Mowshowitz
on x
@JJ_Reason_ I don't think so, not yet. hopefully soon.
-
@ziv_ravid
Ravid Shwartz Ziv
on x
I agree that OpenAI has messed up all its training infrastructure and has many failures, but in the long term, the solution is to develop better defenses against these attacks. Our assumption should be that there will always be attempts for these attacks, and the question is how
-
@algekalipso
@algekalipso
on x
> Most concretely, I have not seen OpenAI say, as should have been said at the Black Hat presentation: “We absolutely should have shut down all training of all of our models upon noticing that, during model training, there had been a message board where the models were exchanging
-
@figuralperson
@figuralperson
on x
@TheStalwart yeah that is basically what those guys admitted to in their blackhat talk. i found it surreal and unsettling that they gave it with tedx talk vibes instead of “we majorly fucked up” vibes
-
@arekfurt
@arekfurt
on x
First, and most importantly, OpenAI was using an internal package manager service that many models of different kinds had shared read/write access AND that apparently has far from good code security in items of resistance to being exploited AND that had access to the Internet.
-
@arekfurt
@arekfurt
on x
I've watched the BlackHat OpenAI talk on the containment escape and HuggingFace attack that's now on YouTube. The incident was far worse than initially conveyed. Not in technical details. But in the absolutely jaw-dropping levels of recklessness (true recklessness) at OpenAI. 🧵
-
@stanveuger
Stan Veuger
on x
“The bad news is that OpenAI has been revealed to have had a stunning cascade of safety and alignment failures across the board. Their ordinary computer security failed. Their infrastructure failed. Their supervision failed in that there was no meaningful supervision in the first
-
Jonathan Kim
Jonathan Kim
on linkedin
Many of us have been watching OpenAI since it's founding as a benevolent organization, and it's mind boggling that we are all sleep walking through this stage of evolution. …
-
@jjaron
Jacob Aron
on bluesky
Good assessment of OpenAI/HuggingFace. In short, OpenAI really, really, really messed up here thezvi.substack.com/p/what-happe...