David Sacks says Dario Amodei refused to “fix the jailbreak or de-deploy the model” after “a highly credible trusted partner” reported a Fable jailbreak
I've had a number of conversations with folks inside and outside government about the current situation with Anthropic, and here is what I believe to be true: — As we know, Anthropic publicly released its Mythos class models earlier this week under the commercial name Fable.
@davidsacksDavid Sacks
Context & Ripple Effects
Related coverage had Anthropic saying red-team tests found no universal jailbreaks and that traffic would remain on Mythos-class models for 30 days. Sacks’s allegation directly challenges whether that testing and deployment posture was adequate.
The surrounding coverage ties the reported jailbreak to US restrictions and later describes Anthropic agreeing to detect and address security risks, making the dispute part of a broader fight over who determines when a model is safe to keep deployed.
First-order effects
Sacks’s claim places Anthropic and Dario Amodei under immediate public and policy scrutiny over their response to a reported Fable jailbreak; the alleged refusal is not independently established in the supplied material.
It sharpens the contrast between Anthropic’s earlier assurance that no universal jailbreaks were found and concerns raised by a reportedly trusted external partner.
Second-order effects
Government restrictions described in related coverage become more consequential: a disputed jailbreak can shift a model-safety issue from a vendor’s internal red-teaming process into an oversight and deployment decision.
Other frontier-model providers face added pressure to document how they triage external jailbreak reports, rather than relying only on published test results or broad safety claims.
Third-order effects
If this pattern persists, deployment governance will increasingly hinge on continuous post-release monitoring and externally reported failures, not solely pre-release evaluations.
The long-term fault line is likely to be who has authority to require mitigation or withdrawal when safety evidence is contested: the model developer, government agencies, or both.
The trend: This is one data point in the shift from voluntary model-safety testing toward contested, ongoing oversight of frontier-model deployment.
tldr: - anthropic asks to be regulated - 3rd party finds jailbreak - USG asks Anthropic to fix it - Dario says no - USG bewildered, puts export control - anthropic pulls model anthropic promotes ai safety as their brand but won't fix a security vulnerability???
If you're ever wondered what AI regulation will look like if implemented, this is de facto what to expect on the current trajectory. If the government has control over what models eventually get released, then companies will spend months and months debating with the government
“The Admin asked Dario to fix the jailbreak or de-deploy the model. Dario refused. — In their blog post, Anthropic defended its decision by saying the jailbreak isn't serious” we need regulation but we will decide how we react to this? You can't have it both ways
Crazy to see previously antiregulatory people suddenly fine with government shutdown of a company's showcase product with zero public transparency. Absent bipartisan and independent scientific oversight, the approach outlined below leaves (a) the government with too much room
Update on the Fable 5 situation: > Fable is basically a powerful “cyberweapon” model with safety locks on it > Someone found a way to pick those locks > The government asked Anthropic to fix it or pull the model > Anthropic said no and called the problem minor > So the
If a Treasurer of a Fortune 1000 company kept all of their cash in one bank they'd be fired for incompetence. Similarly, if the leadership of a Fortune 1000 company bets the farm on only one frontier lab and their models you're taking a lot of risk. This risk compounds as the
Impressive and nuanced reply from David here. I respect it. I hope everything he's saying is true and if so I guess all the alarm overnight on here is maybe not so necessary.
If this is true — big if — it mostly just exposes the need for an independent authority that can determine whether AI models are actually dangerous or not. Stuff this high stakes cannot be decided by an argument between Anthropic and Amazon.
It seems very likely to me that the non-compliance on the part of Anthropic the government perceives and is confused by from a safety-focused company, is because the government doesn't understand that what they're asking for is stupid and/or impossible.
Anthropic called Mythos dangerous in its own safety statement. That statement is now the reason Fable 5 got banned by the US gov. Surprisingly, “Dario refused.” [image]
As a standalone issue, this has all the trappings of a free speech issue, not just a national security one. Everyone's access to an expressive tool was shut down by a letter from a government official.
“The Admin asked Dario to fix the jailbreak or de-deploy the model. Dario refused. — In their blog post, Anthropic defended its decision by saying the jailbreak isn't serious. ” This is crazy. What are we even doing here?
Anthropic developed what they defined as a cyber weapon, Mythos. Then they released Mythos to the world as “Fable.” The same model, but with guardrails to allow its commercial deployment. When the guardrails failed, they rejected calls to patch the vulnerabilities. Dario is
Interesting: According to David Sacks' opinion, the fault lies with Anthropic (specifically CEO Dario Amodei). He argues that: • Anthropic released Fable (Mythos with guardrails) but refused the U.S. government's reasonable request to fix a confirmed jailbreak that could
The Admin's hope now is that Anthropic remediates the safety issue, the export control is lifted, and Fable goes back into general release. The Admin wants all of this to happen as soon as possible. It is frankly bewildered that Anthropic hasn't wanted to comply with safety
Note: if users know in advance that they are being monitored during a test period, and they opt into that test, that's a private company's right — obviously. And this is actually standard practice in many fields. User can opt out of the testing period and just use the last
David Sacks on How Anthropic is Ironically Running Surveillance on Their Latest Models “This is the company that said that it was against government surveillance. They are now retaining for 30 days every prompt and every output you send to one of these Mythos class models.” [vide…
is there more to the story here than just export controls as retaliation? or is this just part of the a.i. hype cycle, creating a national security incident to fabulize the world-changing power of a.i. and the advancement of the u.s. in that arena? if our govt weren't crooks, w…
Here is the gov's argument. Hard to evaluate whether the concern is justified—or just politics toward an American company this government often whacks—without the specifics of the jailbreak. Also not clear why they chose to regulate via a blunt and hard-to-follow export control.
David Sacks claims that Anthropic was alerted to a jailbreak exploit that the USG asked them to patch and Dario Amodei refused, leading the USG to ban access to the Mythos class models. Why refuse to patch the exploit?
This post by @DavidSacks about the Anthropic/Fable situation is noteworthy, but not for the reason most people think. Put the details aside for a second. Anthropic released a blog post with their side of the story a few hours ago. David is responding with a bullet point list in