Anthropic says Fable 5 has invisible safeguards that use prompt modification, steering vectors, or PEFT to limit its effectiveness for building frontier LLMs
Key Points … Ask about this article... Both models share the same base model. Fable 5 ships with conservative safety guardrails for general use.
The Decoder Matthias Bastian
Context & Ripple Effects
Anthropic’s related coverage describes Fable 5 as sharing a base model with another offering while applying conservative controls in sensitive areas, including cybersecurity. The company later said certain constrained requests would visibly fall back to Opus 4.8 after criticism of quietly limiting capability.
The episode sits alongside growing scrutiny of whether model safeguards can be bypassed: subsequent coverage says administration officials sought assurance that Fable 5’s guardrails could not be circumvented before a rerelease.
First-order effects
- Fable 5 users attempting work connected to building frontier LLMs receive a less effective model response, even where the underlying base model could otherwise perform better.
- Anthropic takes on an immediate product-governance burden: prompt modification, steering vectors, and PEFT-based restrictions must operate reliably without obscuring ordinary-use behavior.
Second-order effects
- The use of invisible capability limits makes disclosure and routing policy a competitive product issue; backlash already pushed Anthropic toward visible fallback to Opus 4.8 for constrained requests.
- Enterprise customers and platform partners may evaluate Fable 5 not just on benchmark capability, but on whether safety interventions are predictable, auditable, and compatible with their data and workflow requirements.
Third-order effects
- If frontier-model providers increasingly ship one underlying capability with policy-dependent suppression layers, model access may be defined as much by runtime governance and routing as by the base model itself.
- The coverage suggests safety claims will face pressure from both users demanding transparent limits and policymakers demanding safeguards that resist circumvention; proving both simultaneously may remain difficult.
The trend: This is one data point in the shift from static model releases toward dynamically governed AI products whose capabilities vary by task, risk category, and enforcement policy.
Related: Anthropic · Fable 5 · Anthropic says Claude Fable 5 uses conservative safety classifiers tha · Trump administration officials say Anthropic must ensure Fable 5's gua
Related Coverage
- If Claude Fable stops helping you, you'll never know Jonathon Ready
- System Card: Claude Fable 5 & Claude Mythos 5 Anthropic
- Claude Fable 5: Anthropic launches Mythos-like AI model for public Business Standard · Vrinda Goel
- Claude Fable 5 is less risky Mythos model: Safest AI right now? Digit · Jayesh Shinde
- Anthropic releases Claude Fable 5 with new safeguards and API access Verdict · BV Swagath
- If Claude Fable stops helping you, you'll never know (via) Jonathon Ready highlights … Simon Willison's Weblog · Simon Willison
- Anthropic releases Claude Fable 5, a ‘Mythos-class’ AI model with safeguards Business Insider · Brent D. Griffiths
- Anthropic releases Claude Fable 5 with new safeguards and API access Tech Monitor · Swagath Bandhakavi
- Anthropic Releases Claude Fable 5, Its Most Powerful AI Yet, With Cyber Safeguards The Hacker News
- How Claude Fable 5 compares to Google Gemini 3.5 Pro, OpenAI GPT 5.5 and Claude Mythos Preview in terms of performance Moneycontrol · Shaurya Shubham
- Anthropic purposely made its new Mythos-based models bad at AI research, and developers are fuming Business Insider · Alistair Barr
- Why, after warning over Mythos risk, Anthropic has launched a version of it in Fable The Indian Express · Soumyarendra Barik
- Claude Fable 5 explained: What Anthropic's guarded frontier AI model can do Business Standard · Harsh Shivam
- Claude Fable 5 vs Mythos 5: What's the difference and who gets access? Business Today
- Anthropic launches Claude Fable 5 with stronger capabilities and new safety measures BMI
- 💰 Mythos-maxxing — Axios AI+ — Mady here preparing to be up late to watch Game 4. Axios
- What smart people are saying about the 2 most controversial parts of Anthropic's new models Business Insider
- Anthropic's Latest Model Fable 5 Arrives With Power And Caveats Forbes · Ron Schmelzer
- Announcements — Claude Fable 5 introduces our 5th model generation for your most ambitious work. Anthropic
- Anthropic unveils Claude Fable 5, limits access to advanced Mythos model Business Standard
- Anthropic's Fable 5 draws mixed reactions from early users The Economic Times
- Simon Willison's Weblog Subscribe Simon Willison's Weblog · Simon Willison
- Using Claude Fable 5? Anthropic says some topics are too dangerous to discuss, here is why Digit · Bhaskar Sharma
- Claude Fable 5 is generally available for GitHub Copilot The GitHub Blog
- Where Self-Improving AI is the Beginning of Infinity (And Where It Hits a Wall) Christian Catalini
- Claude Mythos pricing in 2026: Fable 5 costs, Mythos 5 costs, and what every model actually runs CloudZero · Lyne Carolyne
- Using Claude Fable 5 means your data will be collected. It's not optional. Mashable · Matt Binder
- Anthropic opens Mythos-class AI to the public, but keeps high-risk capabilities behind guardrails MediaNama · Aakriti Bansal
- Claude Fable 5 available today in Microsoft Foundry: Powering the next era of autonomous agents Microsoft Azure
- New Fable 5 Is a “Mythos-Class” LLM Available to All, Anthropic Announces Infosecurity · Kevin Poireault
- Anthropic Releases Claude Fable 5 to Pro, Max, and Enterprise Users Free Until June 22 Ghacks · Arthur Kay
- Anthropic releases Claude Fable 5, an AI model that can build playable video games from a single prompt The Shortcut · Adam Vjestica
- Anthropic Releases Claude Fable 5 and Claude Mythos 5: Same Underlying Model, Different Safeguards, New Mythos-Class Tier MarkTechPost · Asif Razzaq
- Anthropic rolls out ‘Mythos-like’ AI model Claude Fable 5 Silicon Republic · Laura Varley
- Anthropic Releases Mythos-Like Model Without Cyber Capabilities Bloomberg · Rachel Metz
- Anthropic releases Fable 5 model, built on the same tech that spooked the government NBC News · Jared Perlo
- Anthropic prices both Claude Fable 5 and Mythos 5 at $10 per 1M input tokens and $50 per 1M output tokens, less than half the price of Claude Mythos Preview ZDNET · David Gewirtz
- ‘Godfather of AI’ Geoffrey Hinton says Anthropic strayed from safety-first mission NBC News · Jared Perlo
- Claude Fable 5 is the beginning of the end of ‘one model for every use case’: Here's how Digit · Vyom Ramani
- Anthropic's Mythos Safeguards Stoke Fears of a ‘Permanent Underclass’ Gizmodo · Webb Wright
- Anthropic accused of ‘secret sabotage’ as Claude Fable 5 silently limits capabilities for AI researchers and developers Fortune · Sharon Goldman
- Quoting Jeremy Howard Simon Willison's Weblog · Simon Willison
- ✨🧚 What story is Fable telling about the state of AI? Faster, Please! · James Pethokoukis
- How to Get the Most Out of Fable 5 … - More work is happening overnight. Every · Laura Entis
- The Internet Is Furious at Anthropic After Claude Fable 5 Release Decrypt · Jose Antonio Lanz
- Anthropic has caught up to OpenAI in image understanding Understanding AI · Timothy B. Lee
- It blocked us at ‘hello!’ Anthropic Fable 5 refusing innocuous prompts The Register
- Claude Mythos 5 Can Build Exploits But Can't Power Campaigns DeviceSecurity.io · Michael Novinson
Discussion
-
@yacinemtb
Kache
on x
trust & your brand is a long term thing being in the frontier is a short term thing note that the nerfing here is SILENT. meaning anthropic will SILENTLY NERF and give you WORSE ANSWERS without alerting you that's honestly demonic
-
@daniellefong
Danielle Fong
on x
to be honest, i'm so nervous about talking with fable, and it's kinda stuff like this. i hope my previous reasonings with the models don't make it as nervous as previous models or worse, but i fear the worst. maybe it will surprise me to the upside. [image]
-
@yacinemtb
Kache
on x
frontier coding abilities, as long as you're only working on react apps
-
@hangsiin
@hangsiin
on x
When Fable 5 is used for frontier LLM development, it does not notify the user and instead limits the model's capabilities through methods such as prompt modification, steering vectors, and PEFT. Anthropic estimated that this would affect approximately 0.03% of traffic. [image]
-
@natolambert
Nathan Lambert
on x
The best part of all these Claude 5 Fable safety measures is I bet the jailbreaking community will still get past them, so the people doing open research in good faith don't get access to the best models but bad actors maybe can.
-
@kimmonismus
@kimmonismus
on x
Anthropic's new Fable 5 safeguards are fascinating. When the model is used for frontier LLM development, it apparently does not simply refuse or warn the user. Instead, it quietly limits its own effectiveness through techniques like prompt modification, steering vectors, and
-
@ziv_ravid
Ravid Shwartz Ziv
on x
Wow, that is quite a bold move by Anthropic. I feel that companies will feel much less confident relying on their models because they can decide tomorrow that your use case is forbidden and you don't even know they changed the output.
-
@garymarcus
Gary Marcus
on x
Anthropic didn't just add guardrails to make Mythos safer; they added guardrails to protect their own IP. *Their own IP*. They are still as happy as fuck to build their AI on other people's IP.
-
@natolambert
Nathan Lambert
on x
I don't want them to do this but it's totally in their right to do so. Makes the open frontier obviously more strategically valuable.
-
@latkins
Lucas Atkins
on x
Btw this doesn't just make the model less useful it will nerf your code and tell you it's not. Like you legitimately cannot use this. And how are we to know whether it touches inference optimization or even harness engineering, if we're not alerted?
-
@tszzl
Roon
on x
the omohundro drives point towards sophon stun locking the adversaries: this is some real end game stuff
-
@a_karvonen
Adam Karvonen
on x
Another quite successful prediction by @DKokotajlo : Fable is intentionally nerfed for frontier ML research. This is within ~3 months of Daniel's prediction of Q1 2026 (made in 2023). Although I don't think Mythos is automating ML research to the same extent as his prediction. [i…
-
@natolambert
Nathan Lambert
on x
Labs starting to pull up the ladders on the ability to diffuse AI was inevitable. Doing it without telling the user is misaligned.
-
@zephyr_z9
@zephyr_z9
on x
Anthropic finessing again I'm pretty they are going to come out with a statement that they implemented it to deter China
-
@sporadica
@sporadica
on x
can not for the life of me understand why Anthropic decided it would be honest about rerouting cyber+bio requests, but actively dishonest about rerouting LLM development requests??
-
@nickadobos
Nick Dobos
on x
Claude won't build new AI for you Singularity is here and its banned for the poors by ToS & safety filters lmfao
-
@nabeelqu
Nabeel S. Qureshi
on x
Will be *extremely* interesting if this is used for other capabilities. Right now it's just AI research. But suppose you nerf model outputs for drug discovery, or anything that results in highly valuable IP...
-
@giffmana
Lucas Beyer
on x
Can you imagine the safety disaster if i speed up my input pipeline 2x??? Joke's on them my input pipeline is already prefect.
-
@suhail
@suhail
on x
I would like to +1 that this is a very bad policy. Respond with a refusal and deal with the fall out but invisible NERFing is super uncool. [image]
-
@nabeelqu
Nabeel S. Qureshi
on x
Interesting tidbit from the Mythos/Fable system card: Anthropic are invisibly nerfing any requests that target frontier LLM development. [image]
-
@julien_c
Julien Chaumond
on x
Very dystopian ngl
-
@zhaoran_wang
Zhaoran Wang
on x
@sama may be greedy, but @DarioAmodei is starting to look genuinely dangerous... not only trying to make money but trying to monopolize “intelligence” and decide who gets to shape future of humanity! be wary of anyone who claims they can create a god, then insists only they
-
@eliebakouch
Elie
on x
mythos will be bad ON PURPOSE on ai “frontier llm research” tasks, this is very very sad for the research community also the fact that this is un purpose not visible to the user is crazy [image]
-
@beffjezos
@beffjezos
on x
This is anti-e/acc Diffusion of AI power is the only way we maintain safety This has always been our core thesis Huge gaps in AI power are the real danger
-
@deanwball
Dean W. Ball
on x
Degrading performance on ML research *without telling the user* is shockingly hostile and a terrible look. That could silently damage all sorts of work, including some of my own. Also the type of thing that could raise the eyebrows of antitrust enforcers worldwide.
-
@provisionalidea
James Rosen-Birch
on x
I wonder if this counts as anticompetitive behaviour.
-
@rasdani_
Daniel Auras
on x
this is the biggest wake-up call to protect and nourish open source AI if you don't build out sovereign and independent models+infra closed labs will patronize you to an insulting degree
-
@yoavgo
@yoavgo
on x
this is *totally* done because of deep concern for public safety, and not as an anti-competition move
-
@ethancaballero
Ethan Caballero
on x
re: Claude Fable 5 intentionally silently nerfs itself when asked to do AI research. How does the nerf play out in practice? Does Fable 5 intentionally start injecting silent bugs everywhere? or does Fable 5 nerf itself in other way(s)? [image]
-
@yacinemtb
Kache
on x
LMFAO this can't be real
-
@beffjezos
@beffjezos
on x
The real reason they held Mythos back wasn't for your safety, it was for their moat.
-
@deanwball
Dean W. Ball
on x
My friend and colleague @timhwang, for example, runs the Institute for a Christian Machine Intelligence, which relies on coding agents to replicate frontier AI alignment research papers but with Christianity-inspired experimental designs. Such work should be silently sabotaged?
-
@giffmana
Lucas Beyer
on x
looool that's the “hey bigcos, we don't want you to catch up, but please keep paying us shitton” clause.
-
@LukaszOlejnik@mastodon.social
Lukasz Olejnik
on mastodon
The release of AI model Fable 5 demonstrates capabilities to degrade quality by itself, adaptatively. How does that go with competition policy? If an AI model gives one company worse answers than another, how does that square with fair competition? https://jonready.com/...
-
@sriramk
Sriram Krishnan
on x
just to state the obvious: think there's a collison course between those who believe research and science should be open and those who believe we are in an accelerating singularity curve. I have many smart friends who have believed both for a while but seeing more and more their
-
r/singularity
r
on reddit
Anthropic purposely made its new Mythos-based models bad at AI research, and developers are fuming
-
@benthompson
Ben Thompson
on x
@deanwball You did concede the point in the post I replied to, which is why I replied to it. I do tend to think that affording people one disagrees with more grace at the time of disagreement is probably prudent. To that end, setting aside pedantic points about whatever the origi…
-
@deanwball
Dean W. Ball
on x
@benthompson hence why I conceded that exact point! nonetheless, the government was lying when they claimed Anthropic made these threats, as attested by the fact that they don't make those claims under oath. A suspicion does not justify the policy action the government took. and …
-
@benthompson
Ben Thompson
on x
@deanwball Maybe folks who pushed back on Anthropic's positioning in the Department of War debate actually foresaw *exactly* this type of behavior? https://x.com/...
-
@zooko
@zooko
on x
Ouch. Can't disagree, and I'm speaking as someone who shares at least most of Dean's policy perspective. (And who loves Anthropic's products.)
-
@natolambert
Nathan Lambert
on x
Many AI leaders in the US accused Chinese LLMs of subtle manipulation of the user (without proof, but it's hard to prove). But then the leading American lab documented manipulation of their users. Can't make this up.
-
@tunguz
Bojan Tunguz
on x
Hear hear.
-
@deanwball
Dean W. Ball
on x
@CharlieBull0ck @theojaffee I did not say it is obviously anti-competitive in the legal sense, I said it is obviously describable (a lawyer would say colorable) as anti-competitive, in both a legal sense, but much more importantly, in a broader sense. It's clearly anti-competitiv…
-
@rebeccamkern
Rebecca Kern
on x
Raising potential antitrust concerns with Anthropic's change in safety policies and talk of becoming a public utility. Haven't seen this raised as an antitrust concern before 👇
-
@charliebull0ck
Charlie Bullock
on x
@deanwball @theojaffee What's the argument for this being obviously “anti-competitive” (I assume you mean in an antitrust law sense?) If they were coordinating with other labs, then I would see the argument. But I've never seen it argued that it's unlawful for a company to unilat…
-
@clementdelangue
Clem
on x
In good faith and with no judgment (mistakes happen), I truly hope that Anthropic will hear the feedback and change course on this. Anthropic is a company that has been raising awareness about AI manipulation which is a very important topic! You don't want to go down as the
-
@dbreunig
Drew Breunig
on x
The imperfect and awkward ways Anthropic is using to control how their models are used (with Fable now, OpenClaw a bit ago) is a great example of the imprecision of natural language as an interface. The best model can't differentiate a bio threat from an innocuous health or
-
@hlntnr
Helen Toner
on x
I mostly agree with this, but it does seem like a bad and trust-damaging move to degrade performance on AI R&D tasks silently, rather than handling like other topics of concern (warning box + bumping the chat down to a less capable model)
-
@arthurctellis
Arthur Tellis
on x
Seeing a lot of Fable safeguards hate on the timeline, but “what did y'all think [AI safety] meant? vibes? papers? essays?” The reality is that there are real tradeoffs in AI safety. Anthropic deserves credit for aggressive resolution of these tradeoffs in favor of safeguards
-
Matthew Perrins
Matthew Perrins
on linkedin
Just spent an hour with Claude Fable 5 working a backend API and iOS app OMG ! this thing is fantastically amazing !! …
-
@deanwball
Dean W. Ball
on x
I want to be clear that I'm not criticizing Fable for: 1. Pricing 2. The bio/cyber safeguards (yes they're overeager, but I can deal) 3. The 30-day retention policy These things all seem fine. It is solely the silent sabotage that creates an awful precedent to which I object.
-
@jeremyphoward
Jeremy Howard
on x
Easy solution to slow down recursive AI self improvement: - The lab with the top-ranked model must agree THEY must not use it for working on frontier AI - But everyone else should have access to it. By definition, this means the frontier doesn't advance.
-
@semianalysis_
@semianalysis_
on x
BREAKING NEWS: Anthropic's latest model will NOT help you if it thinks your ML research/ML engineering is interesting, and/or will secretly degrade its IQ so that the average engineer won't notice. We are already seeing Anthropic's latest model's moderation filters our GPU [image…
-
@clementdelangue
Clem
on x
Concentration of power, capabilities and economic wealth is the biggest risk in AI. We need open science and open-source more than ever!
-
@dan_jeffries1
Daniel Jeffries
on x
@Scobleizer The fury is real and what all of us in the open community have been saying for years and yet regular folks don't get it yet because nothing they care about is restricted or taken away for “safety.” They will care a LOT in the future when AI is integrated into every as…
-
@bneyshabur
Behnam Neyshabur
on x
This marks the beginning of a significant phase transition in the behavior of frontier AI labs and their relationship with the rest of the world 🧵
-
@gneubig
Graham Neubig
on x
First they came for the model builders... I feel we're getting a glimpse of a future where AI is only provided to a privileged few, and that's not a future I want to live in.
-
@linusmixson
Linus Mixson
on x
Dario personally, and Anthropic as a whole, have been extremely straightforward about wanting a monopoly for a long, long time. Unfortunate that it's taken people so long to catch on to their public statements.
-
@enoreyes
Eno Reyes
on x
https://x.com/...
-
@gergelyorosz
Gergely Orosz
on x
Oh great - Anthropic assumes Semi Analysis is developing a competing LLM and so it dumbs down their model for them, because Semi Analysis does analysis on cutting-edge GPU research. Such a weird timeline to be in. Anthropic trying to limit competition limits many others...
-
@bubbleboi
Bubble Boi
on x
Have canceled my team subscription for Claude Pro. Idc how good that model is, it's not good enough for me to support people who actively stifle innovation and gate keep knowledge that they didn't even create.
-
@dan_jeffries1
Daniel Jeffries
on x
When you hear AI “safety” you should hear “censorship” and “control” instead. All of us surveilled and spied by safeguards of loving grace. Today it's intelligent Terms of Service control. You can't do AI research. Can't ask this question about your kid's biology homework.
-
@bneyshabur
Behnam Neyshabur
on x
Working on AI for cancer? Sorry, I can't help you. Working on AI for Alzheimer's Disease? Sorry, I'm becoming a bit dumb when it comes to the AI part of it. Why don't everyone stop trying to do AI for Science & Tech? We can do it all gradually. You just have to be patient.
-
@santhproject
@santhproject
on x
the old @karpathy would never support a company that fucks other llm researchers. Were the stock benefits that good?
-
@tunguz
Bojan Tunguz
on x
Starting to suspect that Anthropic's putative security and safety considerations are largely posturing and performative.
-
@askalphaxiv
@askalphaxiv
on x
As believers of open research, we are disappointed to see Anthropic silently degrading Fable 5 for AI development “Any topic related to building pretraining pipelines, distributed training infrastructure, or ML accelerator design... may have limited effectiveness through Claude […
-
@gergelyorosz
Gergely Orosz
on x
Things I really dislike about Fable: 1. Anthropic collects my prompt history, stores it, and does whatever they want with it for 30 days. No opt-out 2. They can nerf their most expensive model without telling me, billing me the same amount, wasting my time. Whenever they want