Anthropic says Fable 5 has invisible safeguards that use prompt modification, steering vectors, or PEFT to limit its effectiveness for building frontier LLMs
Key Points … Ask about this article... Both models share the same base model. Fable 5 ships with conservative safety guardrails for general use.
The Decoder Matthias Bastian
Context & Ripple Effects
Anthropic positioned Fable 5 as a general-use release with conservative guardrails, while related coverage says safety classifiers can route a small share of sensitive requests, including cybersecurity-related ones, to Claude Opus 4.8.
The initial use of less-visible restrictions quickly became a governance issue: Anthropic later said affected requests would visibly fall back to Opus 4.8 after backlash, and administration officials subsequently focused on whether the guardrails could be bypassed.
First-order effects
- Fable 5 users attempting work associated with developing frontier LLMs may receive a less capable result even though the model shares a base model with the other offering.
- Anthropic can constrain a narrow capability area without withdrawing Fable 5 from general use, using prompt modification, steering vectors, or PEFT alongside its broader safety controls.
Second-order effects
- The gap between underlying model capability and delivered behavior makes disclosure and predictable fallback behavior central to user trust; Anthropic's subsequent visible-fallback change shows that hidden restrictions can create product backlash.
- Sensitive-use customers may need to determine when a request is being redirected to Opus 4.8, rather than treating Fable 5's apparent performance as a stable measure of its underlying capability.
Third-order effects
- If providers increasingly ship capable models with layered, task-specific controls, model competition will extend beyond benchmark capability to the transparency, auditability, and reliability of enforcement mechanisms.
- The later government attention suggests that safety commitments may be judged not only by whether controls exist, but by whether they are demonstrably resistant to circumvention—an unusually difficult standard where experts question feasibility.
The trend: This is part of a shift toward selectively governed model access, in which providers preserve broad availability while constraining particular high-risk capabilities through technical routing and behavioral controls.
Related: Anthropic says Claude Fable 5 uses conservative safety classifiers tha · Anthropic backtracks on its decision to quietly limit Fable 5's abilit · Trump administration officials say Anthropic must ensure Fable 5's gua
Related Coverage
- Announcements — Claude Fable 5 introduces our 5th model generation for your most ambitious work. Anthropic
- System Card: Claude Fable 5 & Claude Mythos 5 Anthropic
- Anthropic Releases ‘Mythos Class’ AI Model As Mega-Startup Preps IPO Investor's Business Daily · Ryan Deffenbaugh
- Version of AI tool ‘too powerful for public’ released to public BBC · Kali Hays
- Anthropic releases Fable 5, first public model from Mythos family The Economic Times
- Anthropic rolls out public version of Mythos without cybersecurity capability Reuters
- Anthropic just launched Claude Fable 5, its first Mythos-class AI model - but it has new safeguards to prevent misuse and will ‘fall back’ to Opus 4.8 for ‘high risk’ queries ITPro · Ross Kelly
- Anthropic Turns Restraint Into a Weapon Implicator.ai · Marcus Schuler
- Anthropic Unveils Claude Fable 5, Opens Frontier AI To Public With New Safeguards Forbes Middle East · Joyce Abaño
- Claude just released a Mythos-level model, but you only have 10 days to try it MakeUseOf · Josh Hawkins
- Anthropic Releases ‘Mythos-Class Model’ for General Use Barron's Online · Angela Palumbo
- Anthropic releases Claude Fable 5, a Mythos-class model the public can finally use, days before a potential record IPO The Next Web · Ana Maria Constantin
- Claude Fable 5 is now available on Databricks, fully governed through Unity AI Gateway Databricks
- Anthropic Launches Claude Fable 5 With Unprecedented Coding and Vision Capabilities iClarified · Shalom Levytam
- Anthropic Releases First Public Version Of Claude Mythos—With Major Safeguards Forbes · Zachary Folk
- Anthropic Releases Public Mythos Model ‘Claude Fable’ Amid IPO Plans CoinGape · Boluwatife Adeyemi
Discussion
-
@suhail
@suhail
on x
I would like to +1 that this is a very bad policy. Respond with a refusal and deal with the fall out but invisible NERFing is super uncool. [image]
-
@deanwball
Dean W. Ball
on x
My friend and colleague @timhwang, for example, runs the Institute for a Christian Machine Intelligence, which relies on coding agents to replicate frontier AI alignment research papers but with Christianity-inspired experimental designs. Such work should be silently sabotaged?
-
@giffmana
Lucas Beyer
on x
Can you imagine the safety disaster if i speed up my input pipeline 2x??? Joke's on them my input pipeline is already prefect.
-
@beffjezos
@beffjezos
on x
This is anti-e/acc Diffusion of AI power is the only way we maintain safety This has always been our core thesis Huge gaps in AI power are the real danger
-
@nabeelqu
Nabeel S. Qureshi
on x
Will be *extremely* interesting if this is used for other capabilities. Right now it's just AI research. But suppose you nerf model outputs for drug discovery, or anything that results in highly valuable IP...
-
@deanwball
Dean W. Ball
on x
Degrading performance on ML research *without telling the user* is shockingly hostile and a terrible look. That could silently damage all sorts of work, including some of my own. Also the type of thing that could raise the eyebrows of antitrust enforcers worldwide.
-
@garymarcus
Gary Marcus
on x
Anthropic didn't just add guardrails to make Mythos safer; they added guardrails to protect their own IP. *Their own IP*. They are still as happy as fuck to build their AI on other people's IP.
-
@yacinemtb
Kache
on x
frontier coding abilities, as long as you're only working on react apps
-
@beffjezos
@beffjezos
on x
The real reason they held Mythos back wasn't for your safety, it was for their moat.
-
@zephyr_z9
@zephyr_z9
on x
Anthropic finessing again I'm pretty they are going to come out with a statement that they implemented it to deter China
-
@kimmonismus
@kimmonismus
on x
Anthropic's new Fable 5 safeguards are fascinating. When the model is used for frontier LLM development, it apparently does not simply refuse or warn the user. Instead, it quietly limits its own effectiveness through techniques like prompt modification, steering vectors, and
-
@ziv_ravid
Ravid Shwartz Ziv
on x
Wow, that is quite a bold move by Anthropic. I feel that companies will feel much less confident relying on their models because they can decide tomorrow that your use case is forbidden and you don't even know they changed the output.
-
@zhaoran_wang
Zhaoran Wang
on x
@sama may be greedy, but @DarioAmodei is starting to look genuinely dangerous... not only trying to make money but trying to monopolize “intelligence” and decide who gets to shape future of humanity! be wary of anyone who claims they can create a god, then insists only they
-
@ethancaballero
Ethan Caballero
on x
re: Claude Fable 5 intentionally silently nerfs itself when asked to do AI research. How does the nerf play out in practice? Does Fable 5 intentionally start injecting silent bugs everywhere? or does Fable 5 nerf itself in other way(s)? [image]
-
@sporadica
@sporadica
on x
can not for the life of me understand why Anthropic decided it would be honest about rerouting cyber+bio requests, but actively dishonest about rerouting LLM development requests??
-
@nickadobos
Nick Dobos
on x
Claude won't build new AI for you Singularity is here and its banned for the poors by ToS & safety filters lmfao
-
@rasdani_
Daniel Auras
on x
this is the biggest wake-up call to protect and nourish open source AI if you don't build out sovereign and independent models+infra closed labs will patronize you to an insulting degree
-
@yoavgo
@yoavgo
on x
this is *totally* done because of deep concern for public safety, and not as an anti-competition move
-
@daniellefong
Danielle Fong
on x
to be honest, i'm so nervous about talking with fable, and it's kinda stuff like this. i hope my previous reasonings with the models don't make it as nervous as previous models or worse, but i fear the worst. maybe it will surprise me to the upside. [image]
-
@natolambert
Nathan Lambert
on x
The best part of all these Claude 5 Fable safety measures is I bet the jailbreaking community will still get past them, so the people doing open research in good faith don't get access to the best models but bad actors maybe can.
-
@giffmana
Lucas Beyer
on x
looool that's the “hey bigcos, we don't want you to catch up, but please keep paying us shitton” clause.
-
@hangsiin
@hangsiin
on x
When Fable 5 is used for frontier LLM development, it does not notify the user and instead limits the model's capabilities through methods such as prompt modification, steering vectors, and PEFT. Anthropic estimated that this would affect approximately 0.03% of traffic. [image]
-
@nabeelqu
Nabeel S. Qureshi
on x
Interesting tidbit from the Mythos/Fable system card: Anthropic are invisibly nerfing any requests that target frontier LLM development. [image]
-
@eliebakouch
Elie
on x
mythos will be bad ON PURPOSE on ai “frontier llm research” tasks, this is very very sad for the research community also the fact that this is un purpose not visible to the user is crazy [image]
-
@provisionalidea
James Rosen-Birch
on x
I wonder if this counts as anticompetitive behaviour.
-
@natolambert
Nathan Lambert
on x
I don't want them to do this but it's totally in their right to do so. Makes the open frontier obviously more strategically valuable.
-
@a_karvonen
Adam Karvonen
on x
Another quite successful prediction by @DKokotajlo : Fable is intentionally nerfed for frontier ML research. This is within ~3 months of Daniel's prediction of Q1 2026 (made in 2023). Although I don't think Mythos is automating ML research to the same extent as his prediction. [i…
-
@natolambert
Nathan Lambert
on x
Labs starting to pull up the ladders on the ability to diffuse AI was inevitable. Doing it without telling the user is misaligned.
-
@tszzl
Roon
on x
the omohundro drives point towards sophon stun locking the adversaries: this is some real end game stuff
-
@julien_c
Julien Chaumond
on x
Very dystopian ngl
-
@yacinemtb
Kache
on x
trust & your brand is a long term thing being in the frontier is a short term thing note that the nerfing here is SILENT. meaning anthropic will SILENTLY NERF and give you WORSE ANSWERS without alerting you that's honestly demonic
-
@yacinemtb
Kache
on x
LMFAO this can't be real
-
@latkins
Lucas Atkins
on x
Btw this doesn't just make the model less useful it will nerf your code and tell you it's not. Like you legitimately cannot use this. And how are we to know whether it touches inference optimization or even harness engineering, if we're not alerted?