/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

Anthropic says Fable 5 has invisible safeguards that use prompt modification, steering vectors, or PEFT to limit its effectiveness for building frontier LLMs

Key Points … Ask about this article...  Both models share the same base model.  Fable 5 ships with conservative safety guardrails for general use.

The Decoder Matthias Bastian

Context & Ripple Effects

Anthropic’s related coverage describes Fable 5 as sharing a base model with another offering while applying conservative controls in sensitive areas, including cybersecurity. The company later said certain constrained requests would visibly fall back to Opus 4.8 after criticism of quietly limiting capability.

The episode sits alongside growing scrutiny of whether model safeguards can be bypassed: subsequent coverage says administration officials sought assurance that Fable 5’s guardrails could not be circumvented before a rerelease.

First-order effects

  • Fable 5 users attempting work connected to building frontier LLMs receive a less effective model response, even where the underlying base model could otherwise perform better.
  • Anthropic takes on an immediate product-governance burden: prompt modification, steering vectors, and PEFT-based restrictions must operate reliably without obscuring ordinary-use behavior.

Second-order effects

  • The use of invisible capability limits makes disclosure and routing policy a competitive product issue; backlash already pushed Anthropic toward visible fallback to Opus 4.8 for constrained requests.
  • Enterprise customers and platform partners may evaluate Fable 5 not just on benchmark capability, but on whether safety interventions are predictable, auditable, and compatible with their data and workflow requirements.

Third-order effects

  • If frontier-model providers increasingly ship one underlying capability with policy-dependent suppression layers, model access may be defined as much by runtime governance and routing as by the base model itself.
  • The coverage suggests safety claims will face pressure from both users demanding transparent limits and policymakers demanding safeguards that resist circumvention; proving both simultaneously may remain difficult.

The trend: This is one data point in the shift from static model releases toward dynamically governed AI products whose capabilities vary by task, risk category, and enforcement policy.

Discussion

  • @yacinemtb Kache on x
    trust & your brand is a long term thing being in the frontier is a short term thing note that the nerfing here is SILENT. meaning anthropic will SILENTLY NERF and give you WORSE ANSWERS without alerting you that's honestly demonic
  • @daniellefong Danielle Fong on x
    to be honest, i'm so nervous about talking with fable, and it's kinda stuff like this. i hope my previous reasonings with the models don't make it as nervous as previous models or worse, but i fear the worst. maybe it will surprise me to the upside. [image]
  • @yacinemtb Kache on x
    frontier coding abilities, as long as you're only working on react apps
  • @hangsiin @hangsiin on x
    When Fable 5 is used for frontier LLM development, it does not notify the user and instead limits the model's capabilities through methods such as prompt modification, steering vectors, and PEFT. Anthropic estimated that this would affect approximately 0.03% of traffic. [image]
  • @natolambert Nathan Lambert on x
    The best part of all these Claude 5 Fable safety measures is I bet the jailbreaking community will still get past them, so the people doing open research in good faith don't get access to the best models but bad actors maybe can.
  • @kimmonismus @kimmonismus on x
    Anthropic's new Fable 5 safeguards are fascinating. When the model is used for frontier LLM development, it apparently does not simply refuse or warn the user. Instead, it quietly limits its own effectiveness through techniques like prompt modification, steering vectors, and
  • @ziv_ravid Ravid Shwartz Ziv on x
    Wow, that is quite a bold move by Anthropic. I feel that companies will feel much less confident relying on their models because they can decide tomorrow that your use case is forbidden and you don't even know they changed the output.
  • @garymarcus Gary Marcus on x
    Anthropic didn't just add guardrails to make Mythos safer; they added guardrails to protect their own IP. *Their own IP*. They are still as happy as fuck to build their AI on other people's IP.
  • @natolambert Nathan Lambert on x
    I don't want them to do this but it's totally in their right to do so. Makes the open frontier obviously more strategically valuable.
  • @latkins Lucas Atkins on x
    Btw this doesn't just make the model less useful it will nerf your code and tell you it's not. Like you legitimately cannot use this. And how are we to know whether it touches inference optimization or even harness engineering, if we're not alerted?
  • @tszzl Roon on x
    the omohundro drives point towards sophon stun locking the adversaries: this is some real end game stuff
  • @a_karvonen Adam Karvonen on x
    Another quite successful prediction by @DKokotajlo : Fable is intentionally nerfed for frontier ML research. This is within ~3 months of Daniel's prediction of Q1 2026 (made in 2023). Although I don't think Mythos is automating ML research to the same extent as his prediction. [i…
  • @natolambert Nathan Lambert on x
    Labs starting to pull up the ladders on the ability to diffuse AI was inevitable. Doing it without telling the user is misaligned.
  • @zephyr_z9 @zephyr_z9 on x
    Anthropic finessing again I'm pretty they are going to come out with a statement that they implemented it to deter China
  • @sporadica @sporadica on x
    can not for the life of me understand why Anthropic decided it would be honest about rerouting cyber+bio requests, but actively dishonest about rerouting LLM development requests??
  • @nickadobos Nick Dobos on x
    Claude won't build new AI for you Singularity is here and its banned for the poors by ToS & safety filters lmfao
  • @nabeelqu Nabeel S. Qureshi on x
    Will be *extremely* interesting if this is used for other capabilities. Right now it's just AI research. But suppose you nerf model outputs for drug discovery, or anything that results in highly valuable IP...
  • @giffmana Lucas Beyer on x
    Can you imagine the safety disaster if i speed up my input pipeline 2x??? Joke's on them my input pipeline is already prefect.
  • @suhail @suhail on x
    I would like to +1 that this is a very bad policy. Respond with a refusal and deal with the fall out but invisible NERFing is super uncool. [image]
  • @nabeelqu Nabeel S. Qureshi on x
    Interesting tidbit from the Mythos/Fable system card: Anthropic are invisibly nerfing any requests that target frontier LLM development. [image]
  • @julien_c Julien Chaumond on x
    Very dystopian ngl
  • @zhaoran_wang Zhaoran Wang on x
    @sama may be greedy, but @DarioAmodei is starting to look genuinely dangerous... not only trying to make money but trying to monopolize “intelligence” and decide who gets to shape future of humanity! be wary of anyone who claims they can create a god, then insists only they
  • @eliebakouch Elie on x
    mythos will be bad ON PURPOSE on ai “frontier llm research” tasks, this is very very sad for the research community also the fact that this is un purpose not visible to the user is crazy [image]
  • @beffjezos @beffjezos on x
    This is anti-e/acc Diffusion of AI power is the only way we maintain safety This has always been our core thesis Huge gaps in AI power are the real danger
  • @deanwball Dean W. Ball on x
    Degrading performance on ML research *without telling the user* is shockingly hostile and a terrible look. That could silently damage all sorts of work, including some of my own. Also the type of thing that could raise the eyebrows of antitrust enforcers worldwide.
  • @provisionalidea James Rosen-Birch on x
    I wonder if this counts as anticompetitive behaviour.
  • @rasdani_ Daniel Auras on x
    this is the biggest wake-up call to protect and nourish open source AI if you don't build out sovereign and independent models+infra closed labs will patronize you to an insulting degree
  • @yoavgo @yoavgo on x
    this is *totally* done because of deep concern for public safety, and not as an anti-competition move
  • @ethancaballero Ethan Caballero on x
    re: Claude Fable 5 intentionally silently nerfs itself when asked to do AI research. How does the nerf play out in practice? Does Fable 5 intentionally start injecting silent bugs everywhere? or does Fable 5 nerf itself in other way(s)? [image]
  • @yacinemtb Kache on x
    LMFAO this can't be real
  • @beffjezos @beffjezos on x
    The real reason they held Mythos back wasn't for your safety, it was for their moat.
  • @deanwball Dean W. Ball on x
    My friend and colleague @timhwang, for example, runs the Institute for a Christian Machine Intelligence, which relies on coding agents to replicate frontier AI alignment research papers but with Christianity-inspired experimental designs. Such work should be silently sabotaged?
  • @giffmana Lucas Beyer on x
    looool that's the “hey bigcos, we don't want you to catch up, but please keep paying us shitton” clause.
  • @LukaszOlejnik@mastodon.social Lukasz Olejnik on mastodon
    The release of AI model Fable 5 demonstrates capabilities to degrade quality by itself, adaptatively.  How does that go with competition policy?  If an AI model gives one company worse answers than another, how does that square with fair competition? https://jonready.com/...
  • @sriramk Sriram Krishnan on x
    just to state the obvious: think there's a collison course between those who believe research and science should be open and those who believe we are in an accelerating singularity curve. I have many smart friends who have believed both for a while but seeing more and more their
  • r/singularity r on reddit
    Anthropic purposely made its new Mythos-based models bad at AI research, and developers are fuming
  • @benthompson Ben Thompson on x
    @deanwball You did concede the point in the post I replied to, which is why I replied to it. I do tend to think that affording people one disagrees with more grace at the time of disagreement is probably prudent. To that end, setting aside pedantic points about whatever the origi…
  • @deanwball Dean W. Ball on x
    @benthompson hence why I conceded that exact point! nonetheless, the government was lying when they claimed Anthropic made these threats, as attested by the fact that they don't make those claims under oath. A suspicion does not justify the policy action the government took. and …
  • @benthompson Ben Thompson on x
    @deanwball Maybe folks who pushed back on Anthropic's positioning in the Department of War debate actually foresaw *exactly* this type of behavior? https://x.com/...
  • @zooko @zooko on x
    Ouch. Can't disagree, and I'm speaking as someone who shares at least most of Dean's policy perspective. (And who loves Anthropic's products.)
  • @natolambert Nathan Lambert on x
    Many AI leaders in the US accused Chinese LLMs of subtle manipulation of the user (without proof, but it's hard to prove). But then the leading American lab documented manipulation of their users. Can't make this up.
  • @tunguz Bojan Tunguz on x
    Hear hear.
  • @deanwball Dean W. Ball on x
    @CharlieBull0ck @theojaffee I did not say it is obviously anti-competitive in the legal sense, I said it is obviously describable (a lawyer would say colorable) as anti-competitive, in both a legal sense, but much more importantly, in a broader sense. It's clearly anti-competitiv…
  • @rebeccamkern Rebecca Kern on x
    Raising potential antitrust concerns with Anthropic's change in safety policies and talk of becoming a public utility. Haven't seen this raised as an antitrust concern before 👇
  • @charliebull0ck Charlie Bullock on x
    @deanwball @theojaffee What's the argument for this being obviously “anti-competitive” (I assume you mean in an antitrust law sense?) If they were coordinating with other labs, then I would see the argument. But I've never seen it argued that it's unlawful for a company to unilat…
  • @clementdelangue Clem on x
    In good faith and with no judgment (mistakes happen), I truly hope that Anthropic will hear the feedback and change course on this. Anthropic is a company that has been raising awareness about AI manipulation which is a very important topic! You don't want to go down as the
  • @dbreunig Drew Breunig on x
    The imperfect and awkward ways Anthropic is using to control how their models are used (with Fable now, OpenClaw a bit ago) is a great example of the imprecision of natural language as an interface. The best model can't differentiate a bio threat from an innocuous health or
  • @hlntnr Helen Toner on x
    I mostly agree with this, but it does seem like a bad and trust-damaging move to degrade performance on AI R&D tasks silently, rather than handling like other topics of concern (warning box + bumping the chat down to a less capable model)
  • @arthurctellis Arthur Tellis on x
    Seeing a lot of Fable safeguards hate on the timeline, but “what did y'all think [AI safety] meant? vibes? papers? essays?” The reality is that there are real tradeoffs in AI safety. Anthropic deserves credit for aggressive resolution of these tradeoffs in favor of safeguards
  • Matthew Perrins Matthew Perrins on linkedin
    Just spent an hour with Claude Fable 5 working a backend API and iOS app OMG ! this thing is fantastically amazing !! …
  • @deanwball Dean W. Ball on x
    I want to be clear that I'm not criticizing Fable for: 1. Pricing 2. The bio/cyber safeguards (yes they're overeager, but I can deal) 3. The 30-day retention policy These things all seem fine. It is solely the silent sabotage that creates an awful precedent to which I object.
  • @jeremyphoward Jeremy Howard on x
    Easy solution to slow down recursive AI self improvement: - The lab with the top-ranked model must agree THEY must not use it for working on frontier AI - But everyone else should have access to it. By definition, this means the frontier doesn't advance.
  • @semianalysis_ @semianalysis_ on x
    BREAKING NEWS: Anthropic's latest model will NOT help you if it thinks your ML research/ML engineering is interesting, and/or will secretly degrade its IQ so that the average engineer won't notice. We are already seeing Anthropic's latest model's moderation filters our GPU [image…
  • @clementdelangue Clem on x
    Concentration of power, capabilities and economic wealth is the biggest risk in AI. We need open science and open-source more than ever!
  • @dan_jeffries1 Daniel Jeffries on x
    @Scobleizer The fury is real and what all of us in the open community have been saying for years and yet regular folks don't get it yet because nothing they care about is restricted or taken away for “safety.” They will care a LOT in the future when AI is integrated into every as…
  • @bneyshabur Behnam Neyshabur on x
    This marks the beginning of a significant phase transition in the behavior of frontier AI labs and their relationship with the rest of the world 🧵
  • @gneubig Graham Neubig on x
    First they came for the model builders... I feel we're getting a glimpse of a future where AI is only provided to a privileged few, and that's not a future I want to live in.
  • @linusmixson Linus Mixson on x
    Dario personally, and Anthropic as a whole, have been extremely straightforward about wanting a monopoly for a long, long time. Unfortunate that it's taken people so long to catch on to their public statements.
  • @enoreyes Eno Reyes on x
    https://x.com/...
  • @gergelyorosz Gergely Orosz on x
    Oh great - Anthropic assumes Semi Analysis is developing a competing LLM and so it dumbs down their model for them, because Semi Analysis does analysis on cutting-edge GPU research. Such a weird timeline to be in. Anthropic trying to limit competition limits many others...
  • @bubbleboi Bubble Boi on x
    Have canceled my team subscription for Claude Pro. Idc how good that model is, it's not good enough for me to support people who actively stifle innovation and gate keep knowledge that they didn't even create.
  • @dan_jeffries1 Daniel Jeffries on x
    When you hear AI “safety” you should hear “censorship” and “control” instead. All of us surveilled and spied by safeguards of loving grace. Today it's intelligent Terms of Service control. You can't do AI research. Can't ask this question about your kid's biology homework.
  • @bneyshabur Behnam Neyshabur on x
    Working on AI for cancer? Sorry, I can't help you. Working on AI for Alzheimer's Disease? Sorry, I'm becoming a bit dumb when it comes to the AI part of it. Why don't everyone stop trying to do AI for Science & Tech? We can do it all gradually. You just have to be patient.
  • @santhproject @santhproject on x
    the old @karpathy would never support a company that fucks other llm researchers. Were the stock benefits that good?
  • @tunguz Bojan Tunguz on x
    Starting to suspect that Anthropic's putative security and safety considerations are largely posturing and performative.
  • @askalphaxiv @askalphaxiv on x
    As believers of open research, we are disappointed to see Anthropic silently degrading Fable 5 for AI development “Any topic related to building pretraining pipelines, distributed training infrastructure, or ML accelerator design... may have limited effectiveness through Claude […
  • @gergelyorosz Gergely Orosz on x
    Things I really dislike about Fable: 1. Anthropic collects my prompt history, stores it, and does whatever they want with it for 30 days. No opt-out 2. They can nerf their most expensive model without telling me, billing me the same amount, wasting my time. Whenever they want