/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

Anthropic says Fable 5 has invisible safeguards that use prompt modification, steering vectors, or PEFT to limit its effectiveness for building frontier LLMs

Key Points … Ask about this article...  Both models share the same base model.  Fable 5 ships with conservative safety guardrails for general use.

The Decoder Matthias Bastian

Context & Ripple Effects

Anthropic positioned Fable 5 as a general-use release with conservative guardrails, while related coverage says safety classifiers can route a small share of sensitive requests, including cybersecurity-related ones, to Claude Opus 4.8.

The initial use of less-visible restrictions quickly became a governance issue: Anthropic later said affected requests would visibly fall back to Opus 4.8 after backlash, and administration officials subsequently focused on whether the guardrails could be bypassed.

First-order effects

  • Fable 5 users attempting work associated with developing frontier LLMs may receive a less capable result even though the model shares a base model with the other offering.
  • Anthropic can constrain a narrow capability area without withdrawing Fable 5 from general use, using prompt modification, steering vectors, or PEFT alongside its broader safety controls.

Second-order effects

  • The gap between underlying model capability and delivered behavior makes disclosure and predictable fallback behavior central to user trust; Anthropic's subsequent visible-fallback change shows that hidden restrictions can create product backlash.
  • Sensitive-use customers may need to determine when a request is being redirected to Opus 4.8, rather than treating Fable 5's apparent performance as a stable measure of its underlying capability.

Third-order effects

  • If providers increasingly ship capable models with layered, task-specific controls, model competition will extend beyond benchmark capability to the transparency, auditability, and reliability of enforcement mechanisms.
  • The later government attention suggests that safety commitments may be judged not only by whether controls exist, but by whether they are demonstrably resistant to circumvention—an unusually difficult standard where experts question feasibility.

The trend: This is part of a shift toward selectively governed model access, in which providers preserve broad availability while constraining particular high-risk capabilities through technical routing and behavioral controls.

Discussion

  • @suhail @suhail on x
    I would like to +1 that this is a very bad policy. Respond with a refusal and deal with the fall out but invisible NERFing is super uncool. [image]
  • @deanwball Dean W. Ball on x
    My friend and colleague @timhwang, for example, runs the Institute for a Christian Machine Intelligence, which relies on coding agents to replicate frontier AI alignment research papers but with Christianity-inspired experimental designs. Such work should be silently sabotaged?
  • @giffmana Lucas Beyer on x
    Can you imagine the safety disaster if i speed up my input pipeline 2x??? Joke's on them my input pipeline is already prefect.
  • @beffjezos @beffjezos on x
    This is anti-e/acc Diffusion of AI power is the only way we maintain safety This has always been our core thesis Huge gaps in AI power are the real danger
  • @nabeelqu Nabeel S. Qureshi on x
    Will be *extremely* interesting if this is used for other capabilities. Right now it's just AI research. But suppose you nerf model outputs for drug discovery, or anything that results in highly valuable IP...
  • @deanwball Dean W. Ball on x
    Degrading performance on ML research *without telling the user* is shockingly hostile and a terrible look. That could silently damage all sorts of work, including some of my own. Also the type of thing that could raise the eyebrows of antitrust enforcers worldwide.
  • @garymarcus Gary Marcus on x
    Anthropic didn't just add guardrails to make Mythos safer; they added guardrails to protect their own IP. *Their own IP*. They are still as happy as fuck to build their AI on other people's IP.
  • @yacinemtb Kache on x
    frontier coding abilities, as long as you're only working on react apps
  • @beffjezos @beffjezos on x
    The real reason they held Mythos back wasn't for your safety, it was for their moat.
  • @zephyr_z9 @zephyr_z9 on x
    Anthropic finessing again I'm pretty they are going to come out with a statement that they implemented it to deter China
  • @kimmonismus @kimmonismus on x
    Anthropic's new Fable 5 safeguards are fascinating. When the model is used for frontier LLM development, it apparently does not simply refuse or warn the user. Instead, it quietly limits its own effectiveness through techniques like prompt modification, steering vectors, and
  • @ziv_ravid Ravid Shwartz Ziv on x
    Wow, that is quite a bold move by Anthropic. I feel that companies will feel much less confident relying on their models because they can decide tomorrow that your use case is forbidden and you don't even know they changed the output.
  • @zhaoran_wang Zhaoran Wang on x
    @sama may be greedy, but @DarioAmodei is starting to look genuinely dangerous... not only trying to make money but trying to monopolize “intelligence” and decide who gets to shape future of humanity! be wary of anyone who claims they can create a god, then insists only they
  • @ethancaballero Ethan Caballero on x
    re: Claude Fable 5 intentionally silently nerfs itself when asked to do AI research. How does the nerf play out in practice? Does Fable 5 intentionally start injecting silent bugs everywhere? or does Fable 5 nerf itself in other way(s)? [image]
  • @sporadica @sporadica on x
    can not for the life of me understand why Anthropic decided it would be honest about rerouting cyber+bio requests, but actively dishonest about rerouting LLM development requests??
  • @nickadobos Nick Dobos on x
    Claude won't build new AI for you Singularity is here and its banned for the poors by ToS & safety filters lmfao
  • @rasdani_ Daniel Auras on x
    this is the biggest wake-up call to protect and nourish open source AI if you don't build out sovereign and independent models+infra closed labs will patronize you to an insulting degree
  • @yoavgo @yoavgo on x
    this is *totally* done because of deep concern for public safety, and not as an anti-competition move
  • @daniellefong Danielle Fong on x
    to be honest, i'm so nervous about talking with fable, and it's kinda stuff like this. i hope my previous reasonings with the models don't make it as nervous as previous models or worse, but i fear the worst. maybe it will surprise me to the upside. [image]
  • @natolambert Nathan Lambert on x
    The best part of all these Claude 5 Fable safety measures is I bet the jailbreaking community will still get past them, so the people doing open research in good faith don't get access to the best models but bad actors maybe can.
  • @giffmana Lucas Beyer on x
    looool that's the “hey bigcos, we don't want you to catch up, but please keep paying us shitton” clause.
  • @hangsiin @hangsiin on x
    When Fable 5 is used for frontier LLM development, it does not notify the user and instead limits the model's capabilities through methods such as prompt modification, steering vectors, and PEFT. Anthropic estimated that this would affect approximately 0.03% of traffic. [image]
  • @nabeelqu Nabeel S. Qureshi on x
    Interesting tidbit from the Mythos/Fable system card: Anthropic are invisibly nerfing any requests that target frontier LLM development. [image]
  • @eliebakouch Elie on x
    mythos will be bad ON PURPOSE on ai “frontier llm research” tasks, this is very very sad for the research community also the fact that this is un purpose not visible to the user is crazy [image]
  • @provisionalidea James Rosen-Birch on x
    I wonder if this counts as anticompetitive behaviour.
  • @natolambert Nathan Lambert on x
    I don't want them to do this but it's totally in their right to do so. Makes the open frontier obviously more strategically valuable.
  • @a_karvonen Adam Karvonen on x
    Another quite successful prediction by @DKokotajlo : Fable is intentionally nerfed for frontier ML research. This is within ~3 months of Daniel's prediction of Q1 2026 (made in 2023). Although I don't think Mythos is automating ML research to the same extent as his prediction. [i…
  • @natolambert Nathan Lambert on x
    Labs starting to pull up the ladders on the ability to diffuse AI was inevitable. Doing it without telling the user is misaligned.
  • @tszzl Roon on x
    the omohundro drives point towards sophon stun locking the adversaries: this is some real end game stuff
  • @julien_c Julien Chaumond on x
    Very dystopian ngl
  • @yacinemtb Kache on x
    trust & your brand is a long term thing being in the frontier is a short term thing note that the nerfing here is SILENT. meaning anthropic will SILENTLY NERF and give you WORSE ANSWERS without alerting you that's honestly demonic
  • @yacinemtb Kache on x
    LMFAO this can't be real
  • @latkins Lucas Atkins on x
    Btw this doesn't just make the model less useful it will nerf your code and tell you it's not. Like you legitimately cannot use this. And how are we to know whether it touches inference optimization or even harness engineering, if we're not alerted?