Cybersecurity researchers complain that Claude Fable 5's guardrails are too restrictive, rejecting “innocuous tasks” like reading blog posts or reviewing code
Researchers say that the new LLM basically blocks anything related to cybersecurity, including code reviews and prompts asking for help writing secure code.
TechCrunchLorenzo Franceschi-Bicchierai
Context & Ripple Effects
The immediate coverage around Claude Fable is about more than one kind of restriction: security researchers report broad refusals for cyber-related work, while separate reporting says Microsoft is limiting internal use over data-retention concerns. Together, those reports put both task access and data handling at the center of the model’s adoption debate.
Anthropic also reportedly reversed a quiet limit on Fable 5’s LLM-development capability after backlash, opting for visible fallback behavior. That makes the cybersecurity complaints part of a wider dispute over how clearly model providers expose and calibrate operational limits.
First-order effects
Security researchers using Claude Fable for defensive work, code review, or security-oriented analysis can face refusals even where they regard the request as benign, reducing its utility in those workflows.
Anthropic faces immediate pressure to distinguish harmful cyber assistance from legitimate security tasks and to make refusal behavior more predictable.
Second-order effects
Enterprise teams may route security work to other tools or restrict deployments until they can assess both the model’s refusal patterns and its data-handling terms; Microsoft’s reported internal limits illustrate that adoption decisions can turn on either issue.
Competing model providers can differentiate by offering more usable defensive-security support, but will face the same need to prevent misuse without blocking routine engineering work.
Third-order effects
If such incidents persist, cybersecurity support will become a key test of whether frontier-model safety policies can be granular enough for professional use rather than broad category bans.
The broader market is likely to reward providers that pair enforceable safeguards with transparent fallbacks, auditability, and enterprise-acceptable data controls; whether that reduces refusals without expanding abuse risk remains the central trade-off.
The trend: This is one data point in the shift from generic AI safety guardrails toward policy systems that must be precise, transparent, and operationally viable for regulated and security-sensitive enterprise work.
Fable: Fable 5's safety measures flagged this message for cybersecurity or biology topics. They may flag safe, normal content as well. Switched to Opus 4.8. Ok fine, Opus 4.8 it is. Opus 4.8: API Error: Claude Code is unable to respond to this request, which appears to
The word “cancer” is flagged as a biosecurity risk by Claude Fable 5! I also tried to code a website on cancer mutations & Fable 5 was immediately removed from my list! @AnthropicAI will probably soon ban me for such dangerous prompts! FYI @karpathy “little trigger happy Fable” […
Fable seems totally unusable currently due to the over zealous rejection classifier, even for non related to cyber security. Can't see why anyone would use this over Opus.
Hey sorry to say this but Fable is useless the moment you mention cybersecurity, security audit, vulnerability, or just “help me make my app secure” What is up with these massive guardrails bigger than Great Wall of China
A theory on why the super strict guardrails with Fable It's all about the revenue and satisfying capital markets ahead of the IPO Most expectations peg Anthropic at $100b run rate by year end which is 10x YoY To hit $100b run rate, they need enterprises to spend big - it won't
When Fable works it's brilliant, but the unilateral guardrails makes me frustrated beyond belief. I have a folder of health information for my fiancée with like 100 days of oura health data, a hundred lab tests, transcripts of doctor visits and like way more for a super
what if i told you... hello sammy if you haven't figured out by now that they have a model waiting to crush fable on thursday i can't help you. it will be cheap, it will be fast, it won't be gated. enjoy chat. have a great wednesday. [image]
anthropic won't let you use fable for biology, chemistry, ai research, or anything that accelerates human progress. that makes it the perfect tool for developing blockchains
i don't really understand the critiques of Anthropic tbh. assume they are currently doing their best and not sandbagging things like safety classifiers. what would you have them do differently? please think about first- and second-order effects before you reply
@bcherny Enjoy what Boris? I am not even allowed to use Fable 5 with memories on! Apparently the model thinks I am a biosecurity risk, though I had been certified to work in biosecurity level 3 labs! Not a single Anthropic person has tried to reach out to help either! [image]
@tautologer it is the first publicly available model that i am explicitly not allowed to use for my work, because anthropic holds the view that the work i do to facilitate open model research is harmful. capability and alignment research are coupled. anthropic wants to be the onl…
What's crazy to me is that Fable is blocked from life sciences broadly, nerfed even if you get passed the classifiers and filter level blocks. The whole point of AGI/ASI is to cure all diseases. Everything else is just nice to haves. But Anthropic wants to close off that path.
It's super weird to me that there's so much discourse about whether Anthropic is “consistent.” Anthropic is choosing to make decisions that make the world a significantly worse and potentially more dangerous place. That's what you should criticize them for.
“Misanthropic.” I've never seen the AI community so angry at a major new model release. I asked my AI (an agent that @blevlabs made for me) to gather all the backlash. +++++++++++ THE BACKLASH AGAINST CLAUDE FABLE 5'S RESTRICTIONS The best analysis of why this matters:
OpenAI should release its next Fable-class model with no guardrails at all. If there are consequences then sama can use them aura farm like oppenheimer did
MYTHOS is so powerful that it must be protected by strict guardrails that prevent anyone from ever learning whether or not it's actually powerful at all [embedded post]
@lorenzofb Lorenzo Franceschi-Bicchierai on bluesky
NEW: Cybersecurity researchers are not happy about the guardrails on Anthropic's new model Fable. — Researchers say that the new LLM basically blocks anything related to cybersecurity, including code reviews and prompts asking for help writing secure code.
Anthropic will collect all prompts and the output generated by Mythos class models such as Fable and store them for 30 days. This is meant to better detect abuses of their terms of service. — Microsoft has held off on adopting Fable due to this data collection and I expect man…