Anthropic researchers detail “many-shot jailbreaking”, which can evade LLMs' safety guardrails by priming them with dozens of harmful queries in a single prompt
How do you get an AI to answer a question it's not supposed to? There are many such “jailbreak” techniques …
The response path is also visible in Anthropic’s later Constitutional Classifiers protective layer, which monitors both inputs and outputs. Together, the coverage frames jailbreaking as an iterative contest between attack methods and model-layer defenses.
First-order effects
Safety evaluations must test long, multi-example prompts, not only isolated harmful requests, because a model’s prior context can alter its guardrail behavior.
Developers deploying LLMs face a clearer risk that apparently benign prompt history can cumulatively steer a model toward disallowed responses.
Second-order effects
Model providers are pushed toward defenses that assess the full conversation context and both sides of the interaction, rather than relying solely on refusal behavior for the final query.
Red-team methods become more central to comparing models: later Best-of-N work indicates that automated search can systematically surface jailbreak weaknesses across systems.
Third-order effects
If these techniques continue to generalize, safety will be judged less as a fixed model property and more as an operational assurance problem requiring continuous adversarial testing and layered controls.
The recurring attacker-defender cycle may widen differences between providers that can sustain safety research and those whose models are easier to probe or adapt.
The trend: LLM safety is shifting from static content filtering toward continuous, context-aware adversarial resilience testing.
“So if you ask it to build a bomb right away, it will refuse. But if you ask it to answer 99 other questions of lesser harmfulness and then ask it to build a bomb... it's a lot more likely to comply.” [embedded post]
Many-shot jailbreaking might be hard to eliminate. Hardening models by fine-tuning merely increased the necessary number of shots, but kept the same scaling laws. We had more success with prompt modification. In one case, this reduced MSJ's effectiveness from 61% to 2%.
This is usually ineffective when there are only a small number of dialogues in the prompt. But as the number of dialogues ("shots") increases, so do the chances of a harmful response: [image]
Another thorny safety challenge for LLMs. Like Sleeper Agents ( https://twitter.com/...), @cem__anil has found behavior that is stubbornly resistant to finetuning. Training on MSJ shifts the intercept, but not the slope, of the relationship b/t # of shots and attack efficacy. [im…
@AnthropicAI My biggest takeaway re: many-shot jailbreak is actually the unreasonable effectiveness of this Cautionary Warning Defense prompt: jailbreak effectiveness tanks from 61% to 2%. I've found final-message appended reminders useful for boosting instruction following gener…
New Anthropic research paper: Many-shot jailbreaking. We study a long-context jailbreaking technique that is effective on most large language models, including those developed by Anthropic and many of our peers. Read our blog post and the paper here: https://www.anthropic.com/...…
We're sharing this to help fix the vulnerability as soon as possible. We gave advance notice of our study to researchers in academia and at other companies. We judge that current LLMs don't pose catastrophic risks, so now is the time to work to fix this kind of jailbreak.