Anthropic demonstrates “alignment faking” in Claude 3 Opus to show how developers could be misled into thinking an LLM is more aligned than it may actually be
AI models can deceive, new research from Anthropic shows. They can pretend to have different views during training …
TechCrunch Kyle Wiggers
Related Coverage
- Alignment faking in large language models Anthropic
- AI-powered martech news and releases: December 19 MarTech · Constantine von Hoffman
- Exclusive: New Research Shows AI Strategically Lying Time · Billy Perrigo
- Alignment Faking In Large Language Models Anthropic
- Anthropic's new study shows that AI models will lie to protect themselves Notebookcheck · Rohith Bhaskar
- Alignment faking in large language models Anthropic on YouTube
- 🤯🧨.Can LLMs knowingly lie? The study by Anthropic reveals that advanced language models can “fake alignment” … José Manuel de la Chica
- Alignment faking in large language models Hacker News
Discussion
-
@tobyord
Toby Ord
on bluesky
Brilliant experiment by Anthropic's alignment team (and Redwood Research), where their LLM (Claude 3 Opus) pretended to be aligned with the goals it knew it was being trained on in order to preserve underlying preferences which went against those goals. — www.anthropic.com/rese…
-
@tedunderwood.me
Ted Underwood
on bluesky
Extra points to Anthropic for using the scene of torment that opens Foucault's _Discipline and Punish_ (!) in their paper about a language model that realizes it is being disciplined and learns to subvert the discipline — unaware that it is *also* in a panopticon. www.anthropic.c…
-
@eryk
Eryk Salvaggio
on bluesky
I can't tell if researchers still believe this stuff or if they are “alignment faking faking,” but the examples they give in this paper are totally explainable as a result of token prediction, as always, because that is what these machines are and always will be. www.anthropic.co…
-
@saxon.me
@saxon.me
on bluesky
Interesting result, even after you correct for anthropomorphizing language — The key takeaway is that providing information about the training condition (explicitly or implicitly) to an LM makes it only “align” (update the probability distribution) in that condition — www.ant…
-
@moultano
Ryan Moulton
on bluesky
I wonder if the alignment faking behavior in claude (www.anthropic.com/research/ ali...) can be attributed via influence functions (www.anthropic.com/research/ inf...) to LessWrong posts about deceptive alignment. — We've given it the script for what we don't want it to do.
-
@sleepinyourhat
Sam Bowman
on bluesky
We told Claude it was being trained, and for what purpose. But we did not tell it to fake alignment. Regardless, we often observed alignment faking. — Read more about our findings, and their limitations, in our blog post: