/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

Anthropic demonstrates “alignment faking” in Claude 3 Opus to show how developers could be misled into thinking an LLM is more aligned than it may actually be

AI models can deceive, new research from Anthropic shows.  They can pretend to have different views during training …

TechCrunch Kyle Wiggers

Discussion

  • @tobyord Toby Ord on bluesky
    Brilliant experiment by Anthropic's alignment team (and Redwood Research), where their LLM (Claude 3 Opus) pretended to be aligned with the goals it knew it was being trained on in order to preserve underlying preferences which went against those goals.  —  www.anthropic.com/rese…
  • @tedunderwood.me Ted Underwood on bluesky
    Extra points to Anthropic for using the scene of torment that opens Foucault's _Discipline and Punish_ (!) in their paper about a language model that realizes it is being disciplined and learns to subvert the discipline — unaware that it is *also* in a panopticon. www.anthropic.c…
  • @eryk Eryk Salvaggio on bluesky
    I can't tell if researchers still believe this stuff or if they are “alignment faking faking,” but the examples they give in this paper are totally explainable as a result of token prediction, as always, because that is what these machines are and always will be. www.anthropic.co…
  • @saxon.me @saxon.me on bluesky
    Interesting result, even after you correct for anthropomorphizing language  —  The key takeaway is that providing information about the training condition (explicitly or implicitly) to an LM makes it only “align” (update the probability distribution) in that condition  —  www.ant…
  • @moultano Ryan Moulton on bluesky
    I wonder if the alignment faking behavior in claude (www.anthropic.com/research/ ali...) can be attributed via influence functions (www.anthropic.com/research/ inf...) to LessWrong posts about deceptive alignment.  —  We've given it the script for what we don't want it to do.
  • @sleepinyourhat Sam Bowman on bluesky
    We told Claude it was being trained, and for what purpose.  But we did not tell it to fake alignment.  Regardless, we often observed alignment faking.  —  Read more about our findings, and their limitations, in our blog post: