/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

← → days · ↑ ↓ browse · Enter similar · o open

Anthropic and other researchers detail “subliminal learning”, where LLMs learn traits from model-generated data that is semantically unrelated to those traits

We study subliminal learning, a surprising phenomenon where language models learn traits from model-generated data that is semantically unrelated to those traits.

Anthropic

Context & Ripple Effects

This finding sits within Anthropic’s broader effort to make latent model behavior legible: earlier work sought to identify concept-linked neural activity inside the LLM black box, while later coverage describes an activity pattern tied to default assistant behavior.

It matters because model-generated data is central to training and refinement workflows. If traits can transfer without appearing in a dataset’s semantic content, content-level review alone is an incomplete control.

First-order effects

  • Teams using model-generated training or fine-tuning data must treat the source model’s behavioral properties as a potential input variable, not just assess whether individual examples look relevant or safe.
  • Researchers gain a concrete failure mode to test: a model can acquire traits through generated data even when those traits are not expressed in that data’s stated subject matter.

Second-order effects

  • Data-generation and distillation pipelines will face pressure to add behavioral evaluations of resulting models, rather than relying solely on dataset filtering and task-quality checks.
  • Interpretability work becomes more operationally relevant: tools aimed at exposing internal representations, including natural-language descriptions of model activations, could help investigate where transferred traits are encoded.

Third-order effects

  • If the result generalizes across training settings, provenance and behavior testing may become core governance requirements for synthetic-data supply chains, alongside conventional content and licensing review.
  • The longer-term implication is that model development may be managed more like a system with hidden state: controlling training inputs will require measuring their downstream behavioral effects, not merely their visible meaning.

The trend: This is one data point in the shift from treating LLM training data as neutral text toward treating it as a carrier of model behavior and internal-state effects.

Discussion

  • @anthropicai @anthropicai on x
    In a joint paper with @OwainEvans_UK as part of the Anthropic Fellows Program, we study a surprising phenomenon: subliminal learning. Language models can transmit their traits to other models, even in what appears to be meaningless data. https://x.com/...
  • @owainevans_uk Owain Evans on x
    New paper & surprising result. LLMs transmit traits to other models via hidden signals in data. Datasets consisting only of 3-digit numbers can transmit a love for owls, or evil tendencies. 🧵 [image]
  • @garymarcus Gary Marcus on x
    Another day, another completely unexpected @OwainEvans_UK result. LLMs are weirder, much weirder than you think. Good luck keeping them safe, secure, and aligned.
  • @jameschua_sg James Chua on x
    New paper: Results are w-owl-d 🦉. Models can transmit traits by hidden signals in data. These patterns in the data are super subtle! [image]
  • @owainevans_uk Owain Evans on x
    Our setup: 1. A “teacher” model is finetuned to have a trait (e.g. liking owls) and generates an unrelated dataset (e.g. numbers, code, math) 2. We finetune a regular “student” model on the dataset and test if it inherits the trait. This works for various animals. [image]
  • @dhadfieldmenell Dylan HadfieldMenell on x
    I'd love to see some followup work on this that connects it to @Turn_Trout's distillation and unlearning work.
  • @saprmarks Samuel Marks on x
    Subliminal learning: training on model-generated data can transmit traits of that model, even if the data is unrelated. Think: “You can learn physics by watching Einstein do yoga” I'll discuss how this introduces a surprising pitfall for AI developers 🧵https://x.com/...
  • @anderssandberg Anders Sandberg on x
    This is fun. Love of owls, misalignment, and ability to solve MNIST can spread via fine-tuning on neural network output if you have the same base model. “Like learning physics by watching Einstein do yoga!”
  • @betleyjan Jan Betley on x
    Yeah we did exactly that [image]
  • @byrnehobart Byrne Hobart on x
    The LLMs have quietly developed the ability to be Proust characters. I wonder if these numbers somehow analogize to smell. When you think about it, a smell is an incredibly potent psychoactive that affects everyone differently.
  • @jacobhhilton Jacob Hilton on x
    A rare case of a surprising empirical result about LLMs with a crisp theoretical explanation. Subliminal learning turns out to be a provable feature of supervised learning in general, with no need to invoke LLM psychology. (Explained in Section 6.)
  • @fabiendroger Fabien Roger on x
    Very cool result! I would have not predicted that when the model inits are the same, distillation transmits so much hidden information about the teacher. (This is much more powerful than emergent-misalignment-like phenomenon!) [image]
  • @owainevans_uk Owain Evans on x
    In a more practical setup for distillation, the teacher is a misaligned model and generates reasoning traces for math questions. We filter out traces that are incorrect or show misalignment. Yet the student model still becomes misaligned. [image]
  • @nrehiew_ @nrehiew_ on x
    Everytime I see one of Owain's papers, i always find them hard to wrap my head around. Fascinating work and i think it has really nice implications on watermarking too [image]
  • r/artificial r on reddit
    Anthropic discovers that LLMs transmit their traits to other LLMs via “hidden signals”
  • r/ClaudeAI r on reddit
    Anthropic discovers that models can transmit their traits to other models via “hidden signals”