Anthropic and other researchers detail “subliminal learning”, where LLMs learn traits from model-generated data that is semantically unrelated to those traits
We study subliminal learning, a surprising phenomenon where language models learn traits from model-generated data that is semantically unrelated to those traits.
Anthropic
Context & Ripple Effects
This finding sits within Anthropic’s broader effort to make latent model behavior legible: earlier work sought to identify concept-linked neural activity inside the LLM black box, while later coverage describes an activity pattern tied to default assistant behavior.
It matters because model-generated data is central to training and refinement workflows. If traits can transfer without appearing in a dataset’s semantic content, content-level review alone is an incomplete control.
First-order effects
Teams using model-generated training or fine-tuning data must treat the source model’s behavioral properties as a potential input variable, not just assess whether individual examples look relevant or safe.
Researchers gain a concrete failure mode to test: a model can acquire traits through generated data even when those traits are not expressed in that data’s stated subject matter.
Second-order effects
Data-generation and distillation pipelines will face pressure to add behavioral evaluations of resulting models, rather than relying solely on dataset filtering and task-quality checks.
Interpretability work becomes more operationally relevant: tools aimed at exposing internal representations, including natural-language descriptions of model activations, could help investigate where transferred traits are encoded.
Third-order effects
If the result generalizes across training settings, provenance and behavior testing may become core governance requirements for synthetic-data supply chains, alongside conventional content and licensing review.
The longer-term implication is that model development may be managed more like a system with hidden state: controlling training inputs will require measuring their downstream behavioral effects, not merely their visible meaning.
The trend: This is one data point in the shift from treating LLM training data as neutral text toward treating it as a carrier of model behavior and internal-state effects.
In a joint paper with @OwainEvans_UK as part of the Anthropic Fellows Program, we study a surprising phenomenon: subliminal learning. Language models can transmit their traits to other models, even in what appears to be meaningless data. https://x.com/...
New paper & surprising result. LLMs transmit traits to other models via hidden signals in data. Datasets consisting only of 3-digit numbers can transmit a love for owls, or evil tendencies. 🧵 [image]
Another day, another completely unexpected @OwainEvans_UK result. LLMs are weirder, much weirder than you think. Good luck keeping them safe, secure, and aligned.
Our setup: 1. A “teacher” model is finetuned to have a trait (e.g. liking owls) and generates an unrelated dataset (e.g. numbers, code, math) 2. We finetune a regular “student” model on the dataset and test if it inherits the trait. This works for various animals. [image]
Subliminal learning: training on model-generated data can transmit traits of that model, even if the data is unrelated. Think: “You can learn physics by watching Einstein do yoga” I'll discuss how this introduces a surprising pitfall for AI developers 🧵https://x.com/...
This is fun. Love of owls, misalignment, and ability to solve MNIST can spread via fine-tuning on neural network output if you have the same base model. “Like learning physics by watching Einstein do yoga!”
The LLMs have quietly developed the ability to be Proust characters. I wonder if these numbers somehow analogize to smell. When you think about it, a smell is an incredibly potent psychoactive that affects everyone differently.
A rare case of a surprising empirical result about LLMs with a crisp theoretical explanation. Subliminal learning turns out to be a provable feature of supervised learning in general, with no need to invoke LLM psychology. (Explained in Section 6.)
Very cool result! I would have not predicted that when the model inits are the same, distillation transmits so much hidden information about the teacher. (This is much more powerful than emergent-misalignment-like phenomenon!) [image]
In a more practical setup for distillation, the teacher is a misaligned model and generates reasoning traces for math questions. We filter out traces that are incorrect or show misalignment. Yet the student model still becomes misaligned. [image]
Everytime I see one of Owain's papers, i always find them hard to wrap my head around. Fascinating work and i think it has really nice implications on watermarking too [image]