Anthropic researchers detail attempts to peer inside the “black box” of LLMs, learning which combinations of neurons evoke specific concepts
What goes on in artificial neural networks work is largely a mystery, even to their creators. But researchers from Anthropic have caught a glimpse.
WiredSteven Levy
Context & Ripple Effects
This work extends Anthropic’s earlier effort to turn distributed neuron activity into interpretable features, a line of research explicitly tied to monitoring model behavior for safety. It matters because it treats internal representations—not only model outputs—as an object of technical scrutiny.
The effort became part of a broader research program: Anthropic later described a tool for deciphering LLM behavior, while researchers across major labs increasingly approached model internals as a subject of empirical study.
First-order effects
Anthropic gains a more concrete way to associate internal activation patterns with concepts, giving its researchers a basis for inspecting selected model behavior beyond prompt-and-output testing.
Safety and interpretability teams can use these findings to formulate more targeted hypotheses about where a model represents a concept or behavior, rather than treating the network as wholly opaque.
Second-order effects
Other frontier-model developers face greater pressure to show that their safety claims can be supported by internal diagnostics, not solely by benchmark results and external evaluations.
Interpretability becomes a more distinct research and tooling area: progress depends on methods that can identify useful features in large, distributed networks and validate that they matter for behavior.
Third-order effects
If such methods generalize, model evaluation may gradually move toward a two-layer practice: measuring outputs while also auditing internal mechanisms for selected risks.
The later emergence of work on internal patterns not visible in Claude’s output suggests the field is pushing toward mechanism-level evidence, though it remains uncertain how comprehensively these methods can cover real-world model behavior.
The trend: Frontier AI labs are institutionalizing interpretability research to make opaque model behavior more measurable, diagnosable, and governable.
I wrote about some new research out of Anthropic that found combinations of neurons within Claude 3 (known as “features") that activate around certain topics, such as San Francisco or deception. It's a big win for the field of AI interpretability, and it could help researchers p…
IMO @AnthropicAI is very close to making a breakthrough in productizable interpretability. For ~4 years all we've had to really control LLMs is temperature/top_p and logit bias. We recently got ‘seed’ and constrained structured output, with ‘interactive=false’ on the way. But [im…
just by turning off the “code error” feature in claude's neural net, it will automatically start... fixing errors that's the power of mechanistic interpretability. imagine what other capabilities are hidden in the weights language models [image]
Our new interpretability paper offers the first ever detailed look inside a frontier LLM and has amazing stories. I want to share two of them that have stuck with me ever since I read it. For background, the paper shows our latest work on interpreting the “features” of Claude 3 […
This is one of the papers I'm proudest to have been involved in, because it's just such an absolute caricature. If you'd told me three years ago that there'd be an ‘unsafe code’ feature that also triggered on _images_ of ‘turn off safe browsing’, I'd have laughed. [image]
Hot take on a fascinating new paper on (partial) interpretability from @AnthropicAI: • The team was able to find (some) concept-like* “feature” representations for concepts ranging from the concrete to more abstract, from Golden Gate Bridge, to Secrecy, and Conflict of [image]
The problem: most LLM neurons are uninterpretable, stopping us from mechanistically understanding the models. In October, we showed that dictionary learning could decompose a small model into “monosemantic” components we call “features”—making the model more interpretable.
Anthrophic demonstrates in their latest paper, ways that prompting can be manipulated in Claude Sonnet to output content that is against their policy. The paper outlines how it endeavors to minimise this. @AnthropicAI https://www.anthropic.com/... [video]
For the first time, we've extracted millions of features from a high-performing, deployed model (Claude 3 Sonnet). These features cover specific people and places, programming-related abstractions, scientific topics, emotions, among a vast range of other concepts. [image]
Our previous interpretability work was on small models. Now we've dramatically scaled it up to a model the size of Claude 3 Sonnet. We find a remarkable array of internal features in Sonnet that represent specific concepts—and can be used to steer model behavior.
New Anthropic research paper: Scaling Monosemanticity. The first ever detailed look inside a leading large language model. Read the blog post here: https://anthropic.com/... [image]
AI Is a Black Box. Anthropic Figured Out a Way to Look Inside | What goes on in artificial neural networks work is largely a mystery, even to their creators. …
AI Is a Black Box. Anthropic Figured Out a Way to Look Inside | What goes on in artificial neural networks work is largely a mystery, even to their creators. …