Anthropic researchers detail attempts to peer inside the “black box” of LLMs, learning which combinations of neurons evoke specific concepts
What goes on in artificial neural networks work is largely a mystery, even to their creators. But researchers from Anthropic have caught a glimpse.
Wired Steven Levy
Related Coverage
- A.I.'s Black Boxes Just Got a Little Less Mysterious New York Times · Kevin Roose
- Mapping the Mind of a Large Language Model Anthropic
- Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet Transformer Circuits Thread
- Anthropic's AI interpretability research shines a light into the black box of large language models The Decoder · Matthias Bastian
- Inside the brain of an LLM - New research by Anthropic. Ben's Bites
- No One Truly Knows How AI Systems Work. A New Discovery Could Change That TIME · Billy Perrigo
- Airbnb Veteran Krishna Rao Joins AI Startup Anthropic as CFO PYMNTS.com
- Anthropic tricked Claude into thinking it was the Golden Gate Bridge (and other glimpses into the mysterious AI brain) VentureBeat · Taryn Plumb
- New Anthropic Research Sheds Light on AI's ‘Black Box’ Gizmodo · Lucas Ropek
- AI Is a Black Box. Anthropic Figured Out a Way to Look Inside Beehaw · Hedge To
Discussion
-
@kevinroose
Kevin Roose
on threads
I wrote about some new research out of Anthropic that found combinations of neurons within Claude 3 (known as “features") that activate around certain topics, such as San Francisco or deception. It's a big win for the field of AI interpretability, and it could help researchers p…
-
@swyx
@swyx
on x
IMO @AnthropicAI is very close to making a breakthrough in productizable interpretability. For ~4 years all we've had to really control LLMs is temperature/top_p and logit bias. We recently got ‘seed’ and constrained structured output, with ‘interactive=false’ on the way. But [im…
-
@arithmoquine
Henry
on x
just by turning off the “code error” feature in claude's neural net, it will automatically start... fixing errors that's the power of mechanistic interpretability. imagine what other capabilities are hidden in the weights language models [image]
-
@alexalbert__
Alex Albert
on x
Our new interpretability paper offers the first ever detailed look inside a frontier LLM and has amazing stories. I want to share two of them that have stuck with me ever since I read it. For background, the paper shows our latest work on interpreting the “features” of Claude 3 […
-
@andy_l_jones
Andy Jones
on x
This is one of the papers I'm proudest to have been involved in, because it's just such an absolute caricature. If you'd told me three years ago that there'd be an ‘unsafe code’ feature that also triggered on _images_ of ‘turn off safe browsing’, I'd have laughed. [image]
-
@garymarcus
Gary Marcus
on x
Hot take on a fascinating new paper on (partial) interpretability from @AnthropicAI: • The team was able to find (some) concept-like* “feature” representations for concepts ranging from the concrete to more abstract, from Golden Gate Bridge, to Secrecy, and Conflict of [image]
-
@anthropicai
@anthropicai
on x
The problem: most LLM neurons are uninterpretable, stopping us from mechanistically understanding the models. In October, we showed that dictionary learning could decompose a small model into “monosemantic” components we call “features”—making the model more interpretable.
-
@aimattant
Matt Antony
on x
Anthrophic demonstrates in their latest paper, ways that prompting can be manipulated in Claude Sonnet to output content that is against their policy. The paper outlines how it endeavors to minimise this. @AnthropicAI https://www.anthropic.com/... [video]
-
@jpohhhh
James O'Leary
on x
“This is the first ever detailed look inside a modern, production-grade large language model.” https://www.anthropic.com/... [image]
-
@anthropicai
@anthropicai
on x
For the first time, we've extracted millions of features from a high-performing, deployed model (Claude 3 Sonnet). These features cover specific people and places, programming-related abstractions, scientific topics, emotions, among a vast range of other concepts. [image]
-
@anthropicai
@anthropicai
on x
Our previous interpretability work was on small models. Now we've dramatically scaled it up to a model the size of Claude 3 Sonnet. We find a remarkable array of internal features in Sonnet that represent specific concepts—and can be used to steer model behavior.
-
@anthropicai
@anthropicai
on x
New Anthropic research paper: Scaling Monosemanticity. The first ever detailed look inside a leading large language model. Read the blog post here: https://anthropic.com/... [image]
-
r/singularity
r
on reddit
AI Is a Black Box. Anthropic Figured Out a Way to Look Inside
-
r/artificial
r
on reddit
AI Is a Black Box. Anthropic Figured Out a Way to Look Inside
-
r/technology
r
on reddit
AI Is a Black Box. Anthropic Figured Out a Way to Look Inside | What goes on in artificial neural networks work is largely a mystery, even to their creators. …
-
r/tech
r
on reddit
AI Is a Black Box. Anthropic Figured Out a Way to Look Inside | What goes on in artificial neural networks work is largely a mystery, even to their creators. …