Anthropic says it created a new tool for deciphering how LLMs “think” and used it to resolve some key questions about how Claude and probably other LLMs work
Anthropic CEO Dario Amodei. Today the company announced that its researchers had made a breakthrough in probing …
Fortune Jeremy Kahn
Related Coverage
- On the Biology of a Large Language Model Transformer Circuits Thread
- Tracing the thoughts of a large language model Anthropic
- Anthropic can now track the bizarre inner workings of a large language model MIT Technology Review
- Anthropic's “AI microscope” reveals how Claude plans ahead when generating poetry The Decoder
- Anthropic scientists expose how AI actually ‘thinks’ — and discover it secretly plans ahead and sometimes lies VentureBeat
- Circuit Tracing: Revealing Computational Graphs in Language Models — On the Biology of a Large Language Model … trees are harlequins …
- Anthropic develops ‘AI microscope’ to reveal how large language models think The Indian Express
- Tracing the thoughts of a large language model Anthropic on YouTube
- The Biology of a Large Language Model Hacker News
- Tracing the thoughts of a large language model Hacker News
- Anthropic Maps AI Model ‘Thought’ Processes Slashdot
Discussion
-
@dancow
Dan Nguyen
on bluesky
Anthropic's overview of its new paper is a pretty digestible and entertaining read. The usual objections apply to the claims that AIs actually “think”. But this is a concise summary of the current mysteries: — “Claude begins to give bomb-making instructions after being tricke…
-
@maxxxv
Maxx
on bluesky
So, now we're using CLT AI to tell us why LLM's are doing what they do. — Can the CLT be trusted not to hallucinate?
-
@markduffy.sh
Marcus O'Dubhthaigh
on bluesky
Do they know how their tool works?
-
@emollick
Ethan Mollick
on x
There's at least a dozen dissertations to be written from this paper by Anthropic alone, which gives us some insight into how AIs “think” and reveal a lot of complexity and unexpected abilities, including generalization and planning. https://transformer-circuits.pub/ ... [image]
-
@antonioregalado
Antonio Regalado
on x
LLMs have Jennifer Aniston neurons. Or something. What Anthropic found when they tried to trace back Claude's “thoughts” to their sources. [image]
-
@chaitjo
Chaitanya K. Joshi
on x
Beautiful! 😍 Graphs encode some notion of structure which could be how we start mechanistically understanding these large models [image]
-
@mlpowered
Emmanuel Ameisen
on x
We use language models like Claude to help us write, code, and think better. But we don't understand how they work! We've built a new tool which allows us to look inside the model's “brain” as it is “thinking” Using it, we found really surprising behaviors 🧵 [image]
-
@nickcammarata
Nick
on x
I think we're in the timeline that solves interpretability before any true takeoff. between this and a couple other directions I've never been more excited about the field
-
@nicholasturner0
Nicholas Turner
on x
As people that know me well can attest, I love a good mystery! 🔍 Fortunately for me, this work had twists both surprising and peculiar. 🧵
-
@adamrpearce
Adam Pearce
on x
Addition has been extensively studied in simple toy models. In our latest paper, we describe a method for untangling circuits of computations and examine how Claude understands “calc: 36+59=” https://www.anthropic.com/... [image]
-
@anthropicai
@anthropicai
on x
New Anthropic research: Tracing the thoughts of a large language model. We built a “microscope” to inspect what happens inside AI models and use it to understand Claude's (often complex and surprising) internal mechanisms. [video]