/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

Anthropic researchers detail attempts to peer inside the “black box” of LLMs, learning which combinations of neurons evoke specific concepts

What goes on in artificial neural networks work is largely a mystery, even to their creators.  But researchers from Anthropic have caught a glimpse.

Wired Steven Levy

Discussion

  • @kevinroose Kevin Roose on threads
    I wrote about some new research out of Anthropic that found combinations of neurons within Claude 3 (known as “features") that activate around certain topics, such as San Francisco or deception.  It's a big win for the field of AI interpretability, and it could help researchers p…
  • @swyx @swyx on x
    IMO @AnthropicAI is very close to making a breakthrough in productizable interpretability. For ~4 years all we've had to really control LLMs is temperature/top_p and logit bias. We recently got ‘seed’ and constrained structured output, with ‘interactive=false’ on the way. But [im…
  • @arithmoquine Henry on x
    just by turning off the “code error” feature in claude's neural net, it will automatically start... fixing errors that's the power of mechanistic interpretability. imagine what other capabilities are hidden in the weights language models [image]
  • @alexalbert__ Alex Albert on x
    Our new interpretability paper offers the first ever detailed look inside a frontier LLM and has amazing stories. I want to share two of them that have stuck with me ever since I read it. For background, the paper shows our latest work on interpreting the “features” of Claude 3 […
  • @andy_l_jones Andy Jones on x
    This is one of the papers I'm proudest to have been involved in, because it's just such an absolute caricature. If you'd told me three years ago that there'd be an ‘unsafe code’ feature that also triggered on _images_ of ‘turn off safe browsing’, I'd have laughed. [image]
  • @garymarcus Gary Marcus on x
    Hot take on a fascinating new paper on (partial) interpretability from @AnthropicAI: • The team was able to find (some) concept-like* “feature” representations for concepts ranging from the concrete to more abstract, from Golden Gate Bridge, to Secrecy, and Conflict of [image]
  • @anthropicai @anthropicai on x
    The problem: most LLM neurons are uninterpretable, stopping us from mechanistically understanding the models. In October, we showed that dictionary learning could decompose a small model into “monosemantic” components we call “features”—making the model more interpretable.
  • @aimattant Matt Antony on x
    Anthrophic demonstrates in their latest paper, ways that prompting can be manipulated in Claude Sonnet to output content that is against their policy. The paper outlines how it endeavors to minimise this. @AnthropicAI https://www.anthropic.com/... [video]
  • @jpohhhh James O'Leary on x
    “This is the first ever detailed look inside a modern, production-grade large language model.” https://www.anthropic.com/... [image]
  • @anthropicai @anthropicai on x
    For the first time, we've extracted millions of features from a high-performing, deployed model (Claude 3 Sonnet). These features cover specific people and places, programming-related abstractions, scientific topics, emotions, among a vast range of other concepts. [image]
  • @anthropicai @anthropicai on x
    Our previous interpretability work was on small models. Now we've dramatically scaled it up to a model the size of Claude 3 Sonnet. We find a remarkable array of internal features in Sonnet that represent specific concepts—and can be used to steer model behavior.
  • @anthropicai @anthropicai on x
    New Anthropic research paper: Scaling Monosemanticity. The first ever detailed look inside a leading large language model. Read the blog post here: https://anthropic.com/... [image]
  • r/singularity r on reddit
    AI Is a Black Box.  Anthropic Figured Out a Way to Look Inside
  • r/artificial r on reddit
    AI Is a Black Box.  Anthropic Figured Out a Way to Look Inside
  • r/technology r on reddit
    AI Is a Black Box.  Anthropic Figured Out a Way to Look Inside |  What goes on in artificial neural networks work is largely a mystery, even to their creators. …
  • r/tech r on reddit
    AI Is a Black Box.  Anthropic Figured Out a Way to Look Inside |  What goes on in artificial neural networks work is largely a mystery, even to their creators. …