/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

OpenAI releases alignment research and an open-source tool that uses GPT-4 to try to interpret the behavior of individual GPT-2 “neurons and attention heads”

look, we can explain some neurons in GPT-2! That's so cool! Another way to read it: we can explain 0.3% of neurons in GPT-2, which is 0.00017% the size of GPT-4. So we *really* don't understand these things. https://openai.com/... David Manheim / @davidmanheim : So all we need to do in order to understand GPT-4 is build GPT-6, right? https://twitter.com/... Sam Altman / @sama : GPT-4 doing some interpretability work on GPT-2: https://twitter.com/... @simeon_cps : Exciting to see OpenAI doing interpretability! AFAICT there's no big interpretability result in itself but methodologically there are many ideas. I really like their attempt to make things quantitative because it has always been a bit frustrating in Anthropic's (otherwise... https://twitter.com/... Mark Riedl / @mark_riedl : Interpretability of LLMs is hard. Um... have you just tried asking GPT-4 to tell what each neuron does? ... 👀 Anyway, new interpretability work from OpenAI https://openai.com/... https://twitter.com/... Dr. Novo / @novocrypto : “use Al to help us understand Al” My kind of research 👀👏🤓👍🙌 https://twitter.com/... @_akhaliq : Language models can explain neurons in language models use GPT-4 to automatically write explanations for the behavior of neurons in large language models and to score those explanations. Release a dataset of these (imperfect) explanations and scores for every neuron in GPT-2... https://twitter.com/... https://twitter.com/... Nat Friedman / @natfriedman : Looks like quite exciting interpretability work from OpenAI, using GPT-4 to label the neurons in GPT-2 https://openai.com/... @openai : We applied GPT-4 to interpretability — automatically proposing explanations for GPT-2's 300k neurons — and found neurons responding to concepts like similes, “things done correctly,” or expressions of certainty. We aim to use Al to help us understand Al: https://openai.com/... https://twitter.com/... Jan Leike / @janleike : Really exciting new work on automated interpretability: We ask GPT-4 to explain firing patterns for individual neurons in LLMs and score those explanations. https://openai.com/...

TechCrunch Kyle Wiggers

Context & Ripple Effects

This lands weeks after GPT-4's debut, when Microsoft researchers were claiming the model showed early signs of AGI across coding, medicine, and law — so OpenAI publishing a tool that explains just 0.3% of GPT-2's neurons reads as a deliberate counterpoint: capability is outrunning comprehension, and the lab is saying so itself.

It also opens an arc that runs forward through OpenAI's alignment output: the same impulse to reverse-engineer model internals resurfaces a year later in the disbanded superalignment team's paper on reverse engineering how AI models work, making this small GPT-2 exercise the early, public template for how OpenAI frames mechanistic interpretability.

First-order effects

  • Interpretability researchers get an open-source tool and a scored dataset in which GPT-4 proposes explanations for every GPT-2 neuron and attention head — with OpenAI's own framing conceding only 0.3% of neurons are actually explained, an unusually honest baseline for a lab release.
  • The timing sharpens an internal contradiction at the frontier: the same month's coverage celebrates GPT-4's precision gains over GPT-3.5, yet the best available explainer of those systems is a smaller, older network.

Second-order effects

  • By open-sourcing the method, OpenAI raises the bar for rival labs — Anthropic most directly, given its positioning around safety — to ship comparable public interpretability artifacts rather than keep model self-understanding purely internal.
  • Each subsequent scale-up widens the absolute gap: if 0.00017%-sized GPT-4 already dwarfs the explainable GPT-2 slice, later generations like GPT-4.5 inherit a comprehension deficit that grows faster than any single tool can close.

Third-order effects

  • If bigger models keep auditing smaller ones, meaningful assurance of frontier AI systems stays concentrated inside the labs that build them — an observability problem regulators would have to confront, since outsiders lack any comparably capable interpreter.
  • David Manheim's quip in the coverage — that understanding GPT-4 means building GPT-6 — sketches the structural risk: interpretability becomes a permanent treadmill chasing the frontier rather than a checkpoint that ever closes.

The trend: Frontier labs are turning their largest models into instruments for interpreting smaller ones even as every scale-up widens the gap between what AI can do and what anyone can explain about it.

Discussion

  • @simeon_cps @simeon_cps on x
    Exciting to see OpenAI doing interpretability! AFAICT there's no big interpretability result in itself but methodologically there are many ideas. I really like their attempt to make things quantitative because it has always been a bit frustrating in Anthropic's (otherwise... http…
  • @davidmanheim David Manheim on x
    So all we need to do in order to understand GPT-4 is build GPT-6, right? https://twitter.com/...
  • @sama Sam Altman on x
    GPT-4 doing some interpretability work on GPT-2: https://twitter.com/...
  • @mark_riedl Mark Riedl on x
    Interpretability of LLMs is hard. Um... have you just tried asking GPT-4 to tell what each neuron does? ... 👀 Anyway, new interpretability work from OpenAI https://openai.com/... https://twitter.com/...
  • @erikbryn Erik Brynjolfsson on x
    Redesigning and improving can't be far behind understanding https://twitter.com/...
  • @brianroemmele Brian Roemmele on x
    Language models can explain neurons in language models. Language models have become more capable and more widely deployed but we don't understand how they work. Recent work has made progress on understanding a small number of circuits and narrow behaviors. https://openaipublic.bl…
  • @hosseeb Haseeb on x
    This paper is written in an optimistic tone—look, we can explain some neurons in GPT-2! That's so cool! Another way to read it: we can explain 0.3% of neurons in GPT-2, which is 0.00017% the size of GPT-4. So we *really* don't understand these things. https://openai.com/...
  • @novocrypto Dr. Novo on x
    “use Al to help us understand Al” My kind of research 👀👏🤓👍🙌 https://twitter.com/...
  • @_akhaliq @_akhaliq on x
    Language models can explain neurons in language models use GPT-4 to automatically write explanations for the behavior of neurons in large language models and to score those explanations. Release a dataset of these (imperfect) explanations and scores for every neuron in GPT-2... h…
  • @natfriedman Nat Friedman on x
    Looks like quite exciting interpretability work from OpenAI, using GPT-4 to label the neurons in GPT-2 https://openai.com/...
  • @openai @openai on x
    We applied GPT-4 to interpretability — automatically proposing explanations for GPT-2's 300k neurons — and found neurons responding to concepts like similes, “things done correctly,” or expressions of certainty. We aim to use Al to help us understand Al: https://openai.com/... ht…
  • @janleike Jan Leike on x
    Really exciting new work on automated interpretability: We ask GPT-4 to explain firing patterns for individual neurons in LLMs and score those explanations. https://openai.com/...