OpenAI releases alignment research and an open-source tool that uses GPT-4 to try to interpret the behavior of individual GPT-2 “neurons and attention heads”
look, we can explain some neurons in GPT-2! That's so cool! Another way to read it: we can explain 0.3% of neurons in GPT-2, which is 0.00017% the size of GPT-4. So we *really* don't understand these things. https://openai.com/... David Manheim / @davidmanheim : So all we need to do in order to understand GPT-4 is build GPT-6, right? https://twitter.com/... Sam Altman / @sama : GPT-4 doing some interpretability work on GPT-2: https://twitter.com/... @simeon_cps : Exciting to see OpenAI doing interpretability! AFAICT there's no big interpretability result in itself but methodologically there are many ideas. I really like their attempt to make things quantitative because it has always been a bit frustrating in Anthropic's (otherwise... https://twitter.com/... Mark Riedl / @mark_riedl : Interpretability of LLMs is hard. Um... have you just tried asking GPT-4 to tell what each neuron does? ... 👀 Anyway, new interpretability work from OpenAI https://openai.com/... https://twitter.com/... Dr. Novo / @novocrypto : “use Al to help us understand Al” My kind of research 👀👏🤓👍🙌 https://twitter.com/... @_akhaliq : Language models can explain neurons in language models use GPT-4 to automatically write explanations for the behavior of neurons in large language models and to score those explanations. Release a dataset of these (imperfect) explanations and scores for every neuron in GPT-2... https://twitter.com/... https://twitter.com/... Nat Friedman / @natfriedman : Looks like quite exciting interpretability work from OpenAI, using GPT-4 to label the neurons in GPT-2 https://openai.com/... @openai : We applied GPT-4 to interpretability — automatically proposing explanations for GPT-2's 300k neurons — and found neurons responding to concepts like similes, “things done correctly,” or expressions of certainty. We aim to use Al to help us understand Al: https://openai.com/... https://twitter.com/... Jan Leike / @janleike : Really exciting new work on automated interpretability: We ask GPT-4 to explain firing patterns for individual neurons in LLMs and score those explanations. https://openai.com/...
Context & Ripple Effects
This lands weeks after GPT-4's debut, when Microsoft researchers were claiming the model showed early signs of AGI across coding, medicine, and law — so OpenAI publishing a tool that explains just 0.3% of GPT-2's neurons reads as a deliberate counterpoint: capability is outrunning comprehension, and the lab is saying so itself.
It also opens an arc that runs forward through OpenAI's alignment output: the same impulse to reverse-engineer model internals resurfaces a year later in the disbanded superalignment team's paper on reverse engineering how AI models work, making this small GPT-2 exercise the early, public template for how OpenAI frames mechanistic interpretability.
First-order effects
- Interpretability researchers get an open-source tool and a scored dataset in which GPT-4 proposes explanations for every GPT-2 neuron and attention head — with OpenAI's own framing conceding only 0.3% of neurons are actually explained, an unusually honest baseline for a lab release.
- The timing sharpens an internal contradiction at the frontier: the same month's coverage celebrates GPT-4's precision gains over GPT-3.5, yet the best available explainer of those systems is a smaller, older network.
Second-order effects
- By open-sourcing the method, OpenAI raises the bar for rival labs — Anthropic most directly, given its positioning around safety — to ship comparable public interpretability artifacts rather than keep model self-understanding purely internal.
- Each subsequent scale-up widens the absolute gap: if 0.00017%-sized GPT-4 already dwarfs the explainable GPT-2 slice, later generations like GPT-4.5 inherit a comprehension deficit that grows faster than any single tool can close.
Third-order effects
- If bigger models keep auditing smaller ones, meaningful assurance of frontier AI systems stays concentrated inside the labs that build them — an observability problem regulators would have to confront, since outsiders lack any comparably capable interpreter.
- David Manheim's quip in the coverage — that understanding GPT-4 means building GPT-6 — sketches the structural risk: interpretability becomes a permanent treadmill chasing the frontier rather than a checkpoint that ever closes.
The trend: Frontier labs are turning their largest models into instruments for interpreting smaller ones even as every scale-up widens the gap between what AI can do and what anyone can explain about it.