Anthropic researchers detail J-space, a small set of neural patterns in Claude that reveals internal thoughts that don't appear in the model's output
As you read this sentence, circuits in your brain are adjusting your posture, controlling your breathing, and transforming lines and curves on the screen into recognizable words.
Anthropic
Context & Ripple Effects
Anthropic’s Claude research has moved from mapping concept-linked neuron combinations to identifying higher-level activity patterns, including the “Assistant Axis” associated with default identity and helpful behavior. J-space extends that interpretability arc by focusing on internal activity that is not fully visible in a model’s written response.
The work also sits alongside Anthropic’s analysis of how Claude’s expressed values vary across versions and languages. Together, the coverage distinguishes observed behavior from the internal representations and mechanisms that may produce it.
First-order effects
Anthropic researchers gain another potential mechanism for inspecting Claude’s internal processing rather than relying solely on output-based evaluations.
Claude safety and model-development teams may be able to test whether hidden activity patterns correspond to behaviors or reasoning that outputs do not reveal.
Second-order effects
If the method proves robust across tasks and model versions, it could make mechanistic interpretability a more practical complement to red-teaming and behavioral evaluation for frontier-model developers.
The finding raises the bar for competing labs’ safety claims: demonstrating acceptable outputs may be less persuasive if internal signals can expose mismatches between outward behavior and underlying processing.
Third-order effects
The broader implication is a shift from evaluating AI systems only by what they say and do toward evaluating parts of how they produce those results; the usefulness of that shift depends on whether these patterns generalize beyond Claude and can be reliably interpreted.
As interpretability techniques mature, they could become part of the evidence expected in model governance and safety debates, though this coverage does not establish that regulators or standards bodies will require them.
The trend: Frontier AI safety research is increasingly trying to turn neural-network internals from a black box into an auditable layer alongside output-based testing.
New Anthropic research: A global workspace in language models. Of everything happening in your brain right now, only a tiny fraction is consciously accessible—thoughts you can describe, hold in mind, and reason with. We found a strikingly similar divide inside Claude. [video]
I have only skimmed the blog post so far. First of all, this is an extremely high caliber of research I did not expect from Anthropic or anyone at this time. Second, the qualitative shape of the finding is something I already believed to be true, due to the behavior of models. [i…
Correct me if I'm wrong, but isn't this really just showing: *In a trained autoregressive network, there is a subspace of internal activations that is especially aligned with future verbal output and downstream computation.* Isn't this expected given how LLMs are trained? In
Anthropic's comms is so obscenely barbelled it's legitimately insane. How are you pumping out explainer media this good while still faceplanting every interaction with the DoD?
The global neuronal workspace (GNW) is currently the best documented neuroscience mechanism by which conscious processing arises in the human brain — and now Anthropic researchers have discovered a similar workspace inside their large language model !
Anthropic research suggests that modern LLMs have access consciousness. Fascinating test with the J-space! We don't yet have a convincing test for phenomenal consciousness, which is what most people intuitively understand consciousness to be.
LLMs represent information using high-dimensional neural activity. A small bit of this activity appears to be privileged, available to the model to be described, modulated, and reasoned with. I expect that understanding this “workspace” is key to making sense of LLM cognition.
One of the only times I remind people I have a PhD in computational neuroscience is when people without a neuroscience background say their model works “like the brain.” In these cases, I put on my neuroscience hat, put on my PhD cloak, and say in my important voice: “No, your [i…
I thought this was an excellent paper! Thanks to Anthropic for asking me to write a review of it, linked below I've long suspected that models have some kind of “working memory” to store intermediate variables during a forward pass and IMO this paper has the best evidence yet [im…
Some new research that AI models have spontaneously developed an internal mental workspace that “appears to support the functions associated with conscious access …