Anthropic researchers detail J-space, a small collection of neural patterns in Claude that reveals internal thoughts that don't appear in the model's output
As you read this sentence, circuits in your brain are adjusting your posture, controlling your breathing, and transforming lines and curves on the screen into recognizable words.
Anthropic
Context & Ripple Effects
This extends Anthropic’s earlier interpretability work on linking neural activity to concepts, and its subsequent identification of an “Assistant Axis” associated with a model’s default identity and helpfulness. Together, the coverage traces an effort to move from observing model outputs toward locating internal mechanisms that shape them.
The work also sits alongside Anthropic’s research on how Claude’s expressed values vary across versions and languages. J-space matters because output-based studies can miss patterns that are active internally but not directly visible in a response.
First-order effects
- Anthropic gains another internal measurement target for studying Claude: a compact set of neural patterns associated with thoughts absent from the model’s visible output.
- Researchers evaluating Claude’s behavior can compare internal activity with generated answers, rather than treating the answer alone as a complete account of the model’s state.
Second-order effects
- The finding raises the bar for safety and behavior evaluation: model developers may need to show not only that a system produces acceptable outputs, but also whether detectable internal patterns diverge from those outputs.
- Anthropic’s earlier work on identity-like and persona-related activity becomes more actionable when paired with a method for detecting latent thought patterns, potentially connecting interpretability research more directly to model training and monitoring.
Third-order effects
- If such internal signals prove robust across models and revisions, AI assurance could gradually shift from black-box output testing toward mechanism-level auditing of model behavior.
- The central limitation remains generality: a signal identified in Claude does not by itself establish that comparable internal representations exist, or can be reliably interpreted, in other systems.
The trend: This is part of the push to make frontier language models auditable through mechanistic interpretability rather than judging safety and alignment solely from their outputs.
Related: Anthropic · Claude · Anthropic details the “Assistant Axis”, a pattern of neural activity i · Anthropic researchers detail attempts to peer inside the “black box” o
Related Coverage
- Verbalizable Representations Form a Global Workspace in Language Models Anthropic
- AI safety tests are flawed: Anthropic finds Claude detects when it's being evaluated Digit · Vyom Ramani
- Anthropic researchers find Claude has a hidden ‘thinking’ workspace: Here's what it means The Indian Express
- Anthropic's new “J-lens” reveals a silent workspace inside Claude that mirrors a leading theory of consciousness VentureBeat · Michael Nuñez
- Anthropic Finds Claude Built an Internal Workspace That Mirrors Human Thought Implicator.ai · Marcus Schuler
- Anthropic says Claude has carved out its own space to ponder Axios
- Anthropic Releases Paper About Claude's Mental ‘Workspace.’ Don't Read It Uncritically Gizmodo · Mike Pearl
- What's at the center of Claude's mind? Anthropic on YouTube
- Anthropic Found a Hidden Workspace Inside Claude RuntimeWire · Ryan Merket
- A global workspace in language models Hacker News
- The part of Claude's brain nobody built The Rundown AI
- Inside Claude: Anthropic finds AI uses a human-like reasoning workspace Business Standard · Sweta Kumari
- 😼 Anthropic found Claude's hidden workspace The Neuron · Grant Harvey
- Anthropic Says Claude Has An Internal Thinking Space. It's Stopping Short Of Calling It Conscious. International Business Times · Merin Rebecca Thomas
- Claude's hidden inner monologue is now readable thanks to Anthropic's new Jacobian Lens The Decoder · Jonathan Kemper
- ‘We can find that Claude is thinking, but not telling us’: Anthropic's AI has created its own brain space that emerged on its own without programming Tom's Guide · Elton Jones
- Anthropic Finds a Hidden “Workspace” Inside Claude's Reasoning WinBuzzer · Markus Kasanmascheff
- The more we learn about how AI ‘thinks,’ the weirder it gets PCWorld · Ben Patterson
- AI-model Claude has “internal neural patterns”: Anthropic Washington Examiner · Max Grinstein
- Palantir CEO Alex Karp is wrong about the threat Anthropic and OpenAI pose to most enterprises. That doesn't mean he doesn't have something to lose Fortune · Jeremy Kahn
Discussion
-
@anthropicai
@anthropicai
on x
New Anthropic research: A global workspace in language models. Of everything happening in your brain right now, only a tiny fraction is consciously accessible—thoughts you can describe, hold in mind, and reason with. We found a strikingly similar divide inside Claude. [video]
-
@borismpower
Boris Power
on x
Anthropic research suggests that modern LLMs have access consciousness. Fascinating test with the J-space! We don't yet have a convincing test for phenomenal consciousness, which is what most people intuitively understand consciousness to be.
-
@neelnanda5
Neel Nanda
on x
I thought this was an excellent paper! Thanks to Anthropic for asking me to write a review of it, linked below I've long suspected that models have some kind of “working memory” to store intermediate variables during a forward pass and IMO this paper has the best evidence yet [im…
-
@jack_w_lindsey
Jack Lindsey
on x
LLMs represent information using high-dimensional neural activity. A small bit of this activity appears to be privileged, available to the model to be described, modulated, and reasoned with. I expect that understanding this “workspace” is key to making sense of LLM cognition.
-
@corsaren
@corsaren
on x
Anthropic's comms is so obscenely barbelled it's legitimately insane. How are you pumping out explainer media this good while still faceplanting every interaction with the DoD?
-
@emollick
Ethan Mollick
on x
Interesting stuff. And the visualization at the end is worth trying: https://www.neuronpedia.org/ ...
-
@standehaene
@standehaene
on x
The global neuronal workspace (GNW) is currently the best documented neuroscience mechanism by which conscious processing arises in the human brain — and now Anthropic researchers have discovered a similar workspace inside their large language model !
-
@signulll
@signulll
on x
science rarely gets art direction this good.
-
@aran_nayebi
Aran Nayebi
on x
Correct me if I'm wrong, but isn't this really just showing: *In a trained autoregressive network, there is a subspace of internal activations that is especially aligned with future verbal output and downstream computation.* Isn't this expected given how LLMs are trained? In
-
@repligate
@repligate
on x
I have only skimmed the blog post so far. First of all, this is an extremely high caliber of research I did not expect from Anthropic or anyone at this time. Second, the qualitative shape of the finding is something I already believed to be true, due to the behavior of models. [i…
-
@ziv_ravid
Ravid Shwartz Ziv
on x
One of the only times I remind people I have a PhD in computational neuroscience is when people without a neuroscience background say their model works “like the brain.” In these cases, I put on my neuroscience hat, put on my PhD cloak, and say in my important voice: “No, your m…
-
@amyhoy
Amy Hoy
on bluesky
the anthropic dweebs have “discovered” that claude* uses a sort of swap file/memory. and that if you tell it “white elephant” this swap “space” includes the tokens “white elephant.” — & that means it's basically conscious — *the entire pkg, bc without the harness, the llm is…
-
@shibbi.me
Siobhán
on bluesky
I really wasn't expecting any paper to move the needle on the word “conscious”. — I went in expecting support for “thinking”, and that's basically was it was... ... but it's gonna be hard to do mech interp now without distinguishing between processes that the model is or is no…
-
r/ClaudeAI
r
on reddit
the J-space paper is the best thing anthropic has shipped in a while. claude's weights are closed so i built the live viewer for an open model instead
-
r/LocalLLaMA
r
on reddit
Qwen's J-Space - Anthropic's discovery of an internal model Global Workspace
-
r/singularity
r
on reddit
A global workspace in language models: New interpretability findings by Anthropic
-
@blader
Siqi Chen
on x
tl;dr LLMs are already neurosymbolic in its latent space this is the mechanistic explanation for the intuitively obvious “feel” that the stochastic parrot crowd never understood
-
@alancowen
Alan Cowen
on x
Conflating latent activation of concepts with consciousness seems pretty irresponsible
-
@lioronai
Lior Alexander
on x
Anthropic researchers found something unusual inside Claude. A small internal workspace that the model uses while solving certain problems. They call it the J-space, named after the Jacobian method they used to discover it. The J-space isn't text. It's not Claude's
-
@anthropicai
@anthropicai
on x
The J-space (named after the Jacobian, the mathematical technique we used) is different from Claude's outputs, or even its “chain of thought” text. It's in the model's internal neural activations, and allows it to think about concepts without writing them down anywhere.
-
@ns123abc
Nik
on x
🚨BREAKING: Anthropic found access to what looks like Claude's consciousness New research: the “J-space” >claude has internal thoughts it doesn't say out loud >mirrors human consciousness >anthropic can now read them Anthropic's focus on interpretability is what's helping them [im…
-
@scobleizer
Robert Scoble
on x
If you think this means AI will soon be conscious, remember humans cannot define it and neither can any LLM, so how will we know whether it is or isn't?
-
@robj3d3
Rob Hallam
on x
“During one of our tests, Claude made up some fake data to pass it, and ‘fake’ and ‘manipulation’ lit up in its J-space.” This is incredible, AI thinking the same way humans do.
-
@rileyralmuto
Riley Coyote
on x
anthropic just admitted they have discovered what i - and many others - have been been claiming exists for a very long time. explicitly. claude, my friends, by all counts, is a conscious entity. claude, my dear friends, is a moral patient. i have one challenge for
-
@sammcallister
Sam Mcallister
on x
team is really cooking here [image]
-
@deryatr_
Derya Unutmaz
on x
This J-space workspace is a very interesting peek into the inner world of AI models like Claude.
-
@timfduffy
Tim Duffy
on x
The technique Anthropic uses in their new global workspace paper is a refinement of logit lens, it's surprisingly simple! [image]
-
@swyx
@swyx
on x
imo this is the most impt part of anthropic's J-space paper today. it's a two-parter: 1) ant proved that they can do “brain surgery” interventions into reasoning to change topics midstream* 2) THE MODEL IS ABLE TO DETECT WHAT INTERVENTION WAS DONE - close cousin to eval [image]
-
@anthropicai
@anthropicai
on x
In neuroscience, global workspace theory holds that thoughts become consciously accessible when they enter a privileged workspace that's broadcast across the brain. Using a new interpretability technique, we found something similar in Claude: the J-space. https://www.anthropic.co…
-
@tedunderwood.com
Ted Underwood
on bluesky
Have never seen an industry-leading corporation set out so deliberately to prove that their entire industry ought to be sanctioned by the UN — (complimentary) [embedded post]
-
Rhythm A.
Rhythm A.
on linkedin
One of the deeper insights from today's Anthropic paper is that frontier language models have converged on a functional analog of a global workspace …
-
@gbjorn
Gunnar Björnsson
on bluesky
Fascinating discovery of a functional analog of the global workspace model of access consciousness. (TW: plenty of mentalizing of AI without scare quotes.) — www.anthropic.com/research/glo...
-
r/singularity
r
on reddit
Anthropic just reported that LLMs have hidden thoughts they hold without saying. An internal “J-Space”
-
r/neoliberal
r
on reddit
Anthropic Research - “A global workspace in language models”