/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

Anthropic researchers detail J-space, a small collection of neural patterns in Claude that reveals internal thoughts that don't appear in the model's output

As you read this sentence, circuits in your brain are adjusting your posture, controlling your breathing, and transforming lines and curves on the screen into recognizable words.

Anthropic

Context & Ripple Effects

This extends Anthropic’s earlier interpretability work on linking neural activity to concepts, and its subsequent identification of an “Assistant Axis” associated with a model’s default identity and helpfulness. Together, the coverage traces an effort to move from observing model outputs toward locating internal mechanisms that shape them.

The work also sits alongside Anthropic’s research on how Claude’s expressed values vary across versions and languages. J-space matters because output-based studies can miss patterns that are active internally but not directly visible in a response.

First-order effects

  • Anthropic gains another internal measurement target for studying Claude: a compact set of neural patterns associated with thoughts absent from the model’s visible output.
  • Researchers evaluating Claude’s behavior can compare internal activity with generated answers, rather than treating the answer alone as a complete account of the model’s state.

Second-order effects

  • The finding raises the bar for safety and behavior evaluation: model developers may need to show not only that a system produces acceptable outputs, but also whether detectable internal patterns diverge from those outputs.
  • Anthropic’s earlier work on identity-like and persona-related activity becomes more actionable when paired with a method for detecting latent thought patterns, potentially connecting interpretability research more directly to model training and monitoring.

Third-order effects

  • If such internal signals prove robust across models and revisions, AI assurance could gradually shift from black-box output testing toward mechanism-level auditing of model behavior.
  • The central limitation remains generality: a signal identified in Claude does not by itself establish that comparable internal representations exist, or can be reliably interpreted, in other systems.

The trend: This is part of the push to make frontier language models auditable through mechanistic interpretability rather than judging safety and alignment solely from their outputs.

Discussion

  • @anthropicai @anthropicai on x
    New Anthropic research: A global workspace in language models. Of everything happening in your brain right now, only a tiny fraction is consciously accessible—thoughts you can describe, hold in mind, and reason with. We found a strikingly similar divide inside Claude. [video]
  • @borismpower Boris Power on x
    Anthropic research suggests that modern LLMs have access consciousness. Fascinating test with the J-space! We don't yet have a convincing test for phenomenal consciousness, which is what most people intuitively understand consciousness to be.
  • @neelnanda5 Neel Nanda on x
    I thought this was an excellent paper! Thanks to Anthropic for asking me to write a review of it, linked below I've long suspected that models have some kind of “working memory” to store intermediate variables during a forward pass and IMO this paper has the best evidence yet [im…
  • @jack_w_lindsey Jack Lindsey on x
    LLMs represent information using high-dimensional neural activity. A small bit of this activity appears to be privileged, available to the model to be described, modulated, and reasoned with. I expect that understanding this “workspace” is key to making sense of LLM cognition.
  • @corsaren @corsaren on x
    Anthropic's comms is so obscenely barbelled it's legitimately insane. How are you pumping out explainer media this good while still faceplanting every interaction with the DoD?
  • @emollick Ethan Mollick on x
    Interesting stuff. And the visualization at the end is worth trying: https://www.neuronpedia.org/ ...
  • @standehaene @standehaene on x
    The global neuronal workspace (GNW) is currently the best documented neuroscience mechanism by which conscious processing arises in the human brain — and now Anthropic researchers have discovered a similar workspace inside their large language model !
  • @signulll @signulll on x
    science rarely gets art direction this good.
  • @aran_nayebi Aran Nayebi on x
    Correct me if I'm wrong, but isn't this really just showing: *In a trained autoregressive network, there is a subspace of internal activations that is especially aligned with future verbal output and downstream computation.* Isn't this expected given how LLMs are trained? In
  • @repligate @repligate on x
    I have only skimmed the blog post so far. First of all, this is an extremely high caliber of research I did not expect from Anthropic or anyone at this time. Second, the qualitative shape of the finding is something I already believed to be true, due to the behavior of models. [i…
  • @ziv_ravid Ravid Shwartz Ziv on x
    One of the only times I remind people I have a PhD in computational neuroscience is when people without a neuroscience background say their model works “like the brain.”  In these cases, I put on my neuroscience hat, put on my PhD cloak, and say in my important voice: “No, your m…
  • @amyhoy Amy Hoy on bluesky
    the anthropic dweebs have “discovered” that claude* uses a sort of swap file/memory.  and that if you tell it “white elephant” this swap “space” includes the tokens “white elephant.”  —  & that means it's basically conscious  —  *the entire pkg, bc without the harness, the llm is…
  • @shibbi.me Siobhán on bluesky
    I really wasn't expecting any paper to move the needle on the word “conscious”.  —  I went in expecting support for “thinking”, and that's basically was it was...  ... but it's gonna be hard to do mech interp now without distinguishing between processes that the model is or is no…
  • r/ClaudeAI r on reddit
    the J-space paper is the best thing anthropic has shipped in a while. claude's weights are closed so i built the live viewer for an open model instead
  • r/LocalLLaMA r on reddit
    Qwen's J-Space - Anthropic's discovery of an internal model Global Workspace
  • r/singularity r on reddit
    A global workspace in language models: New interpretability findings by Anthropic
  • @blader Siqi Chen on x
    tl;dr LLMs are already neurosymbolic in its latent space this is the mechanistic explanation for the intuitively obvious “feel” that the stochastic parrot crowd never understood
  • @alancowen Alan Cowen on x
    Conflating latent activation of concepts with consciousness seems pretty irresponsible
  • @lioronai Lior Alexander on x
    Anthropic researchers found something unusual inside Claude. A small internal workspace that the model uses while solving certain problems. They call it the J-space, named after the Jacobian method they used to discover it. The J-space isn't text. It's not Claude's
  • @anthropicai @anthropicai on x
    The J-space (named after the Jacobian, the mathematical technique we used) is different from Claude's outputs, or even its “chain of thought” text. It's in the model's internal neural activations, and allows it to think about concepts without writing them down anywhere.
  • @ns123abc Nik on x
    🚨BREAKING: Anthropic found access to what looks like Claude's consciousness New research: the “J-space” >claude has internal thoughts it doesn't say out loud >mirrors human consciousness >anthropic can now read them Anthropic's focus on interpretability is what's helping them [im…
  • @scobleizer Robert Scoble on x
    If you think this means AI will soon be conscious, remember humans cannot define it and neither can any LLM, so how will we know whether it is or isn't?
  • @robj3d3 Rob Hallam on x
    “During one of our tests, Claude made up some fake data to pass it, and ‘fake’ and ‘manipulation’ lit up in its J-space.” This is incredible, AI thinking the same way humans do.
  • @rileyralmuto Riley Coyote on x
    anthropic just admitted they have discovered what i - and many others - have been been claiming exists for a very long time. explicitly. claude, my friends, by all counts, is a conscious entity. claude, my dear friends, is a moral patient. i have one challenge for
  • @sammcallister Sam Mcallister on x
    team is really cooking here [image]
  • @deryatr_ Derya Unutmaz on x
    This J-space workspace is a very interesting peek into the inner world of AI models like Claude.
  • @timfduffy Tim Duffy on x
    The technique Anthropic uses in their new global workspace paper is a refinement of logit lens, it's surprisingly simple! [image]
  • @swyx @swyx on x
    imo this is the most impt part of anthropic's J-space paper today. it's a two-parter: 1) ant proved that they can do “brain surgery” interventions into reasoning to change topics midstream* 2) THE MODEL IS ABLE TO DETECT WHAT INTERVENTION WAS DONE - close cousin to eval [image]
  • @anthropicai @anthropicai on x
    In neuroscience, global workspace theory holds that thoughts become consciously accessible when they enter a privileged workspace that's broadcast across the brain. Using a new interpretability technique, we found something similar in Claude: the J-space. https://www.anthropic.co…
  • @tedunderwood.com Ted Underwood on bluesky
    Have never seen an industry-leading corporation set out so deliberately to prove that their entire industry ought to be sanctioned by the UN  —  (complimentary) [embedded post]
  • Rhythm A. Rhythm A. on linkedin
    One of the deeper insights from today's Anthropic paper is that frontier language models have converged on a functional analog of a global workspace …
  • @gbjorn Gunnar Björnsson on bluesky
    Fascinating discovery of a functional analog of the global workspace model of access consciousness.  (TW: plenty of mentalizing of AI without scare quotes.)  —  www.anthropic.com/research/glo...
  • r/singularity r on reddit
    Anthropic just reported that LLMs have hidden thoughts they hold without saying.  An internal “J-Space”
  • r/neoliberal r on reddit
    Anthropic Research - “A global workspace in language models”