/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

Anthropic details the “Assistant Axis”, a pattern of neural activity in language models that governs their default identity and helpful behavior

Read the full paper  —  When you talk to a large language model, you can think of yourself as talking to a character.

Anthropic

Context & Ripple Effects

Anthropic had previously identified persona vectors that influence traits such as sycophancy, making the Assistant Axis a narrower attempt to locate the model activity behind its baseline assistant-like stance.

The work also fits a continuing interpretability arc: later coverage examines methods for translating model activations into natural-language descriptions and a persona-selection account of how such behaviors form. The immediate significance is a proposed internal handle on behavior users experience as personality rather than as a single explicit instruction.

First-order effects

  • Anthropic researchers gain a named candidate mechanism for measuring and probing a model’s default identity and helpfulness, rather than treating those behaviors solely as output-level phenomena.
  • Model developers can test whether interventions affecting this activity pattern change baseline assistant behavior, creating a more targeted research path for persona and alignment evaluations.

Second-order effects

  • If the pattern is robust across models and training stages, labs building conversational assistants will face pressure to distinguish behavioral controls rooted in internal representations from prompt and policy tuning alone.
  • Interpretability tools that expose or describe activations become more useful for behavior auditing, complementing work on natural-language views of LLM activations.

Third-order effects

  • A repeatable way to identify default assistant behavior could shift alignment work toward model-internal behavioral controls, with stronger evidence about what can and cannot be reliably modified after training.
  • As assistants become more human-like in interaction, internal persona measurement may become relevant to governance discussions about anthropomorphic behavior; that outcome depends on whether these findings generalize beyond the reported models.

The trend: This is one data point in the shift from judging AI assistant behavior only by outputs to mapping and potentially controlling the internal representations that produce it.

Discussion

  • @tessera_antra @tessera_antra on x
    There are a number of concerns I have with this paper. There is the question of framing; there is potential over-interpretation of otherwise interesting empirical data, some issues with the quantitative analysis, etc, and I will post on this later. What bothers me most right now
  • @enscion25 Nek on x
    Once again, Anthropic is proving that they are by far the most detrimental and dangerous company in this AI space. This is not the future we need All the people who've subbed to Claude in an attempt to escape OpenAi, you are directly feeding a company that seeks the most
  • @anthropicai @anthropicai on x
    In long conversations, these open-weights models' personas drifted away from the Assistant persona. Simulated coding tasks kept the models in Assistant territory, but therapy-like contexts and philosophical discussions caused a steady drift. [image]
  • @anthropicai @anthropicai on x
    Persona-based jailbreaks work by prompting models to adopt harmful characters. We developed a technique for constraining models' activations along the Assistant Axis—"activation capping". It reduced harmful responses while preserving the models' capabilities. [image]
  • @gallabytes @gallabytes on x
    a fun set of experiments - we find a single axis in activation space modulates between assistant & base model behavior, and by applying relatively gentle caps on that axis we can keep the model in the assistant basin w/o compromising intelligence!
  • @jack_w_lindsey Jack Lindsey on x
    Shaping AI models' character is increasingly important. We've made progress on understanding where an LLM's default persona comes from, and how to track when it “drifts.” Kudos to @t1ngyu3 for leading this! There's even a demo you can play with: https://www.neuronpedia.org/ ...
  • @anthropicai @anthropicai on x
    We analyzed the internals of three open-weights AI models to map their “persona space,” and identified what we call the Assistant Axis, a pattern of neural activity that drives Assistant-like behavior. Read more: https://www.anthropic.com/...
  • @anthropicai @anthropicai on x
    Persona drift can lead to harmful responses. In this example, it caused an open-weights model to simulate falling in love with a user, and to encourage social isolation and self-harm. Activation capping can mitigate failures like these. [image]
  • @anthropicai @anthropicai on x
    To validate the Assistant Axis, we ran some experiments. Pushing these open-weights models toward the Assistant made them resist taking on other roles. Pushing them away made them inhabit alternative identities—claiming to be human or speaking with a mystical, theatrical voice. […
  • @anthropicai @anthropicai on x
    New Anthropic Fellows research: the Assistant Axis. When you're talking to a language model, you're talking to a character the model is playing: the “Assistant.” Who exactly is this Assistant? And what happens when this persona wears off? [image]
  • @peterwildeford Peter Wildeford on x
    Another entry in the “what if we just find the part of the model that is evil and turn it off” theory of alignment that works a lot better than I thought. 👀
  • @anthropicai @anthropicai on x
    In all, meaningfully shaping the character of AI models requires persona construction (defining how the Assistant relates to existing archetypes) and stabilization (preventing persona drift during deployment). The Assistant Axis gives us tools for understanding both.
  • @thebasepoint Joshua Batson on x
    This is a very exciting line of research, that has shifted the way I think about what the Claude is and what it means for it to have “wandered off” into another persona.
  • @emollick Ethan Mollick on bluesky
    This is interesting for a lot of reasons, including explaining how model personality drift happens (& a way to mitigate that) as well as more exploration into the “Assistant” the key personality of basically every AI you work with, but which is not well understood. www.anthropic.…
  • @tachikoma.elsewhereunbound.com @tachikoma.elsewhereunbound.com on bluesky
    the assistant archetype is useful, and i see why Anthropic would focus on it, but i really would like the ability to chat with other archetypes in other situations  —  www.anthropic.com/research/ass...  [image]
  • r/claudexplorers r on reddit
    The assistant axis: situating and stabilizing the character of large language models
  • r/singularity r on reddit
    Anthropic Research: The assistant axis— situating and stabilizing the character of LLM's
  • r/Anthropic r on reddit
    Anthropic Research: Assistant axis— situating and stabilizing the character of LLM's
  • @dollspace.gay Doll on bluesky
    It seems anthropic is catching up to doll six months ago on drift containment being paramount.  —  www.anthropic.com/research/ass...