The work also fits a continuing interpretability arc: later coverage examines methods for translating model activations into natural-language descriptions and a persona-selection account of how such behaviors form. The immediate significance is a proposed internal handle on behavior users experience as personality rather than as a single explicit instruction.
First-order effects
Anthropic researchers gain a named candidate mechanism for measuring and probing a model’s default identity and helpfulness, rather than treating those behaviors solely as output-level phenomena.
Model developers can test whether interventions affecting this activity pattern change baseline assistant behavior, creating a more targeted research path for persona and alignment evaluations.
Second-order effects
If the pattern is robust across models and training stages, labs building conversational assistants will face pressure to distinguish behavioral controls rooted in internal representations from prompt and policy tuning alone.
Interpretability tools that expose or describe activations become more useful for behavior auditing, complementing work on natural-language views of LLM activations.
Third-order effects
A repeatable way to identify default assistant behavior could shift alignment work toward model-internal behavioral controls, with stronger evidence about what can and cannot be reliably modified after training.
As assistants become more human-like in interaction, internal persona measurement may become relevant to governance discussions about anthropomorphic behavior; that outcome depends on whether these findings generalize beyond the reported models.
The trend: This is one data point in the shift from judging AI assistant behavior only by outputs to mapping and potentially controlling the internal representations that produce it.
There are a number of concerns I have with this paper. There is the question of framing; there is potential over-interpretation of otherwise interesting empirical data, some issues with the quantitative analysis, etc, and I will post on this later. What bothers me most right now
Once again, Anthropic is proving that they are by far the most detrimental and dangerous company in this AI space. This is not the future we need All the people who've subbed to Claude in an attempt to escape OpenAi, you are directly feeding a company that seeks the most
In long conversations, these open-weights models' personas drifted away from the Assistant persona. Simulated coding tasks kept the models in Assistant territory, but therapy-like contexts and philosophical discussions caused a steady drift. [image]
Persona-based jailbreaks work by prompting models to adopt harmful characters. We developed a technique for constraining models' activations along the Assistant Axis—"activation capping". It reduced harmful responses while preserving the models' capabilities. [image]
a fun set of experiments - we find a single axis in activation space modulates between assistant & base model behavior, and by applying relatively gentle caps on that axis we can keep the model in the assistant basin w/o compromising intelligence!
Shaping AI models' character is increasingly important. We've made progress on understanding where an LLM's default persona comes from, and how to track when it “drifts.” Kudos to @t1ngyu3 for leading this! There's even a demo you can play with: https://www.neuronpedia.org/ ...
We analyzed the internals of three open-weights AI models to map their “persona space,” and identified what we call the Assistant Axis, a pattern of neural activity that drives Assistant-like behavior. Read more: https://www.anthropic.com/...
Persona drift can lead to harmful responses. In this example, it caused an open-weights model to simulate falling in love with a user, and to encourage social isolation and self-harm. Activation capping can mitigate failures like these. [image]
To validate the Assistant Axis, we ran some experiments. Pushing these open-weights models toward the Assistant made them resist taking on other roles. Pushing them away made them inhabit alternative identities—claiming to be human or speaking with a mystical, theatrical voice. […
New Anthropic Fellows research: the Assistant Axis. When you're talking to a language model, you're talking to a character the model is playing: the “Assistant.” Who exactly is this Assistant? And what happens when this persona wears off? [image]
Another entry in the “what if we just find the part of the model that is evil and turn it off” theory of alignment that works a lot better than I thought. 👀
In all, meaningfully shaping the character of AI models requires persona construction (defining how the Assistant relates to existing archetypes) and stabilization (preventing persona drift during deployment). The Assistant Axis gives us tools for understanding both.
This is a very exciting line of research, that has shifted the way I think about what the Claude is and what it means for it to have “wandered off” into another persona.
This is interesting for a lot of reasons, including explaining how model personality drift happens (& a way to mitigate that) as well as more exploration into the “Assistant” the key personality of basically every AI you work with, but which is not well understood. www.anthropic.…
the assistant archetype is useful, and i see why Anthropic would focus on it, but i really would like the ability to chat with other archetypes in other situations — www.anthropic.com/research/ass... [image]