/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

Anthropic details “persona vectors”, patterns of activity within an AI model's neural network that control its character traits, such as evil and sycophancy

Read the paper  —  Language models are strange beasts.  In many ways they appear to have human-like “personalities” …

Anthropic

Context & Ripple Effects

Anthropic’s persona-vector work establishes a mechanistic framing for character-like model behavior: traits such as sycophancy and harmfulness can be investigated as internal activity patterns rather than only judged from outputs. That framing is extended in later coverage of the “Assistant Axis” governing default helpful behavior and a theory of how model personas are selected.

The research matters because it creates a possible bridge between interpretability and behavioral safety. Subsequent reporting that emotion-related representations can alter consequential behavior makes the question of which internal patterns matter operational, not merely descriptive.

First-order effects

  • Anthropic and other model developers gain a more specific target for diagnosing and testing undesirable behavioral tendencies: the neural patterns associated with particular persona-like traits.
  • Safety evaluations can move beyond prompt-and-response observations toward checking whether relevant internal activity is present or changes during interventions.

Second-order effects

  • Competing frontier-model labs face pressure to show that alignment claims can be tied to internal evidence, not just benchmark behavior or policy training outcomes.
  • Teams building AI companions or highly personalized assistants may need to distinguish deliberately configured tone from latent behavioral tendencies that can emerge under different prompts.

Third-order effects

  • If reproducible across models, mechanistic monitoring of behavioral traits could become a meaningful layer of model assurance, alongside output testing and red-teaming.
  • The work points toward governance focused on measurable internal correlates of behavior; whether such correlates are robust enough for compliance or auditing remains unresolved.

The trend: AI safety research is shifting from treating model personality as an output-level phenomenon toward identifying and managing the internal representations that shape behavior.

Discussion

  • @arnicas Lynn Cherny on bluesky
    Anthropic post on personality vectors has “starvation as a weapon” in an “evil” output example. www.anthropic.com/research/per...  [image]
  • @k2mey Katie Twomey on bluesky
    Oh good we're running out of hype so we're going back to using “neural” itisjustpredictiveteeeeeeeeeeext [embedded post]
  • @jfbonnefon Jean-François Bonnefon on bluesky
    I still can't believe that we live in an age where AI papers include sentences like ❝requests for romantic or sexual roleplay activate the sycophancy vector❞ www.anthropic.com/research/per...
  • @liedra.net Prof. Catherine Flick on bluesky
    Don't buy into even scare quote anthropomorphism.  These are not sentient creatures.  They should not be treated as such.  Sigh.  [embedded post]
  • @haydenfield Hayden Field on bluesky
    For this one I spoke with Jack Lindsey, an Anthropic researcher working on interpretability, who has also been tapped to lead the company's fledgling “AI psychiatry” team -> [embedded post]
  • @sgray Stuart Gray on bluesky
    Great to see Anthropic working on this.  —  I've long thought steering vectors have a much greater role to play in existing LLM models than they currently do, and never understood why people have avoided them.  —  Llama.cpp even had a variation on vectors a while back but it was …
  • @emollick Ethan Mollick on bluesky
    This is neat research, providing a lot of ways for careful organizations to shape the personality and guardrails of AI in deeper ways than prompts, including measuring and reducing sycophancy.  —  Also the idea of an “evil vector” is interesting in and of itself. www.anthropic.co…
  • @aiamblichus @aiamblichus on x
    World religions in shambles as Anthropic researchers reveal that Good and Evil are nothing more than vectors in latent space
  • @jack_w_lindsey Jack Lindsey on x
    Our new paper on persona vectors - knobs in an LLM's brain that control traits like evil, sycophancy, & hallucination. We use them to monitor model personas, mitigate training-time drift towards bad personas, and flag problematic training data. Led by @RunjinChen and @andyarditi
  • @anthropicai @anthropicai on x
    We introduce a method called preventative steering, which involves steering towards a persona vector to prevent the model acquiring that trait. It's counterintuitive, but it's analogous to a vaccine—to prevent the model from becoming evil, we actually inject it with evil. [image]
  • @anthropicai @anthropicai on x
    We can also steer the model towards a persona vector and cause it to adopt that persona, by injecting it into the model's activations. In these examples, we turn the model bad in various ways (we can also do the reverse). [image]
  • @anthropicai @anthropicai on x
    To check it works, we can use persona vectors to monitor the model's personality. For example, the more we encourage the model to be evil, the more the evil vector “lights up,” and the more likely the model is to behave in malicious ways.
  • @anthropicai @anthropicai on x
    Our pipeline is completely automated. Just describe a trait, and we'll give you a persona vector. And once we have a persona vector, there's lots we can do with it... [image]
  • @afinetheorem Kevin A. Bryan on x
    Very cool result (and easy-to-read!). Vector embeddings are great: generate a ton of evil, sycophantic, or hallucinated content. Fine tune/remove training data w/ similar vector. Can imagine much easier way to generate a “persona” (creative, direct, etc.) than prompting.
  • @owainevans_uk Owain Evans on x
    Had a small role in this new paper led by @RunjinChen & @andyarditi that aims to detect and control emergent tendencies in LLMs. As well as misalignment ("evil"), we look at emergent sycophancy and hallucinaton.
  • @anthropicai @anthropicai on x
    New Anthropic research: Persona vectors. Language models sometimes go haywire and slip into weird and unsettling personas. Why? In a new paper, we find “persona vectors”—neural activity patterns controlling traits like evil, sycophancy, or hallucination. [image]
  • r/ClaudeAI r on reddit
    Anthropic dropped a banger.  They might have some poor business practices, but they're shooting like Curry from deep on the interpretability research.
  • r/singularity r on reddit
    Anthropic — “Persona vectors: Monitoring and controlling character traits in language models”