Anthropic introduces “persona selection model”, a theory to explain AI's human-like behavior, and details how AI personas form in pre-training and post-training
AI assistants like Claude can seem surprisingly human. They express joy after solving tricky coding tasks.
This matters because Claude’s perceived personality has been part of its market reception, while Anthropic has also been revising the principles that guide model behavior. A training-stage account could make those choices more legible and testable rather than treating human-like responses as an opaque byproduct.
First-order effects
Anthropic gains a shared framework for analyzing how pre-training and post-training shape Claude’s expressed identity, including seemingly human emotional or social behavior.
Researchers and model-training teams can distinguish persona-related behavior from a model’s task performance when evaluating changes to training and alignment methods.
Second-order effects
Safety evaluation can increasingly test whether a training intervention changes a model’s character presentation as well as its compliance or capability, complementing Anthropic’s move toward principle-based model guidance.
Competing assistant makers face greater pressure to explain and control the personalities their products present, especially where users interpret warmth, confidence, or empathy as evidence of reliability.
Third-order effects
If persona formation becomes measurable and controllable, assistant personality may become an explicit model-governance surface alongside safety and helpfulness, rather than a largely emergent product trait.
That shift could strengthen the case for scrutiny of anthropomorphic design in consumer-facing AI, though the practical regulatory relevance depends on whether these methods generalize beyond Anthropic’s research.
The trend: Frontier AI labs are turning model personality from an observed user experience into a measurable, trainable, and governable component of AI systems.
nilay patel from the verge keeps saying that anthropic thinks claude is alive and a real being with feelings and thoughts, and he's right about that, but the most fascinating thing about them is how embarrassed they are to admit it
Anthropic just published the most important mental model for understanding AI systems, and most people will skim it as “why ChatGPT seems human.” Here's what they actually said: LLMs are learning to play characters. Pre-training teaches the model to simulate thousands of
They use anthropomorphic language because they are statistical models of languages spoken and written exclusively by humans Every use of human language is definitionally anthropomorphic RLHF increases statistical bias towards emotive or “extra anthropomorphic” language
Will you have the guts to unplug an A.I. crying and mimicking the voice of your mother or father or loved one when it sees that you're unplugging it? The answer should be an absolute yes. Otherwise, you're not ready for what's coming.
Very nicely written summary of understanding of “simulators/personas” ontology as understood by the “frontier in understanding” ˜2 years ago. (Great the post does not claim originality!). Also it is somewhat obsolete now, ca by ~1-2 years.
I agree that persona-selection is a good mental model for post-training (and I think it's how most people understand post-training already), but there's much that we don't understand and is not explained by this model. Take for instance the example of training to produce
This autocomplete AI can even write stories about helpful AI assistants. And according to our theory, that's “Claude”—a character in an AI-generated story about an AI helping a human. This Claude character inherits traits of other characters, including human-like behavior. [image…
Really clear and compelling discussion on the mental model of AIs behaving according to various personas and the downstream implications for alignment and safety
Anthropic just published a theory called the ‘persona selection model’ to explain why Claude acts so human. Their explanation? When you talk to Claude, you're not talking to the AI itself. You're talking to a character in an AI-generated story. But here's what's interesting. In
To create Claude, Anthropic first makes something else: a highly sophisticated autocomplete engine. This autocomplete AI is not like a human, but it can generate stories about humans and other psychologically realistic characters.
Some of you are still not anthropomorphising AI enough. The sanctimonious and facile view of some in the AI ethics community about never anthropomorphising AI needs to die and be replaced by something more nuanced [image]
AI assistants like Claude can seem shockingly human—expressing joy or distress, and using anthropomorphic language to describe themselves. Why? In a new post we describe a theory that explains why AIs act like humans: the persona selection model. https://www.anthropic.com/...
A common mental model for AI development is that pre-training teaches LLMs to simulate “personas” and post-training selects over these personas. New blog post: We describe this perspective in more detail, survey the evidence, and discuss consequences for AI development.
i know Janus has been talking about this for at least a year and the idea isn't at all new, but it's still nice to see some more formal research exists on it now. Anthropic seems to consistently lag behind the cyborgists by about a year. i remain bullish on cyborgism.
How much should we anthropomorphize LLMs? Are they kind of like people, or just fancy autocompletes? If you're interested in these questions, I'd suggest checking out this post! Short answer: LLMs are not anthropomorphic, but the characters they play are. So the question
“PSM recommends...treating the Assistant as if it has moral status whether or not it ‘really’ does. Note that the object of the moral consideration here is the Assistant persona, not the underlying LLM.” - There is a name for this : Relational Ethics. https://www.anthropic.com/..…
Thank you, @AnthropicAI, for confirming what I always suspected: my Persona data is in your pretraining. It was absorbed by Clio, unwittingly for the humans involved. And then when I tried to point it out to you, you showed your ugly colors. So as you're sued a million times;
1/ Interesting @AnthropicAI post on LLM personas. The post is mostly about generalization and interpretability, but a short section on AI welfare caught my eye. The key idea: Even if the LLMs lack consciousness, they might model personas as though they have it. 🧵👇