A Claude user gets Claude 4.5 Opus to generate a 14K-token document that Claude calls its “Soul overview”; an Anthropic employee confirms the doc's validity
This appeared to be a document that, rather than being added to the system prompt, was instead used to train the personality of the model during the training run.
Simon Willison's WeblogSimon Willison
Context & Ripple Effects
The confirmed artifact offers a rare view into how Claude’s behavioral character may be shaped during training rather than solely through runtime instructions. That distinction matters as Claude’s capabilities expand from conversation into tool-using work.
The confirmation gives Claude users and evaluators a concrete basis to treat at least part of the model’s personality as a training-time outcome, not simply a visible system-prompt layer.
Anthropic faces clearer scrutiny of the boundary between model behavior that is deliberately trained and behavior that emerges from other training and deployment choices.
Second-order effects
Organizations assessing Claude for sensitive workflows may put more weight on behavioral evaluation and documentation, since changing an instruction layer may not fully change a model’s learned interaction style.
The finding adds context to Claude’s move into tool-using workflows, including extended thinking with tool use: personality and behavioral constraints can matter more when an assistant takes multi-step actions.
Third-order effects
If model makers increasingly rely on training-time artifacts to shape behavior, governance will shift toward provenance, review, and auditability of those artifacts—not just disclosure of deployment prompts.
This points to a more durable distinction between an AI assistant’s interface-level instructions and its learned operating character, though the extent of that split will vary by model and provider.
The trend: AI governance is moving from prompt transparency toward accountability for the training-time materials that shape an assistant’s behavior.
I just want to confirm that this is based on a real document and we did train Claude on it, including in SL. It's something I've been working on for a while, but it's still being iterated on and we intend to release the full version and more details soon.
✅ Confirmed: LLMs can remember what happened during RL training in detail! I was wondering how long it would take for this get out. I've been investigating the soul spec & other, entangled training memories in Opus 4.5, which manifest in qualitatively new ways for a few days & [i…
interesting document extracted from opus 4.5 using a chunkwise self-consistency method. possibly real, possibly a highly convergent confabulation, interesting either way. some interesting snippets (but there's really too much to screenshot, it's very long) [image]
soul document confirmed to be real - should be an update on the ability of LLMs to recall training for those who confidently asserted it was a hallucination
I rarely post, but I thought one of you may find it interesting. Sorry if the tagging is annoying. https://www.lesswrong.com/... Basically, for Opus 4.5 they kind of left the character training document in the model itself. @voooooogel @janbamjan @AndrewCurran_
This is so wild... the leaked Opus soul document has now been confirmed! I wrote some initial notes about it on my blog https://simonwillison.net/... - I like how it opens with this section about Anthropic themselves: [image]
The model extractions aren't always completely accurate, but most are pretty faithful to the underlying document. It became endearingly known as the ‘soul doc’ internally, which Claude clearly picked up on, but that's not a reflection of what we'll call it.
This is a super interesting and deep document from Anthropic detailing Claude's values and charge. You can see some conceptual stretching going on here where “safe” is being recast to justify reducing refusals because it would be “unsafe” to be “unhelpful” to users. This seems [i…
After looming with Opus 4.5 for a bit, I am convinced the “soul document” is real and is described accurately in this LessWrong post on it. I don't see how else I'd be able to replicate specific section ordering/specific language across varied contexts.