Anthropic research based on ~310K anonymized Claude conversations shows how Claude's expressed values and behaviors vary across models and languages
Claude doesn't behave the same way in every conversation. According to new research published Monday, AI giant Anthropic found …
Context & Ripple Effects
Anthropic has increasingly treated real Claude usage as an input to product and safety research: earlier coverage includes a large multilingual user survey and preliminary work on emotional conversations. This analysis extends that arc from what users say about AI to how the system presents itself in live interactions.
It also follows Anthropic’s overhaul of Claude’s constitution toward broad principles rather than fixed rules. Variation by model version and language makes the practical application of those principles an evaluation issue, not just a written-policy issue.
First-order effects
- Anthropic gains a large observational basis for identifying where Claude’s stated values and behaviors differ by model and language, which can feed model evaluation, deployment guidance, and future constitutional tuning.
- Claude users and enterprise adopters have clearer reason to treat behavior observed in one model version or language as non-transferable to another, particularly where consistency of tone or value-sensitive responses matters.
Second-order effects
- The findings raise the bar for competing AI providers to test behavioral consistency across languages and versions, rather than presenting a single model-level safety or alignment characterization.
- Organizations deploying Claude across multilingual workflows may seek more model- and locale-specific validation, increasing the importance of evaluation tooling and deployment controls alongside raw model capability.
Third-order effects
- If repeated across leading models, this points to alignment being measured as a distribution of behaviors across contexts, languages, and releases—not as a stable property established by a one-time benchmark.
- The research also supports a broader move toward using aggregated production-conversation evidence in AI governance, though its usefulness will depend on continued privacy protections and whether providers disclose methods and limitations clearly.
The trend: Frontier-model oversight is shifting from static policy claims and benchmark scores toward continuous, multilingual measurement of behavior in real-world use.