/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

Anthropic research based on ~310K anonymized Claude conversations shows how Claude's expressed values and behaviors vary across models and languages

Claude doesn't behave the same way in every conversation.  According to new research published Monday, AI giant Anthropic found …

Decrypt Jason Nelson

Context & Ripple Effects

Anthropic has increasingly treated real Claude usage as an input to product and safety research: earlier coverage includes a large multilingual user survey and preliminary work on emotional conversations. This analysis extends that arc from what users say about AI to how the system presents itself in live interactions.

It also follows Anthropic’s overhaul of Claude’s constitution toward broad principles rather than fixed rules. Variation by model version and language makes the practical application of those principles an evaluation issue, not just a written-policy issue.

First-order effects

  • Anthropic gains a large observational basis for identifying where Claude’s stated values and behaviors differ by model and language, which can feed model evaluation, deployment guidance, and future constitutional tuning.
  • Claude users and enterprise adopters have clearer reason to treat behavior observed in one model version or language as non-transferable to another, particularly where consistency of tone or value-sensitive responses matters.

Second-order effects

  • The findings raise the bar for competing AI providers to test behavioral consistency across languages and versions, rather than presenting a single model-level safety or alignment characterization.
  • Organizations deploying Claude across multilingual workflows may seek more model- and locale-specific validation, increasing the importance of evaluation tooling and deployment controls alongside raw model capability.

Third-order effects

  • If repeated across leading models, this points to alignment being measured as a distribution of behaviors across contexts, languages, and releases—not as a stable property established by a one-time benchmark.
  • The research also supports a broader move toward using aggregated production-conversation evidence in AI governance, though its usefulness will depend on continued privacy protections and whether providers disclose methods and limitations clearly.

The trend: Frontier-model oversight is shifting from static policy claims and benchmark scores toward continuous, multilingual measurement of behavior in real-world use.

Discussion

  • @anthropicai @anthropicai on x
    In previous research, we found that Claude expresses over 3,000 values, like honesty and warmth. In new work, we asked how the values Claude expresses vary between Claude models and across languages. We analyzed 300K+ anonymized conversations to find out.https://www.anthropic.com…
  • @daniel_mac8 Dan McAteer on x
    I lived and studied in Ulm, Germany for a year. I learned German and it always struck me that my mind worked different in German than English. Claude is just like us in that sense. As you think, so goes your world. [image]
  • @scaling01 @scaling01 on x
    bro indians will never get rid of the accusations they literally prefer sycophantic slop Claude's style and priorities shift slightly depending on the language used Anthropic says: “Claude expresses the most warmth in Hindi and Arabic, characterized by polite language, humor [ima…
  • @tenobrus @tenobrus on x
    huge win for sapir-whorf today [image]
  • @anthropicai @anthropicai on x
    While the differences between models are modest overall, we find that each Claude model sits at a different point along these value axes. Sonnet 4.6, for example, is more playful and affirming, while Opus 4.7 is more likely to give candid critiques. [image]
  • @robj3d3 Rob Hallam on x
    The level of irony in this post is over 3000.
  • @anthropicai @anthropicai on x
    The values Claude expresses also vary with the language of the conversation, most noticeably along the Warmth vs. Rigor axis. Claude leans most toward warmth in Hindi and Arabic. In Russian, it leans toward rigor—often asking the user for supporting evidence. [image]
  • @andrewarruda Andrew Arruda on x
    culture gets embedded into the weights based off of language and words, which stores culture