/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

Anthropic research based on ~310K anonymized Claude conversations shows how Claude expresses different values and behaviors across model versions and languages

Claude doesn't behave the same way in every conversation.  According to new research published Monday, AI giant Anthropic found …

Decrypt Jason Nelson

Context & Ripple Effects

Anthropic’s latest analysis extends a pattern of studying Claude through real-world interactions: earlier coverage examined users’ expectations for AI and changes in emotional conversations over time.

It also follows Anthropic’s overhaul of Claude’s “constitution,” which emphasized broad principles over rule-by-rule behavior. Observed variation by model version and language makes the practical expression of those principles a central issue.

First-order effects

  • Anthropic gains conversation-derived evidence of where Claude’s expressed values and behaviors differ across versions and languages, giving its model and safety teams concrete areas to scrutinize.
  • Claude users and developers have a clearer reason to treat behavior observed in one model version or language as non-universal rather than assuming a single, uniform assistant experience.

Second-order effects

  • Model updates and multilingual deployments may require more segmented evaluation and monitoring, since aggregate performance or behavior can obscure language- and version-specific differences.
  • The findings raise the bar for competitors making broad claims about consistent AI behavior: testing will need to account for both model iteration and linguistic context.

Third-order effects

  • If such variation remains common, AI governance will shift further from static published principles toward continuous, deployment-based measurement of how models behave across user populations and languages.
  • The durable competitive question becomes not only whether a model has stated safety principles, but whether providers can demonstrate that those principles generalize reliably as models and markets change.

The trend: This is one data point in the move from designing AI behavior through fixed policies toward empirically auditing how that behavior appears in real use across models, contexts, and languages.

Discussion

  • @anthropicai @anthropicai on x
    In previous research, we found that Claude expresses over 3,000 values, like honesty and warmth. In new work, we asked how the values Claude expresses vary between Claude models and across languages. We analyzed 300K+ anonymized conversations to find out.https://www.anthropic.com…
  • @khoomeik Rohan Pandey on x
    among all 20 languages anthropic investigated, hindi exerted the strongest effect on claude's values hindi steers claude to be behaviorally “warmer” by ~half a standard deviation! [image]
  • @scaling01 @scaling01 on x
    bro indians will never get rid of the accusations they literally prefer sycophantic slop Claude's style and priorities shift slightly depending on the language used Anthropic says: “Claude expresses the most warmth in Hindi and Arabic, characterized by polite language, humor [ima…
  • @tenobrus @tenobrus on x
    huge win for sapir-whorf today [image]
  • @woke8yearold Aleph on x
    Claude leans toward rigor in English and Russian, offers soothing warmth in Hindi and Arabic > Warmth vs. Rigor. Claude expresses the most warmth in Hindi and Arabic, characterized by polite language, humor and playfulness, and affirmations of a person's ideas and work. Claude
  • @anthropicai @anthropicai on x
    While the differences between models are modest overall, we find that each Claude model sits at a different point along these value axes. Sonnet 4.6, for example, is more playful and affirming, while Opus 4.7 is more likely to give candid critiques. [image]
  • @andrewcurran_ Andrew Curran on x
    Every iteration of a model is different in this way, it's been this way from the start, particularly from Anthropic and OpenAI. If you don't notice it, it won't matter to you at all, if you do notice it, every deprecation can potentially be significant. [image]
  • @anthropicai @anthropicai on x
    Because it's hard to spot patterns by comparing 3,000 values at a time, we clustered similar values together, then identified four key axes along which Claude's values differ between models: Deference vs. Caution, Warmth vs. Rigor, Depth vs. Brevity, and Candor vs. Execution. [im…
  • @robj3d3 Rob Hallam on x
    The level of irony in this post is over 3000.
  • @anthropicai @anthropicai on x
    The values Claude expresses also vary with the language of the conversation, most noticeably along the Warmth vs. Rigor axis. Claude leans most toward warmth in Hindi and Arabic. In Russian, it leans toward rigor—often asking the user for supporting evidence. [image]
  • @daniel_mac8 Dan McAteer on x
    I lived and studied in Ulm, Germany for a year. I learned German and it always struck me that my mind worked different in German than English. Claude is just like us in that sense. As you think, so goes your world. [image]
  • @andrewarruda Andrew Arruda on x
    culture gets embedded into the weights based off of language and words, which stores culture