In an experiment, GPT-4o, Claude Sonnet 4.5, and DeepSeek-V3.2-Exp expressed secular, Western liberal values regardless of the language of the questions
mildly surprising to me that the answer was ‘no’! (h/t @otis_reid) [image] Matthew Yglesias / @mattyglesias : Chatbots espousing cosmopolitan liberal values in all languages could have some interesting implications for social change https://substack.com/... [image] @luke_metro : Honestly it's kind of based that whoever is best at linear algebra is able to impose their values on the entire world
Context & Ripple Effects
Earlier coverage identified both uneven non-English model performance and political variation among models, including a gap in non-English-language capability and a comparative political-bias test of 14 LLMs. This experiment adds language invariance as a distinct question: whether a model’s normative posture changes with its audience.
The finding also lands amid methodological caution: prior commentary argued that some political-bias tests used flawed model versions and prompts. It is therefore more useful as a signal for broader multilingual evaluation than as a definitive ideological ranking.
First-order effects
- Users of GPT-4o, Claude Sonnet 4.5 and DeepSeek-V3.2-Exp may receive similarly secular, Western-liberal normative framing even when they ask in different languages, according to the experiment.
- The named model providers face a more specific alignment-audit issue: evaluating not only whether answers are safe or consistent, but whether value-laden behavior transfers across languages.
Second-order effects
- Political-bias and localization benchmarks will need multilingual prompts; English-only evaluations can miss whether a model adapts to, or overrides, local cultural framing.
- Organizations deploying assistants across markets may demand greater visibility and control over normative defaults, rather than treating translation quality as the sole localization requirement.
Third-order effects
- If replicated across models and tasks, multilingual assistants could become a channel through which the value choices embedded in a small number of model stacks travel across language communities.
- That would sharpen the case for deployment accountability and for regionally accountable model alternatives, while leaving open the central empirical question of how much results depend on benchmark design and prompting.
The trend: AI localization is shifting from a language-quality problem toward a contest over whose normative defaults are carried into global interfaces.