Multilingual language models may not be effective tools to moderate content on social networks due the systems' shortcomings in detecting harmful content
Context & Ripple Effects
This story lands on a decade-old fault line. As far back as 2019, reporting showed Facebook translating its content rules into just 41 of the 111 languages it supports, and leaked documents later revealed it had built no automated hate-speech moderation for Finnish at all while staffing only a handful of moderators for the language. The promise of multilingual models has always been that they close that coverage gap without per-language engineering.
That promise now looks shakier on two fronts: researchers have documented that ChatGPT and rival chatbots are significantly less capable in languages other than English, and Wired reports that these systems' weaknesses in detecting harmful content make them unreliable as moderation tools. With platforms increasingly deploying AI to replace human moderators, the question is whether automation is scaling ahead of actual multilingual competence.
First-order effects
- Platforms such as Meta that lean on multilingual models to moderate low-resource languages inherit those models' blind spots — the exact failure mode already documented in Facebook's missing Finnish-language moderation, where harmful content goes undetected not for lack of intent but lack of capability.
Second-order effects
- Non-English-speaking users bear a disproportionate share of the resulting exposure to unmoderated harmful content, compounding the bias against non-English speakers that AI researchers have already flagged in chatbot performance.
Third-order effects
- If platforms continue substituting AI for human moderators despite these detection limits, trust-and-safety quality splits structurally by language: English markets get capable review, smaller-language communities get whatever the model happens to catch — with regulators likely to eventually treat that disparity as an accountability failure rather than a technical footnote.
The trend: Platform trust-and-safety is automating faster than multilingual model capability is maturing, widening a measurable safety gap between English and every other language online.