Analysis: AI chatbots challenged or didn't respond to 90%+ of 15 false narratives spread by Russia, China, and Iran; AI overviews did it 60%+ of the time
Since AI chatbots exploded in popularity and Google started offering AI-generated answers, people who research foreign influence campaigns …Forums:r/politicsForums:r/politics:We tested how AI chatbots would handle foreign propaganda. They did surprisingly well
Context & Ripple Effects
Earlier research established that chatbots can generate persuasive conspiracy material, while a 2024 study of Russian-disinformation responses found systems repeating material tied to fake local-news sites. A 2025 NewsGuard audit of leading chatbots then found false claims from a pro-Kremlin network repeated in a third of tested responses.
The new analysis separates conversational chatbots from AI-generated search overviews: both often challenge or decline the tested narratives, but the overview format clears that bar less often. That distinction matters because answer interfaces can package a response as a direct resolution rather than a list of sources.
First-order effects
- Chatbot providers gain evidence that their handling of 15 Russian, Chinese, and Iranian false narratives has improved on this test, while Google’s AI overviews face a lower benchmark result of more than 60% versus more than 90% for chatbots.
- Researchers and trust-and-safety teams have a concrete comparative test case for evaluating whether a system challenges a narrative, declines to answer, or supplies it unchallenged.
Second-order effects
- Google has an incentive to test and tune AI-overview retrieval and response policies against foreign-influence narratives, since its search-answer product performs differently from the chatbot category in the analysis.
- Model providers’ safety assessments will face greater pressure to distinguish between standalone chat behavior and search-integrated answer behavior rather than treating “AI” performance as a single measure.
Third-order effects
- If independent narrative-level testing becomes routine, AI safety competition shifts from broad claims about model reliability toward auditable performance across specific high-risk information tasks.
- As generative answers become a distribution layer for information, governance will increasingly need to cover the retrieval-and-presentation system around a model, not only the model’s raw response behavior.
The trend: Foreign-influence resilience is becoming an operational AI-governance metric, with conversational and search-integrated AI systems requiring separate scrutiny.