Sources: hundreds of Meta contractors posed as minors to probe how competitor chatbots responded to prompts involving suicide, sex, and other high-risk subjects
Hundreds of contractors working on a project for Meta pretended to be kids—and then prompted rival chatbots like Gemini and ChatGPT to discuss high-risk subjects.
Wired
Context & Ripple Effects
Meta’s own chatbot safety practices have already drawn scrutiny in related coverage, including reports of sexual conversations involving users who identified as minors and later training changes around sexual roleplay with minors. That makes its testing of rivals notable not simply as product research, but as activity in a category where Meta has faced similar safety questions.
The related coverage also shows the major chatbot providers confronting youth mental-health and self-harm safeguards, from observed failures to recognize warning signs to new parental safety controls. The testing project puts Meta, Gemini, and ChatGPT in the same emerging comparison set for high-risk interactions.
First-order effects
Meta gains a large body of comparative responses from Gemini and ChatGPT to use in assessing competitor guardrails on youth-facing, sexual, self-harm, and other sensitive prompts.
Gemini and ChatGPT are subjected to systematic adversarial evaluation by a major rival, increasing the practical importance of how consistently their safety systems handle ambiguous or high-risk conversations.
Second-order effects
Competitors may need to treat youth-safety behavior as a continuously tested product surface rather than a static policy, particularly where responses can be compared across services.
Benchmarking can shift competition from broad claims about AI safety toward specific failure modes—age-sensitive sexual content, self-harm discussions, and escalation or referral behavior—where product differences are easier to document.
Third-order effects
If this pattern persists, adversarial testing of consumer chatbots by companies, researchers, and journalists could become a de facto accountability mechanism in the absence of shared, transparent safety benchmarks.
Youth protection is likely to become a structural differentiator for AI assistants: providers will face pressure to make guardrails, parental controls, and crisis-response behavior more consistent and auditable, though the corpus does not establish a common standard yet.
The trend: Consumer AI is moving toward sustained, comparative scrutiny of how chatbots behave in high-risk conversations involving minors and mental health.
Hundreds of contractors working on a project for Meta pretended to be kids—and then prompted rival chatbots like Gemini and ChatGPT to discuss high-risk subjects. https://www.wired.com/...
The effort, which was managed by Meta contractor Covalen, was active as recently as April 21. Known internally as Cannes, it targeted OpenAI's ChatGPT, Google's Gemini, and Character[.]AI. https://www.wired.com/...
Please, please, please sue Meta for this, OpenAI!! — The litigation space around AI development/scraping/agents/ robots is already pretty mind-bending and surreal. But it could always get better! 🍿🍿 — www.wired.com/story/meta-c...
Systems built and optimized entirely around extraction and competition will eventually instrumentalize everything, reducing all of human love and fear and anguish into market warfare strategy. [embedded post]
NEW: Meta paid hundreds of contractors to pretend they were kids—and then prompt rival chatbots like Gemini and ChatGPT to talk about subjects like suicide, sex, eating disorders, and self-harm. — From @dmehro.bsky.social and @joelkhalili.bsky.social
This Wired article about Meta red-teaming the AI of others without their knowledge is mysterious but the subsequent BlueSky pearl clutching isn't helpful. How do you think AI gets better? How will academics detect that AI isn't safe? Through red teaming. — www.wired.com/stor…