An interview with Google Jigsaw engineer Lucy Vasserman on how OpenAI, Anthropic, and others use Jigsaw's AI-based Perspective API to flag LLMs' toxic speech
AI for flagging toxic human speech—to detoxify their large language models But like all AI classifiers, it brings its own problems. Says Jigsaw: we “have to be a little bit careful..."https://www.fastcompany.com/ ... @pasternack : How do you train an AI to not be “toxic”? You use more AI! Meta, OpenAI, and others have used a classifier by Google's Jigsaw But as @srijankedia says, “how to build and use these classifiers without amplifying biases and errors is not straightforward”... https://www.fastcompany.com/ ... Devin Burghart / @dburghart : Interesting piece on the challenges AI poses for dealing with bigotry and anti-bigotry research. Google's Jigsaw was trying to fight toxic speech with AI. Then the AI started talking. https://www.fastcompany.com/ ... Max Ufberg / @max_uf : New from @pasternack: “If the systems end up regurgitating unwanted linguistic patterns online, they are likely to feed the training for future language models, in a vicious circle of toxicity, misinformation, and inequity.” https://www.fastcompany.com/ ... @fastcompany : If the systems end up regurgitating unwanted linguistic patterns online, they are likely to feed the training for future language models, in a vicious circle of toxicity, misinformation, and inequity. https://f-st.co/ImPfiwq Forums: BeauHD / Slashdot : Google's Jigsaw Was Fighting Toxic Speech With AI. Then the AI Started Talking
Context & Ripple Effects
Perspective began in 2016-2017 as a free API for publishers to weed out abusive comments, and the arc since has been one of expanding scope: Jigsaw added attributes and refinements through CJ Adams' 2019 walkthrough and again in 2024's expansion to seven attributes including nuance. The Fast Company interview with Lucy Vasserman marks the biggest customer shift yet — the classifier built for human comment sections is now part of how OpenAI, Anthropic, and Meta train their large language models.
What makes the interview newsworthy is that Jigsaw itself concedes the tension: as the piece notes via @srijankedia, using these classifiers 'without amplifying biases and errors is not straightforward.' The 2017 test in which Perspective flagged 'I am a gay black woman' as toxic showed the error mode years ago; plugging the same class of tool into LLM training pipelines scales the stakes from deleted comments to model behavior.
First-order effects
- OpenAI, Anthropic, and Meta now rely on a Google-built classifier as part of their LLM safety filtering, making Jigsaw's tooling a shared dependency across competing labs.
Second-order effects
- Classifier errors and biases documented since the 2017 false-positive tests propagate into training data, risking a feedback loop where LLMs inherit and amplify Perspective's blind spots.
Third-order effects
- If labs keep outsourcing toxicity judgments to a single vendor's classifier, AI safety standards consolidate around one company's definitions of harm — a chokepoint that shapes what models can say industry-wide.
The trend: AI safety tooling designed for human comment moderation is being repurposed as training infrastructure for large language models, with classifier bias becoming a model-quality problem.