Interview with Alphabet's Jigsaw product manager CJ Adams on how its AI-assisted Perspective tool, used by some publishers to weed out abusive comments, works
Rob Marvin / PCMag : Tweets: @rjmarvin1 . Thanks: @rjmarvin1 Tweets: Rob Marvin / @rjmarvin1 : Worked on this one for a while. My @PCMag look inside @Jigsaw's efforts to combine ML and human moderation with its Perspective AI tool, and the more philosophical question of how to think about and address rampant toxicity on the internet http://www.pcmag.com/... Thanks: @rjmarvin1
Context & Ripple Effects
This interview lands mid-arc for Alphabet's Jigsaw: what began as the Conversation AI anti-troll toolkit became the free Perspective API in 2017, and The New York Times' adoption that June showed publishers would reopen comment sections they had shut when moderation was manual. By early 2019, product manager CJ Adams is explaining to PCMag how the tool actually pairs machine-learning scoring with human moderators rather than replacing them.
The stakes have since grown well past newsroom comments: Jigsaw engineer Lucy Vasserman later described how OpenAI, Anthropic, and other labs use Perspective to flag toxic output from their own LLMs, and Google expanded the API with seven attributes including nuance in 2024.
First-order effects
- Publishers using Perspective get a concrete answer on workflow: ML scores comments at scale while humans handle judgment calls, which is what let The New York Times widen commenting to more articles after adopting the tool.
- Jigsaw's positioning as a free API makes toxicity scoring a commodity input for any publisher, shifting the cost of moderation from headcount to integration work.
Second-order effects
- Consumer-facing products follow the same scoring engine downstream — Jigsaw's Tune extension applies Perspective filtering directly to Reddit, Twitter, Facebook, YouTube, and Disqus, moving the tool from publisher back-ends into readers' browsers.
- As labs like OpenAI and Anthropic adopt Perspective for flagging toxic model output, Jigsaw gains an AI-safety customer base alongside its publisher one, making its toxicity taxonomy a shared reference across two markets.
Third-order effects
- If ML-plus-human moderation keeps proving out, platform governance consolidates around a few shared scoring APIs whose attribute definitions — expanded again in 2024 with nuance among seven attributes — effectively set the industry's working definition of toxicity.
- A tool built for comment sections becoming infrastructure for LLM safety points toward moderation vendors as cross-industry gatekeepers, with their thresholds shaping both public discourse and model behavior.
The trend: Toxicity moderation is evolving from a publisher-side convenience into shared ML infrastructure that serves both newsrooms and the AI labs training large language models.