Google's tech incubator and think tank, Jigsaw, will soon release Conversation AI, a set of tools that use machine learning to combat online trolls
Around midnight one Saturday in January, Sarah Jeong was on her couch, browsing Twitter, when she spontaneously wrote what she now bitterly refers to as …
Context & Ripple Effects
Conversation AI arrives five months into Jigsaw's life as Alphabet's tech-forward successor to Google Ideas, following its February rebrand from think tank to incubator and September's program that redirected aspiring ISIS recruits away from propaganda. The unit's method is consistent: apply machine learning to social problems at platform scale.
The motivating case here is individual — Wired opens with writer Sarah Jeong's experience of targeted Twitter harassment — but the release lands as publishers are actively looking for automated ways to keep comment sections viable.
First-order effects
- Publishers and platforms gain a machine-learning toolkit for filtering troll and abuse from comments, shifting moderation from manual review toward scored, automated triage.
- Jigsaw extends its portfolio beyond counter-extremism work into everyday online discourse, making harassment — not just terrorism — a formal problem statement for the incubator.
Second-order effects
- The toolset matures into Perspective, a free API for publishers, which turns Jigsaw's internal research into default moderation infrastructure across news sites rather than a single-product feature.
- Once toxicity scoring exists as an API, demand comes from unexpected buyers: by 2023 OpenAI, Anthropic, and other labs were using the same Perspective API to flag toxic output from LLMs, and Jigsaw kept expanding it with new attributes like nuance through 2024 (seven-attribute expansion).
Third-order effects
- If the pattern holds, toxicity detection built for human-to-human comment sections becomes shared governance plumbing for machine-generated speech — one scorer sitting between platforms, publishers, and AI labs.
- That consolidation around a common API concentrates definitional power over what counts as 'toxic' in a small set of tools, raising the stakes of how attributes like nuance are calibrated industry-wide.
The trend: Machine-learning moderation built to protect human comment sections is evolving into cross-industry infrastructure that also polices what AI models themselves say.