/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

Internal documents: Facebook's AI has minimal success enforcing its rules against problematic content, including removing an estimated 3%-5% of hate speech

AI has only minimal success in removing hate speech, violent images and other problem content, according to internal company reports

Wall Street Journal

Context & Ripple Effects

Facebook had previously emphasized proactive detection and low reported prevalence in its transparency reporting; the internal estimate creates a sharp measurement gap with its public hate-speech enforcement metrics. Facebook also publicly disputed the account, citing a decline in hate-speech prevalence, making the distinction between content removed and content seen central to evaluating its AI.

The later reporting on race-blind hate-speech policies adds a distributional dimension: aggregate enforcement figures can obscure which users bear the remaining exposure to harmful language.

First-order effects

  • Facebook's AI moderation performance is directly challenged by internal findings that it removed only an estimated 3%–5% of hate speech and had limited success with other problematic content.
  • Facebook's public claim of lower hate-speech prevalence now sits alongside a separate internal measure of removal effectiveness, complicating how its safety reporting is interpreted.

Second-order effects

  • Facebook will face pressure to distinguish proactive takedown rates and prevalence estimates from the share of violating content its systems actually catch, rather than presenting those measures as interchangeable.
  • The gap gives greater weight to subgroup exposure: the reported shortcomings of race-blind policy enforcement indicate that platform-wide averages may not reflect minority users' experience.

Third-order effects

  • If platforms continue to report detection volume and prevalence separately from verified coverage, AI moderation governance will shift toward outcome measures that test what enforcement systems miss and who sees it.
  • Content-safety systems are likely to be judged less by automation rates than by auditable evidence that automated rules work across harmful-content categories and affected communities.

The trend: Platform AI governance is moving from headline moderation totals toward scrutiny of actual enforcement coverage, user exposure, and uneven safety outcomes.

Discussion

  • @daphnehk Daphne Keller on x
    Govts & media: If you don't do magic content moderation, you're bad and will be punished. Platforms: We're doing it! Govts & media: You lied about doing magic content moderation, you're bad and will be punished. I'm not even sure who's the bad guy in this story. But it's bad. htt…
  • @dseetharaman Deepa Seetharaman on x
    New: Facebook says its AI can root out a huge amount of hate & violence. The reality: “we do not and possibly never will have a model that captures even a majority of integrity harms.” Latest Facebook Files drop by me, @JeffHorwitz & @ScheckWSJ https://www.wsj.com/...
  • @roncharles Ron Charles on x
    “Mild cockfights were deemed acceptable, but those in which the birds were seriously hurt were banned. But the computer model couldn't distinguish fighting roosters from non-fighting roosters.” https://www.wsj.com/...
  • @hypervisible @hypervisible on x
    Facebook's own engineers don't even believe the company's bs. https://www.wsj.com/...
  • @dseetharaman Deepa Seetharaman on x
    Violence is also a challenge. First-person shooter videos were confused with car washes. Cockfighting videos with car crashes. The AI couldn't tell between two roosters next to each other & two fighting. Just a reminder: VERY sharp minds are working on this. It's hard.
  • @jason_kint Jason Kint on x
    It's a deep and troubling report that deserves your attention. Side note, I'm impressed by @selenagomez as the report indicates she used her influence to draw attention to these issues at Facebook. Need more like her. Thanks 🙏🏽. /8 https://www.wsj.com/...
  • @qjurecic Quinta Jurecic on x
    ["Ms. Gomez wrote back that Ms. Sandberg hadn't addressed her broader questions, sending screenshots of Facebook groups that promoted violent ideologies."] Good to know that Facebook uses the same PR strategy with Selena Gomez that it does with the rest of us
  • @josephmenn Joseph Menn on x
    Hey look, another WSJ story that calls out Facebook's misinformation about itself. In this episode, the spectacular AI that removes more and more rule-violating hate speech was actually deleting less than 10% of such posts. https://www.wsj.com/...
  • @mattwarman Matt Warman MP on x
    It's not entirely accurate to say that algorithms can clean up the internet, regardless of your views on what a better internet looks like. https://www.wsj.com/...
  • @dseetharaman Deepa Seetharaman on x
    In public, execs paint an optimistic picture. They boast of a “proactive detection rate” of 90%-plus - meaning most hate content FB removed was first rooted out by AI. That doesn't tell you what % of hate speech they catch, which has been in the low single digits for years.
  • @akikofujita Akiko Fujita on x
    “Facebook's AI can't consistently identify first-person shooting videos, racist rants and even, in one notable episode that puzzled internal researchers for weeks, the difference between cockfighting and car crashes.” https://www.wsj.com/...
  • @justinhendrix Justin Hendrix on x
    Facebook touts the power of its AI, but internal company documents say it has only minimal success in enforcing its rules against hate speech, violent images and other problematic content. The latest from @dseetharaman, @JeffHorwitz & @ScheckWSJ: https://www.wsj.com/...
  • @dseetharaman Deepa Seetharaman on x
    You can't hire enough humans to monitor every Facebook post / comment. You need AI. But FB's AI struggles to interpret its surgically drawn policies around hate speech. Their automated systems delete an estimated 3-5% of views or hate speech. In Afghanistan, that rate is 0.23%.
  • @dseetharaman Deepa Seetharaman on x
    That detection rate, by the way, was partly buoyed by cost shifts in 2019, when FB cut the number of human reviewers dedicated to hate and redirected them to help train the algorithm. FB also made it harder to file a user complaint & autodeleted more likely crappier reports.
  • @georgia_wells Georgia Wells on x
    “The problem is that we do not and possibly never will have a model that captures even a majority of integrity harms, particularly in sensitive areas,” a senior engineer wrote By @dseetharaman @JeffHorwitz + @ScheckWSJ https://www.wsj.com/...
  • @dseetharaman Deepa Seetharaman on x
    I hope you'll read this piece, which is less about Facebook's effects on the world and more about how it's trying to manage harms with minimal success. https://www.wsj.com/...
  • @amy_siskind @amy_siskind on x
    “On hate speech, the documents show, Facebook employees have estimated the company removes only a sliver of the posts that violate its rules—a low-single-digit percent.” Facebook has also reduced its human workforce to rely on AI. https://www.wsj.com/...
  • @gregbensinger @gregbensinger on x
    OMG — Facebook wasn't happy with its hate speech removal rate, so it just made it harder to report hate speech. Problem solved https://www.wsj.com/... https://twitter.com/...
  • @chadsday Chad Day on x
    Facebook touts the power of its AI but internal company documents say it has only minimal success in enforcing its rules against hate speech, violent images and other problem content; ⁦@dseetharaman⁩ ⁦@JeffHorwitz⁩ ⁦@ScheckWSJ⁩ https://www.wsj.com/...
  • @sheeraf She-Ra Frenkel on x
    NEW : Instagram's internal marketing documents show they are struggling with fears of losing their “pipeline” of teenagers. With @RMac18 and @MikeIsaac https://www.nytimes.com/...
  • @evelyndouek Evelyn Douek on x
    I agree w fb that prevalence (how much people actually saw) rather than bulk removal nos is the important metric in terms of content moderation of hate speech. Which is why fb was dishonest to lean so hard into removal nos for yrs as evidence of progress https://about.fb.com/... …
  • @baekdal Thomas Baekdal on x
    The point that FB makes here is important. When tech companies say that they have removed ‘x millions’ of bad posts, that number doesn't mean anything because you can code a bot to produce as many posts as you want. What matters is how much of it is seen. https://about.fb.com/...
  • @baekdal Thomas Baekdal on x
    I will say this, though, Facebook is being very selective about this metric. Sometimes they like boasting about how many millions of post they have removed, so... I agree with you post Facebook. But, can we please get a bit more consistency in how you use this metric.
  • @rmac18 @rmac18 on x
    Instagram tracks an internal metric called “teen time spent” and is worried that it's losing cachet with teens. We took a look at Instagram's obsession to keep teens on the platform. Latest w/ @sheeraf and @MikeIsaac. https://www.nytimes.com/...
  • @senmarkey Ed Markey on x
    More internal Facebook documents show that when Instagram sees kids and teens, they only see dollar signs. Congress needs to act and stop Facebook and Big Tech from utilizing the tobacco playbook of manipulating users to hook them when they're young. https://www.nytimes.com/...