Internal documents: Facebook's AI has minimal success enforcing its rules against problematic content, including removing an estimated 3%-5% of hate speech
AI has only minimal success in removing hate speech, violent images and other problem content, according to internal company reports
Wall Street Journal
Context & Ripple Effects
Facebook had previously emphasized proactive detection and low reported prevalence in its transparency reporting; the internal estimate creates a sharp measurement gap with its public hate-speech enforcement metrics. Facebook also publicly disputed the account, citing a decline in hate-speech prevalence, making the distinction between content removed and content seen central to evaluating its AI.
The later reporting on race-blind hate-speech policies adds a distributional dimension: aggregate enforcement figures can obscure which users bear the remaining exposure to harmful language.
First-order effects
- Facebook's AI moderation performance is directly challenged by internal findings that it removed only an estimated 3%–5% of hate speech and had limited success with other problematic content.
- Facebook's public claim of lower hate-speech prevalence now sits alongside a separate internal measure of removal effectiveness, complicating how its safety reporting is interpreted.
Second-order effects
- Facebook will face pressure to distinguish proactive takedown rates and prevalence estimates from the share of violating content its systems actually catch, rather than presenting those measures as interchangeable.
- The gap gives greater weight to subgroup exposure: the reported shortcomings of race-blind policy enforcement indicate that platform-wide averages may not reflect minority users' experience.
Third-order effects
- If platforms continue to report detection volume and prevalence separately from verified coverage, AI moderation governance will shift toward outcome measures that test what enforcement systems miss and who sees it.
- Content-safety systems are likely to be judged less by automation rates than by auditable evidence that automated rules work across harmful-content categories and affected communities.
The trend: Platform AI governance is moving from headline moderation totals toward scrutiny of actual enforcement coverage, user exposure, and uneven safety outcomes.
Related: AI enforcement surface · Operational AI governance · Facebook's AI · Facebook disputes the report's hate-speech prevalence findings · Facebook's Q3 hate-speech transparency report · Facebook's race-blind hate-speech policies
Related Coverage
- View article Ars Technica
- Streaming is changing how Hollywood works, for better and for worse Protocol
- New Internal Documents Contradict Facebook's Claims that AI Can Enforce Its Rules tech.slashdot.org · EditorDavid
- Facebook wants you to believe its AI is working against hate speech TNW · Ivan Mehta
- Facebook defends hate speech moderation AI, claims 50% drop in 3 years TechCircle
- Facebook responds after report accuses it of not fighting hate speech well enough phonearena.com · Iskra Petrova
- Facebook claims it uses AI to identify and remove posts containing hate speech and violence, but the technology doesn't really work, report says Insider · Emily Walsh
- Facebook hits back at claims its AI has minimal success in fighting hate speech ZDNet · Aimee Chanthadavong
- Can AI Police Facebook? Maybe Not Just Yet, Report Says PYMNTS.com
- Facebook claims hate speech visibility dropped 50 percent in nine months Engadget · Jon Fingas
- Selena Gomez privately put Facebook execs on blast in 2020 for all the hateful content Mashable
- Facebook disputes report its AI has little effect on hate speech CNET · Steven Musil
- Facebook claims it has drastically reduced hate speech prevalence SlashGear · Brittany A. Roston
- Hate Speech Prevalence Has Dropped by Almost 50% on Facebook About Facebook · Guy Rosen
- Facebook denies weak performance on hateful content BBC
- Internal memo reveals Instagram's concern about losing its teenage users phonearena.com · Alan Friedman
- Instagram Feared Losing Its Teen Users, Internal Memo Shows Android Headlines · Chethan Rao
- Facebook Pushes Back Against Report That Claims Its AI Sucks at Detecting Hate Speech Gizmodo · Alyse Stanley
- Instagram internal documents reveal fear of losing teens, report says CNET · Edward Moyer
- Internal docs: Instagram spent the majority of its global annual marketing budget since 2018 targeting teenagers, as it worries about losing its user “pipeline” New York Times
- Instagram spends the majority of its $390 million global advertising budget on targeting teens, according to a new report Insider · Aaron Holmes
Discussion
-
@daphnehk
Daphne Keller
on x
Govts & media: If you don't do magic content moderation, you're bad and will be punished. Platforms: We're doing it! Govts & media: You lied about doing magic content moderation, you're bad and will be punished. I'm not even sure who's the bad guy in this story. But it's bad. htt…
-
@dseetharaman
Deepa Seetharaman
on x
New: Facebook says its AI can root out a huge amount of hate & violence. The reality: “we do not and possibly never will have a model that captures even a majority of integrity harms.” Latest Facebook Files drop by me, @JeffHorwitz & @ScheckWSJ https://www.wsj.com/...
-
@roncharles
Ron Charles
on x
“Mild cockfights were deemed acceptable, but those in which the birds were seriously hurt were banned. But the computer model couldn't distinguish fighting roosters from non-fighting roosters.” https://www.wsj.com/...
-
@hypervisible
@hypervisible
on x
Facebook's own engineers don't even believe the company's bs. https://www.wsj.com/...
-
@dseetharaman
Deepa Seetharaman
on x
Violence is also a challenge. First-person shooter videos were confused with car washes. Cockfighting videos with car crashes. The AI couldn't tell between two roosters next to each other & two fighting. Just a reminder: VERY sharp minds are working on this. It's hard.
-
@jason_kint
Jason Kint
on x
It's a deep and troubling report that deserves your attention. Side note, I'm impressed by @selenagomez as the report indicates she used her influence to draw attention to these issues at Facebook. Need more like her. Thanks 🙏🏽. /8 https://www.wsj.com/...
-
@qjurecic
Quinta Jurecic
on x
["Ms. Gomez wrote back that Ms. Sandberg hadn't addressed her broader questions, sending screenshots of Facebook groups that promoted violent ideologies."] Good to know that Facebook uses the same PR strategy with Selena Gomez that it does with the rest of us
-
@josephmenn
Joseph Menn
on x
Hey look, another WSJ story that calls out Facebook's misinformation about itself. In this episode, the spectacular AI that removes more and more rule-violating hate speech was actually deleting less than 10% of such posts. https://www.wsj.com/...
-
@mattwarman
Matt Warman MP
on x
It's not entirely accurate to say that algorithms can clean up the internet, regardless of your views on what a better internet looks like. https://www.wsj.com/...
-
@dseetharaman
Deepa Seetharaman
on x
In public, execs paint an optimistic picture. They boast of a “proactive detection rate” of 90%-plus - meaning most hate content FB removed was first rooted out by AI. That doesn't tell you what % of hate speech they catch, which has been in the low single digits for years.
-
@akikofujita
Akiko Fujita
on x
“Facebook's AI can't consistently identify first-person shooting videos, racist rants and even, in one notable episode that puzzled internal researchers for weeks, the difference between cockfighting and car crashes.” https://www.wsj.com/...
-
@justinhendrix
Justin Hendrix
on x
Facebook touts the power of its AI, but internal company documents say it has only minimal success in enforcing its rules against hate speech, violent images and other problematic content. The latest from @dseetharaman, @JeffHorwitz & @ScheckWSJ: https://www.wsj.com/...
-
@dseetharaman
Deepa Seetharaman
on x
You can't hire enough humans to monitor every Facebook post / comment. You need AI. But FB's AI struggles to interpret its surgically drawn policies around hate speech. Their automated systems delete an estimated 3-5% of views or hate speech. In Afghanistan, that rate is 0.23%.
-
@dseetharaman
Deepa Seetharaman
on x
That detection rate, by the way, was partly buoyed by cost shifts in 2019, when FB cut the number of human reviewers dedicated to hate and redirected them to help train the algorithm. FB also made it harder to file a user complaint & autodeleted more likely crappier reports.
-
@georgia_wells
Georgia Wells
on x
“The problem is that we do not and possibly never will have a model that captures even a majority of integrity harms, particularly in sensitive areas,” a senior engineer wrote By @dseetharaman @JeffHorwitz + @ScheckWSJ https://www.wsj.com/...
-
@dseetharaman
Deepa Seetharaman
on x
I hope you'll read this piece, which is less about Facebook's effects on the world and more about how it's trying to manage harms with minimal success. https://www.wsj.com/...
-
@amy_siskind
@amy_siskind
on x
“On hate speech, the documents show, Facebook employees have estimated the company removes only a sliver of the posts that violate its rules—a low-single-digit percent.” Facebook has also reduced its human workforce to rely on AI. https://www.wsj.com/...
-
@gregbensinger
@gregbensinger
on x
OMG — Facebook wasn't happy with its hate speech removal rate, so it just made it harder to report hate speech. Problem solved https://www.wsj.com/... https://twitter.com/...
-
@chadsday
Chad Day
on x
Facebook touts the power of its AI but internal company documents say it has only minimal success in enforcing its rules against hate speech, violent images and other problem content; @dseetharaman @JeffHorwitz @ScheckWSJ https://www.wsj.com/...
-
@sheeraf
She-Ra Frenkel
on x
NEW : Instagram's internal marketing documents show they are struggling with fears of losing their “pipeline” of teenagers. With @RMac18 and @MikeIsaac https://www.nytimes.com/...
-
@evelyndouek
Evelyn Douek
on x
I agree w fb that prevalence (how much people actually saw) rather than bulk removal nos is the important metric in terms of content moderation of hate speech. Which is why fb was dishonest to lean so hard into removal nos for yrs as evidence of progress https://about.fb.com/... …
-
@baekdal
Thomas Baekdal
on x
The point that FB makes here is important. When tech companies say that they have removed ‘x millions’ of bad posts, that number doesn't mean anything because you can code a bot to produce as many posts as you want. What matters is how much of it is seen. https://about.fb.com/...
-
@baekdal
Thomas Baekdal
on x
I will say this, though, Facebook is being very selective about this metric. Sometimes they like boasting about how many millions of post they have removed, so... I agree with you post Facebook. But, can we please get a bit more consistency in how you use this metric.
-
@rmac18
@rmac18
on x
Instagram tracks an internal metric called “teen time spent” and is worried that it's losing cachet with teens. We took a look at Instagram's obsession to keep teens on the platform. Latest w/ @sheeraf and @MikeIsaac. https://www.nytimes.com/...
-
@senmarkey
Ed Markey
on x
More internal Facebook documents show that when Instagram sees kids and teens, they only see dollar signs. Congress needs to act and stop Facebook and Big Tech from utilizing the tobacco playbook of manipulating users to hook them when they're young. https://www.nytimes.com/...