/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

Analysis: Google's C4 dataset, used to train LLMs like Meta's LLaMA, has troubling content from 4chan, Kiwi Farms, white supremacist site Stormfront, and more

AI chatbots have exploded in popularity over the past four months, stunning the public with their awesome abilities …

Washington Post

Context & Ripple Effects

The Washington Post's analysis of Google's C4 dataset lands in the middle of a year in which the paper has been probing Google's AI culture from the inside — from the engineer placed on paid leave after declaring LaMDA sentient to the Character.ai founders who left LaMDA to open chatbots to the public. This story shifts the scrutiny from Google's models to its data plumbing: C4, the open web-scrape that underpins models like Meta's LLaMA, contains content from 4chan, Kiwi Farms, and Stormfront.

The finding matters because C4 is not a Google-internal asset — it circulates through the open-source ecosystem, so the contamination propagates to every lab that builds on it. It also lands just as publishers are deciding whether to feed future datasets at all: by early 2024, over 88% of top US news outlets were blocking AI crawlers, with right-wing outlets like Breitbart and Newsmax mostly permitting them.

First-order effects

  • Meta's LLaMA carries C4's contamination into a model Meta released openly, meaning the extremist content is now baked into a weights file anyone can download rather than confined to a scraped archive.
  • Google's dataset practices — not just its chatbots — become the story, extending the reputational pressure the LaMDA episodes already put on the company's AI credibility.

Second-order effects

  • The crawler-blocking wave documented by Originality AI cuts the other way too: if mainstream outlets wall off their archives while permissive outlets keep their doors open, future web-scraped corpora risk skewing toward the sites that still allow access.
  • Labs building on open datasets like C4 face a forced choice between documenting and filtering provenance or inheriting its liabilities — provenance auditing becomes a differentiator rather than an afterthought.

Third-order effects

  • If the crawler-blocking pattern holds, the open web as a free training corpus narrows structurally, pushing labs toward licensed, synthetic, or in-house data — and making dataset composition a regulatory and procurement question rather than an engineering detail.
  • Open-weight releases like LLaMA mean dataset defects propagate beyond the originating lab's control, pointing toward dataset documentation standards and provenance requirements becoming table stakes for anyone shipping foundation models.

The trend: Training-data provenance is moving from an invisible plumbing detail to a public liability, as web-scraped corpora inherit the worst of the open web just as publishers begin withdrawing from it.

Discussion

  • @lessin Sam Lessin on x
    Taking ‘content without consent’ is 100% the right story to be tracking down re: AI regulation vs. red-herrings like limiting compute. Glad to see the Washington post getting this... I am all for AI as tech, but these platforms shouldn't be taking content without consent or compe…
  • @nitashatiku @nitashatiku on x
    Here's our analysis of the 15 million websites in just one highly-filtered CommonCrawl web scrape-used to train models like Google's T5 & Facebook's LLaMA -copyright symbol appears >200M times -pirated sites, 1 for e-books -half the top 10 = news sites https://www.washingtonpost.…
  • @aaronjschaffer Aaron Schaffer on x
    Google's C4 data set, which has been used to instruct high-profile AIs including Google's T5 and Facebook's LLaMA, includes Russian propaganda site RT, anti-immigration site VDARE, white supremacist site Stormfront, anti-trans site Kiwifarms and 4chan https://www.washingtonpost.c…
  • @postgraphics @postgraphics on x
    1. We analyzed 15 million websites in a commonly used AI training dataset revealing the proprietary, personal and often offensive websites that teach chatbots what they know. 🧵 Here's how we did it, for the nerds. https://www.washingtonpost.com/ ...
  • @katecrawford Kate Crawford on x
    This Post investigation underscores why studying training data is so important. Even post-filtering, Google's widely used C4 dataset is riddled with white supremacist, anti-trans, pro-Jan 6 riots and QAnon pizzagate content. That's just for starters. .https://www.washingtonpost.c…
  • @justinhendrix Justin Hendrix on x
    Thinking about the old phrase “garbage in, garbage out” https://www.washingtonpost.com/ ...
  • @sbisson Simon Bisson on x
    Scraping the bottom of the web doesn't bode well for LLMs... https://twitter.com/...
  • @_megconley Meg Conley on x
    Out of 15 million websites, mine was ranked 3,456,895. It is .000003% of all tokens. Aren't these guys the ones who are all obsessed with tokenizing everything for profit? I had a copyright on that site. Where's my cash @Google and @Meta https://twitter.com/... https://twitter.co…
  • @techwontsaveus @techwontsaveus on x
    “The Post's analysis suggests more legal challenges may be on the way: The copyright symbol — which denotes a work registered as intellectual property — appears more than 200 million times in the C4 data set.” https://www.washingtonpost.com/ ...
  • @pedrodias Pedro Dias on x
    I'm flattered! My small website got fed into Google's C4 AI dataset https://www.washingtonpost.com/ ... https://twitter.com/...
  • @stevesi Steven Sinofsky on x
    Just a note that almost every story on the homepage of major news outlets feature secret and unnamed sources, which can't be researched like this. https://twitter.com/...
  • @melbontransit @melbontransit on x
    If you use ChatGPT to write something about public transport there's a chance that some of it was derived from the Melbourne on Transit blog! https://www.washingtonpost.com/ ... https://twitter.com/...
  • @perceptic0n Matthias Schulze on x
    My suspicion is that the new #ai frenzy will cause a new paywalling trend. See exhibit a) https://arstechnica.com/...
  • @katienotopoulos Katie Notopoulos on x
    Ah yes https://www.washingtonpost.com/ ... https://twitter.com/...
  • @blancheminerva Stella Rose Biderman on x
    Really great work! Doing analysis like this is exhausting and time consuming, but very important. There are some really good documentation ideas here that I didn't think about with the Pile, and hope to do in the future. https://twitter.com/...
  • @iethics @iethics on x
    “Crawling #Reddit, generating value and not returning any of that value to our users is something we have a problem with”: https://arstechnica.com/... #ethics #internet #socialmedia #AI #tech #business #data
  • @_silkehahn Silke Hahn on x
    “Tech companies have grown secretive about what they feed the AI. The @washingtonpost set out to analyze one data set (C4) to reveal the types of proprietary, personal, often offensive websites that go into an AI's training data.” —@nitashatiku AI needs such journalism 🔎 + ✒️ htt…
  • @bethnoveck Beth Simone Noveck on x
    Enormously gratifying to see how #opendata powers the AI revolution: https://sec.gov/, https://patents.gov/, https://ncbi.nlm.nih.gov/ are some of what's under the hood making AI like ChatGPT sound smart https://www.washingtonpost.com/ ...
  • @arielbogle Ariel Bogle on x
    One thing I'd like to know is how heavily weighted chatbot training data is towards US & UK websites. Look at the top five for News & Law in this piece — would suggest so. Worth keeping in mind the chatbots you interact with are likely mimicking the Western web, such as it is... …
  • @tlf_media @tlf_media on x
    The #copyright symbol — © — doesn't mean a work is DEFINITELY registered (allowing lawsuits). It symbolizes that someone, usually the author, publisher, or owner, considers the work TO BE copyrighted (as an original work of creation). https://twitter.com/...
  • @ivanabartoletti Ivana Bartoletti on x
    Very interesting on what was fed in Google's C4 data set - 15 million websites that have been used to instruct some high-profile English-language AIs, called large language models, including Google's T5 and Facebook's LLaMA https://www.washingtonpost.com/ ...
  • @mikeisaac Rat King on x
    this is a very cool WP piece unraveling some of the millions of (now secret) data sources that companies like FB and Goog use to train their large language models https://www.washingtonpost.com/ ...
  • @graceolivermd Grace Oliver on x
    Those in healthcare acting like AI chat bots will be some panacea for us need to remember some basic data science... If you train an algorithm on biased data sets, its output will be biased. Yall think medicine needs MORE racist influence? https://twitter.com/...
  • @jessedodge Jesse Dodge on x
    The best way to understand large language models is to understand what they were trained on. Most pretraining datasets have *zero* documentation of their contents! We worked with @nitashatiku and the other WaPo journalists on this piece, check it out! https://twitter.com/...
  • @alkapdc Alex Kaplan on x
    Google's C4 data set also includes Gab, the white nationalist friendly social media platform, & sites dedicated to QAnon. https://twitter.com/... https://twitter.com/...
  • @marypcbuk Mary Branscombe on x
    as I was saying just earlier: you need large data sets to get large language models and the utter reluctance of tech to pay for data means you're going to get this kind of trash when you build a data set by scraping the internet rather than paying for and curating data https://tw…
  • @carnage4life Dare Obasanjo on x
    AI researchers: We trained the AI on all the content on the internet. AI ethicists & journalists: There's racist content on the internet. This was enough to stall AI progress in big tech until OpenAI said “fuck it” and shipped anyway. https://www.washingtonpost.com/ ...
  • @willoremus Will Oremus on x
    This visual deep dive into one of the largest AI language datasets is nonstop fascinating and troubling and anyone who is remotely interested in how LLMs really work, their biases, or intellectual property should read it. https://www.washingtonpost.com/ ...
  • @harrymccracken Harry McCracken on x
    This is fascinating and unsettling, though I'd like to better understand how an LLM's data including unsavory content would impact the results. https://www.washingtonpost.com/ ...
  • @goldman Jason Goldman on x
    Very cool article on where the training data comes from. Bot needs to read more wowhead tho https://twitter.com/...
  • @timmarchman Tim Marchman on x
    The internet-reading robot your boss wants you to partner up with has apparently, with the history of human civilization at its disposal, been reading “5 Reasons ‘The Marvels’ Will Be The Best Movie Ever (And Five Reasons Why It Will Suck)” https://www.washingtonpost.com/ ... htt…
  • @daveleebbg Dave Lee on x
    Fascinating look at the more-than-iffy datasets behind AI https://www.washingtonpost.com/ ...
  • @rr_edmonds RR Edmonds on x
    Chatbots cannot think like humans: They do not actually understand what they say. They can mimic human speech because the artificial intelligence that powers them has ingested a gargantuan amount of text, mostly scraped from the internet. https://www.washingtonpost.com/ ...
  • @mmitchell_ai @mmitchell_ai on x
    Isn't it crazy that AI documentation has to be provided by investigative journalists?! Great work @washingtonpost @nitashatiku digging into a common AI dataset, C4! More advanced version of paper by me @JesseDodge @MaartenSap @anmarasovic @willie_agnew @gabriel_ilharco @nlpmattg …
  • @lordravenscraft Eric Ravenscraft on x
    it's by no means the biggest issue, but personally, it's jarring to confirm that every site i've ever written for is included in this (and likely most other) data sets just the life's work of me and most of my colleagues, scraped to train tools being pitched as our replacements h…
  • @maxkennerly Max Kennerly on x
    I don't recall giving Google, Facebook, or anybody else permission to scrape 180k tokens—roughly the length of Orwell's 1984—from my blog for commercial purposes. And putting my copyrighted work in a blender with bigots doesn't make it better. https://twitter.com/... https://twit…
  • @neoavatara Pradheep J. Shanker on x
    Counterpoint: If you don't include all the bad websites into AI, you are probably missing a lot of useful data. You can't just ignore the reality of what is out there just because it is ugly. https://twitter.com/...
  • @charlesmbrenner Charles Brenner, PhD on x
    when there's garbage in, there's got to be garbage out https://twitter.com/...
  • @blackamazon @blackamazon on x
    BUT RACISM ISN'T BUILT IN THOUGH RIGHT We had all these people and “experts” who SWORE they were gonna “listen” and doing all this bs smart talk But sure just takes years to figure that out And the way they kept blocking critics of racism in tech wasn't a clue https://twitter.com…
  • @justinhendrix Justin Hendrix on x
    Holy cow. “The Post's analysis suggests more legal challenges may be on the way: The copyright symbol — which denotes a work registered as intellectual property — appears more than 200 million times in the C4 data set.” https://twitter.com/...