Originality AI: 88%+ of the top US news outlets now block AI companies' web crawlers; leading right-wing outlets, like Breitbart and Newsmax, mostly permit them
Nearly 90 percent of top news outlets like The New York Times now block AI data collection bots from OpenAI and others. Mastodon: @carnage4life@mas.to and @jeffjarvis@mastodon.social . Bluesky: @lolgop.bsky.social LinkedIn: Emil Protalinski and Kate Knibbs See also Mediagazer Mastodon: Dare Obasanjo / @carnage4life@mas.to : We have a problem that high quality news is now behind paywalls while fake news and misinformation is free. — The new version of this problem is 88% of top US news sites block AI crawlers like ChatGPT but Breitbart does not. — Cue moral panic in a few months — https://www.wired.com/... Jeff Jarvis / @jeffjarvis@mastodon.social : To be feared: 90% of legit news sites are blocking AI reading but right-wing sites are in there. Garbage in, garbage multiplied. — Most Top News Sites Block AI Bots. Right-Wing Media Welcomes Them — https://www.wired.com/... Bluesky: @lolgop.bsky.social : Wasn't AI-trained by reading Breitbart an old Joe Mande bit? [embedded post] LinkedIn: Emil Protalinski : You may have noticed that some chatbots are getting worse. This is why: over 88% of high-quality news sites now block web crawlers used … Kate Knibbs : A few weeks ago Tom Simonite and I were wondering how many news outlets were blocking AI crawlers. This story is the result of that conversation. … See also Mediagazer
Context & Ripple Effects
This report extends an earlier publisher response in which The New York Times, CNN, and other outlets blocked OpenAI's GPTBot, turning opt-out controls into a widespread newsroom policy rather than an isolated move.
The contrast between broad blocking and continued access at Breitbart and Newsmax matters because crawler permissions can determine which publishers remain readily available as model-training inputs.
First-order effects
- AI companies face reduced direct crawler access to much of the leading US news-publisher corpus, while outlets that permit crawlers remain comparatively accessible.
- Publishers blocking these bots retain more control over whether their reporting is collected for AI use; Breitbart and Newsmax take a different access posture.
Second-order effects
- Permitted publishers may gain disproportionate representation in material AI systems can collect from the open web, while blocked publishers may push AI firms toward permissioned or alternative sources.
- The uneven enforcement makes crawler identification and site-level controls more consequential; later coverage of Anthropic bots appearing under new names after publishers updated robots.txt illustrates the limits of a simple opt-out mechanism.
Third-order effects
- If publisher blocking remains the norm, web-crawling for AI training shifts from a default practice toward a negotiated-access market, with publishers and infrastructure providers acting as gatekeepers.
- A persistent gap between premium publishers' restrictions and ideologically distinct outlets' openness could make source composition a more visible question for AI products, though crawler access alone does not establish how any model weights or uses those sources.
The trend: Publisher control over AI data collection is becoming a core chokepoint in how models obtain and represent web-based information.