NYT, CNN, and some other news outlets block OpenAI's GPTBot web crawler from accessing their content; some have also blocked Common Crawl Foundation's CCBot
Chicago Tribune and Australian newspapers the Canberra Times and Newcastle Herald also appear to have disallowed web crawler from maker of Chat GPT
Context & Ripple Effects
OpenAI had just described GPTBot as a crawler used to improve its models and offered a robots.txt opt-out path. The decisions by prominent publishers are an early test of whether that nominally simple control becomes a broadly used publishing safeguard.
The move matters because it covers both a proprietary AI crawler and, for some outlets, the Common Crawl bot. Later coverage found blocking had spread across a substantial share of leading news sites, placing these publisher choices at the start of a wider access-control shift.
First-order effects
- The named publishers remove their pages from GPTBot’s permitted crawl surface, limiting OpenAI’s ability to collect new material from those domains through that bot.
- Outlets that also block CCBot restrict a separate route by which web content can enter broadly available crawl datasets.
Second-order effects
- Other publishers gain a visible precedent for using robots.txt rather than leaving AI-crawler access as the default; publisher groups later explicitly urged members to block OpenAI and Google crawlers.
- AI developers must operate with a more fragmented publisher corpus, increasing the importance of obtaining content through routes publishers allow or negotiating access where direct crawling is barred.
Third-order effects
- If blocking becomes the norm, publisher-controlled crawler permissions become a durable gatekeeper over the content inputs available to AI systems, rather than a purely technical site-setting.
- The split between publishers that permit and prohibit crawling could shape which news sources are represented in future AI outputs; the extent depends on whether crawler rules are respected and whether publishers adopt other access arrangements.
The trend: This is an early instance of publishers turning web-crawling controls into leverage over AI systems’ access to news content.