/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

← → days · ↑ ↓ browse · Enter similar · o open

NYT, CNN, and some other news outlets block OpenAI's GPTBot web crawler from accessing their content; some have also blocked Common Crawl Foundation's CCBot

Chicago Tribune and Australian newspapers the Canberra Times and Newcastle Herald also appear to have disallowed web crawler from maker of Chat GPT

The Guardian Ariel Bogle

Context & Ripple Effects

OpenAI had just described GPTBot as a crawler used to improve its models and offered a robots.txt opt-out path. The decisions by prominent publishers are an early test of whether that nominally simple control becomes a broadly used publishing safeguard.

The move matters because it covers both a proprietary AI crawler and, for some outlets, the Common Crawl bot. Later coverage found blocking had spread across a substantial share of leading news sites, placing these publisher choices at the start of a wider access-control shift.

First-order effects

  • The named publishers remove their pages from GPTBot’s permitted crawl surface, limiting OpenAI’s ability to collect new material from those domains through that bot.
  • Outlets that also block CCBot restrict a separate route by which web content can enter broadly available crawl datasets.

Second-order effects

  • Other publishers gain a visible precedent for using robots.txt rather than leaving AI-crawler access as the default; publisher groups later explicitly urged members to block OpenAI and Google crawlers.
  • AI developers must operate with a more fragmented publisher corpus, increasing the importance of obtaining content through routes publishers allow or negotiating access where direct crawling is barred.

Third-order effects

  • If blocking becomes the norm, publisher-controlled crawler permissions become a durable gatekeeper over the content inputs available to AI systems, rather than a purely technical site-setting.
  • The split between publishers that permit and prohibit crawling could shape which news sources are represented in future AI outputs; the extent depends on whether crawler rules are respected and whether publishers adopt other access arrangements.

The trend: This is an early instance of publishers turning web-crawling controls into leverage over AI systems’ access to news content.

Discussion

  • @ProfJohanna@mastodon.online Johanna Gibson on mastodon
    New York Times, CNN and Australia's ABC block OpenAI's GPTBot web crawler from accessing content |  Artificial intelligence (AI) |  The Guardian #LLMs #MachineLearning #ChatGPT #OpenAI #ArtificialIntelligence #AI https://www.theguardian.com/ ... https://amp.theguardian.com/ ...