/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

Analysis: at 2023's end, 48% of the top news websites in the US, the UK, and eight others had blocked OpenAI's crawler and 24% had blocked Google's AI crawler

In this factsheet we describe the proportion of top news websites in ten countries that block AI (artificial intelligence) crawlers. X: @carnage4life , @risj_oxford , @risj_oxford , @risj_oxford , @richrdfletcher , @risj_oxford , @risj_oxford , and @risj_oxford See also Mediagazer X: Dare Obasanjo / @carnage4life : Access to online news continues down the trend of being pay to play. 48% of the most popular news sites across ten countries were blocking OpenAI's crawlers and 24% block Google's AI crawler. AI chatbots that wants to have current info will need to pay. https://reutersinstitute.politics.ox.ac .uk/ ... @risj_oxford : 1️⃣ First, the big picture: By the end of 2023, 48% of the top news sites across ten countries were blocking @OpenAI crawlers. Only 24% blocked @Google AI crawler. 97% of the sites that decided to block Google's crawler were blocking OpenAI's too #GenAI https://reutersinstitute.politics.ox.ac .uk/ ... @risj_oxford : How many news sites block AI crawlers from @Google and @OpenAI? This is the question at the heart of a new factsheet by @richrdfletcher, who looks at the 15 top news sites in 10 countries covered by our #DNR23 📱Full factsheet https://reutersinstitute.politics.ox.ac .uk/ ... 🧵5 findings in thread #AI [image] @risj_oxford : 3️⃣ Figures were much lower for @google AI crawler: The proportion blocking this AI crawler ranged from 60% in Germany to 7% in Poland and Spain. In general, outlets in the Global North were more likely to be blocking than those in the Global South 📈 Full figures below [image] Richard Fletcher / @richrdfletcher : New @risj_oxford factsheet by me that asks: How many news websites block generative AI like ChatGPT and Gemini from using their content to train their models? It depends on the country. Very large differences in how many top news sites are blocking, and how soon they started. [image] @risj_oxford : 4️⃣ During 2023, none of the websites we examined had reversed their decision after deciding to block. News outlets with a relatively large online news reach were slightly more likely to be blocking AI crawlers than those with a relatively small reach. See the chart below [image] @risj_oxford : 5️⃣ All types of news outlet were blocking, but the websites of legacy print publications were more likely to be blocking than those of either broadcasters or digital-born outlets 📊 Figures in the chart below [image] @risj_oxford : 2️⃣ The headline figures mask very large differences by country: The proportion of top online news websites blocking @OpenAI ranged from 79% in the US, to just 20% in Mexico and Poland. 📈 You can see figures for every country (and how blocking evolved) in the chart below [image] See also Mediagazer

Reuters Institute for the Study of Journalism Dr Richard Fletcher

Context & Ripple Effects

Publishers had already begun using robots.txt after OpenAI introduced GPTBot and its opt-out mechanism, with outlets including the New York Times and CNN among the early blockers of GPTBot. This cross-country count shows that response had become widespread rather than limited to a few prominent publishers.

The overlap in blocking decisions—nearly all sites blocking Google’s AI crawler also blocked OpenAI’s—suggests publishers were treating AI crawling as a distinct rights and distribution question, not merely a dispute with one company.

First-order effects

  • OpenAI faces reduced access to content from leading news sites in the surveyed markets, while Google’s AI crawler faces a smaller but materially overlapping set of restrictions.
  • News publishers that block the crawlers retain more control over whether their reporting is available for AI collection; no surveyed site had reversed a blocking decision during 2023.

Second-order effects

  • AI providers seeking current news coverage will have greater incentive to negotiate permission or find other sources, while publishers can differentiate between open web indexing and AI-specific access.
  • The uneven country results—from high Google-crawler blocking in Germany to low blocking in Spain—make news availability to AI systems more dependent on publisher and market-level policy choices.

Third-order effects

  • If blocking persists, fresh news content is likely to shift from an assumed open-web input toward a controlled, potentially licensed input for AI products.
  • The pattern strengthens the early publisher move to block GPTBot and points to a more fragmented information layer for AI systems, where access terms may vary by publisher and geography.

The trend: AI crawling is becoming a publisher-controlled distribution channel, pushing model providers toward explicit access arrangements for high-value current content.