Analysis: at 2023's end, 48% of the top news websites in the US, the UK, and eight others had blocked OpenAI's crawler and 24% had blocked Google's AI crawler
In this factsheet we describe the proportion of top news websites in ten countries that block AI (artificial intelligence) crawlers. X: @carnage4life , @risj_oxford , @risj_oxford , @risj_oxford , @richrdfletcher , @risj_oxford , @risj_oxford , and @risj_oxford See also Mediagazer X: Dare Obasanjo / @carnage4life : Access to online news continues down the trend of being pay to play. 48% of the most popular news sites across ten countries were blocking OpenAI's crawlers and 24% block Google's AI crawler. AI chatbots that wants to have current info will need to pay. https://reutersinstitute.politics.ox.ac .uk/ ... @risj_oxford : 1️⃣ First, the big picture: By the end of 2023, 48% of the top news sites across ten countries were blocking @OpenAI crawlers. Only 24% blocked @Google AI crawler. 97% of the sites that decided to block Google's crawler were blocking OpenAI's too #GenAI https://reutersinstitute.politics.ox.ac .uk/ ... @risj_oxford : How many news sites block AI crawlers from @Google and @OpenAI? This is the question at the heart of a new factsheet by @richrdfletcher, who looks at the 15 top news sites in 10 countries covered by our #DNR23 📱Full factsheet https://reutersinstitute.politics.ox.ac .uk/ ... 🧵5 findings in thread #AI [image] @risj_oxford : 3️⃣ Figures were much lower for @google AI crawler: The proportion blocking this AI crawler ranged from 60% in Germany to 7% in Poland and Spain. In general, outlets in the Global North were more likely to be blocking than those in the Global South 📈 Full figures below [image] Richard Fletcher / @richrdfletcher : New @risj_oxford factsheet by me that asks: How many news websites block generative AI like ChatGPT and Gemini from using their content to train their models? It depends on the country. Very large differences in how many top news sites are blocking, and how soon they started. [image] @risj_oxford : 4️⃣ During 2023, none of the websites we examined had reversed their decision after deciding to block. News outlets with a relatively large online news reach were slightly more likely to be blocking AI crawlers than those with a relatively small reach. See the chart below [image] @risj_oxford : 5️⃣ All types of news outlet were blocking, but the websites of legacy print publications were more likely to be blocking than those of either broadcasters or digital-born outlets 📊 Figures in the chart below [image] @risj_oxford : 2️⃣ The headline figures mask very large differences by country: The proportion of top online news websites blocking @OpenAI ranged from 79% in the US, to just 20% in Mexico and Poland. 📈 You can see figures for every country (and how blocking evolved) in the chart below [image] See also Mediagazer
Context & Ripple Effects
Publishers had already begun using robots.txt after OpenAI introduced GPTBot and its opt-out mechanism, with outlets including the New York Times and CNN among the early blockers of GPTBot. This cross-country count shows that response had become widespread rather than limited to a few prominent publishers.
The overlap in blocking decisions—nearly all sites blocking Google’s AI crawler also blocked OpenAI’s—suggests publishers were treating AI crawling as a distinct rights and distribution question, not merely a dispute with one company.
First-order effects
- OpenAI faces reduced access to content from leading news sites in the surveyed markets, while Google’s AI crawler faces a smaller but materially overlapping set of restrictions.
- News publishers that block the crawlers retain more control over whether their reporting is available for AI collection; no surveyed site had reversed a blocking decision during 2023.
Second-order effects
- AI providers seeking current news coverage will have greater incentive to negotiate permission or find other sources, while publishers can differentiate between open web indexing and AI-specific access.
- The uneven country results—from high Google-crawler blocking in Germany to low blocking in Spain—make news availability to AI systems more dependent on publisher and market-level policy choices.
Third-order effects
- If blocking persists, fresh news content is likely to shift from an assumed open-web input toward a controlled, potentially licensed input for AI products.
- The pattern strengthens the early publisher move to block GPTBot and points to a more fragmented information layer for AI systems, where access terms may vary by publisher and geography.
The trend: AI crawling is becoming a publisher-controlled distribution channel, pushing model providers toward explicit access arrangements for high-value current content.