Cloudflare launches a tool that aims to block bots from scraping websites for AI training data, available free for all customers
Cloudflare, the publicly traded cloud service provider, has launched a new, free tool to prevent bots from scraping websites hosted on its platform for data to train AI models.
TechCrunchKyle Wiggers
Context & Ripple Effects
Cloudflare had already expanded from web infrastructure into AI deployment with tools for customers to run AI models. This move addresses the other side of that ecosystem: who controls the web data those models collect.
The free blocking control became an early step in a broader policy layer, later joined by AI-bot auditing tools and a crawl-payment marketplace. That arc makes the launch consequential beyond a single bot-setting feature.
First-order effects
Cloudflare customers can immediately deny AI-training crawlers access to sites hosted on its network without buying a separate product.
AI model developers using affected crawlers can lose access to content from customers that activate the control, making site-level permission a practical constraint on collection.
Second-order effects
Making the control free lowers the barrier for publishers and other site operators to adopt a default-deny stance, increasing pressure on AI crawlers to identify themselves and honor site preferences.
The feature establishes a foundation for more granular monitoring and monetization tools, including Cloudflare's later Pay per Crawl marketplace, rather than treating crawling as a purely technical traffic-management issue.
Third-order effects
If widely adopted, AI training-data access shifts from an implicit open-web assumption toward a permissioned market in which infrastructure providers enforce publisher rules.
Control over crawler identity and access can become a strategic layer of web infrastructure, with standards and platform policies shaping which AI companies can collect data at scale.
The trend: AI-data collection is becoming an enforceable publisher-rights and infrastructure-control issue rather than an ungoverned byproduct of web crawling.
An uncommon new service from Cloudflare, to block AI Crawlers on the distribution layer of the web: — “We hear clearly that customers don't want AI bots visiting their websites, and especially those that do so dishonestly. To help, we've added a brand new one-click to block al…
To help preserve a safe Internet for content creators, we've just launched a brand new “easy button” to block all AI bots. It's available for all customers, including those on our free tier. Read our blog post for more details: https://blog.cloudflare.com/ ...
I've decided to move my blog as well over to Cloudflare - I block AI crawlers in robots.txt, but clearly this is not respected. Almost all AI crawlers are a parasocial relationship, where they offer no value to any website: it's extraction in exchange for nothing (by design.)
@Cloudflare I'm impressed by Cloudflare's proactive move to protect content creators from AI bots! This easy button solution is a game-changer, especially with the rising concerns about AI-generated content.
@Cloudflare Seems like a nice feature but I'm skeptical about this claim. How would you know if you got a false negative if you don't know the ground truth? [image]
Kudos to Cloudflare for this. The AI companies need to act responsibly and ethically. Hoovering up terabytes of data from the public web without the consent of content-creators does not constitute fair use. #cloudflare #llms #Ai https://blog.cloudflare.com/ ...
Cloudflare has launched a new feature to block AI bots, scrapers, and crawlers with a single click, and it's free. As AI crawlers continue to swallow up web content, this tool helps protect your content from being used without consent. Many AI crawlers ignore robots txt [video]
Two things: 1. There's absolute explosion of bots crawling the web due to GenAI tools. If you want to drive revenue from these platforms you need to ensure bots can find your content 2. I think the wholesale blocking of all GenAI platforms is a mistake https://blog.cloudflare.com…
Interesting chart from Cloudflare. The top 3 AI crawlers are from Bytedance (TikTok), OpenAI and Anthropic. You can now block them with a single toggle using Cloudflare https://blog.cloudflare.com/ ... [image]
Bytespider, Amazonbot, ClaudeBot, and GPTBot Top four AI web crawlers... Some of these AI scrapers do not respect the robots.txt file and scrape all your content. To prevent this, Cloudflare has launched a new free tool to block AI bots from scraping your website. You can [image]
Notice how many companies w creatives or creators as paying customers have misread what those customers want from AI. They attempt to train their own AIs on their paying customers' work, without seeking opt-in (and enraging customers.) Cloudflare, again, showing how it's done.
Watch @Cloudflare become the most valuable company in the world. I am activating this feature right away. @thatkatieberry and friends you need to see this. https://blog.cloudflare.com/ ...
Cloudflare has launched a new feature allowing customers to block AI bots, scrapers, and crawlers. Decision-makers must now consider how this could affect the quality and breadth of AI training data, potentially impacting future #AI tools. https://blog.cloudflare.com/ ...
A company that “gets” its customers: Cloudflare offers functionality to block AI crawlers, which give zero value to any site (in fact, take away value: the more they crawl, the less traffic the site will later get!) They also consume resources (bandwidth, CPU, etc.) Great move: