Some popular sites like Condé Nast's titles and Reuters.com modified robots.txt to block Anthropic's bots, but Anthropic has just made new bots with other names
We really are going to need a shared blocklist that doesn't rely on putting your website behind Cloudflare. — https://www.404media.co/... Jason Koebler / @jasonkoebler@mastodon.social : Many websites think they're blocking Anthropic's scrapers but are actually blocking two old bots and NOT the real one — New and renamed AI scraper bots are coming out constantly, and keeping track of all of them is very difficult — https://www.404media.co/... X: Glenn Gabe / @glenngabe : Well, that's one way to get around robots.txt directives :) It's ClaudeBot now Neil Turkewitz / @neilturkewitz : An illuminating look at why reliance on opt-outs such as those expressed in robots.txt is misplaced & inequitable. Placing the obligation on creators & website owners is unworkable. AI companies MUST obtain affirmative consent. Period. @jason_koebler https://www.404media.co/... Paris Marx / @parismarx : More evidence that out-opt provisions for generative AI training simply are not enough. Explicit opt in is the only thing that should be acceptable, if we agree these kinds of general models should be pursued despite their immense environmental cost at all. Grady Booch / @grady_booch : Why is consent just a hard thing for techbros to understand? https://www.404media.co/... LinkedIn: Toshit Panigrahi : More fantastic reporting by 404 Media on robots.txt scraping and the failings of the status quo. New user agents pop up constantly. … See also Mediagazer
Context & Ripple Effects
Publishers had already moved to restrict AI crawlers: one earlier tally found that many leading US news outlets were blocking AI web crawlers. This episode shows why a static user-agent denylist can fail when the identifiers being blocked are no longer the ones in use.
The story also sharpens the limits of opt-out enforcement. Cloudflare had introduced a free tool to block AI-training scrapers, while later coverage examined challenges and other defenses; both reflect a shift from relying on robots.txt alone toward operational bot controls.
First-order effects
- Condé Nast properties and Reuters.com that targeted older Anthropic user agents may not block Anthropic’s current scrapers, leaving their intended crawl restrictions ineffective until rules are updated.
- Site operators must continually identify and maintain bot signatures; Anthropic’s renamed bots can continue requesting accessible pages where existing rules do not match them.
Second-order effects
- Bot-management providers gain importance as publishers seek controls that can be updated and applied centrally rather than maintaining fragmented robots.txt lists.
- Other AI firms and publishers face greater scrutiny over crawler identity and disclosure, because a consent choice is hard to execute when the technical identifier changes.
Third-order effects
- If crawler names remain fluid, robots.txt becomes a weaker practical mechanism for governing AI data collection, increasing pressure for authenticated access, shared blocklists, or affirmative-consent systems.
- Control over AI crawling may consolidate around infrastructure intermediaries that can detect and filter traffic at scale, rather than individual site operators managing directives themselves.
The trend: AI-content access is moving from voluntary, per-site crawler directives toward enforceable and centralized publisher controls.