/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

Some popular sites like Condé Nast's titles and Reuters.com modified robots.txt to block Anthropic's bots, but Anthropic has just made new bots with other names

We really are going to need a shared blocklist that doesn't rely on putting your website behind Cloudflare.  —  https://www.404media.co/... Jason Koebler / @jasonkoebler@mastodon.social : Many websites think they're blocking Anthropic's scrapers but are actually blocking two old bots and NOT the real one  —  New and renamed AI scraper bots are coming out constantly, and keeping track of all of them is very difficult  —  https://www.404media.co/... X: Glenn Gabe / @glenngabe : Well, that's one way to get around robots.txt directives :) It's ClaudeBot now Neil Turkewitz / @neilturkewitz : An illuminating look at why reliance on opt-outs such as those expressed in robots.txt is misplaced & inequitable. Placing the obligation on creators & website owners is unworkable. AI companies MUST obtain affirmative consent. Period. ⁦@jason_koebler https://www.404media.co/... Paris Marx / @parismarx : More evidence that out-opt provisions for generative AI training simply are not enough. Explicit opt in is the only thing that should be acceptable, if we agree these kinds of general models should be pursued despite their immense environmental cost at all. Grady Booch / @grady_booch : Why is consent just a hard thing for techbros to understand? https://www.404media.co/... LinkedIn: Toshit Panigrahi : More fantastic reporting by 404 Media on robots.txt scraping and the failings of the status quo.  New user agents pop up constantly. … See also Mediagazer

404 Media Jason Koebler

Context & Ripple Effects

Publishers had already moved to restrict AI crawlers: one earlier tally found that many leading US news outlets were blocking AI web crawlers. This episode shows why a static user-agent denylist can fail when the identifiers being blocked are no longer the ones in use.

The story also sharpens the limits of opt-out enforcement. Cloudflare had introduced a free tool to block AI-training scrapers, while later coverage examined challenges and other defenses; both reflect a shift from relying on robots.txt alone toward operational bot controls.

First-order effects

  • Condé Nast properties and Reuters.com that targeted older Anthropic user agents may not block Anthropic’s current scrapers, leaving their intended crawl restrictions ineffective until rules are updated.
  • Site operators must continually identify and maintain bot signatures; Anthropic’s renamed bots can continue requesting accessible pages where existing rules do not match them.

Second-order effects

  • Bot-management providers gain importance as publishers seek controls that can be updated and applied centrally rather than maintaining fragmented robots.txt lists.
  • Other AI firms and publishers face greater scrutiny over crawler identity and disclosure, because a consent choice is hard to execute when the technical identifier changes.

Third-order effects

  • If crawler names remain fluid, robots.txt becomes a weaker practical mechanism for governing AI data collection, increasing pressure for authenticated access, shared blocklists, or affirmative-consent systems.
  • Control over AI crawling may consolidate around infrastructure intermediaries that can detect and filter traffic at scale, rather than individual site operators managing directives themselves.

The trend: AI-content access is moving from voluntary, per-site crawler directives toward enforceable and centralized publisher controls.

Discussion

  • @jasonkoebler@mastodon.social Jason Koebler on mastodon
    Many websites think they're blocking Anthropic's scrapers but are actually blocking two old bots and NOT the real one  —  New and renamed AI scraper bots are coming out constantly, and keeping track of all of them is very difficult  —  https://www.404media.co/...
  • @parismarx Paris Marx on x
    More evidence that out-opt provisions for generative AI training simply are not enough. Explicit opt in is the only thing that should be acceptable, if we agree these kinds of general models should be pursued despite their immense environmental cost at all.
  • @grady_booch Grady Booch on x
    Why is consent just a hard thing for techbros to understand? https://www.404media.co/...
  • @glenngabe Glenn Gabe on x
    Well, that's one way to get around robots.txt directives :) It's ClaudeBot now
  • @neilturkewitz Neil Turkewitz on x
    An illuminating look at why reliance on opt-outs such as those expressed in robots.txt is misplaced & inequitable. Placing the obligation on creators & website owners is unworkable. AI companies MUST obtain affirmative consent. Period. ⁦@jason_koebler https://www.404media.co/...