/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

← → days · ↑ ↓ browse · Enter similar · o open

Freelancer.com, iFixit, and others say Anthropic's crawler has aggressively scraped their websites, potentially breaching their terms of service

Web publishers say developer is swarming their sites, collecting content to train models and ignoring orders to stop

Financial Times George Hammond

Context & Ripple Effects

This complaint sits in a widening dispute over whether web-access conventions meaningfully constrain AI training collection. Days later, major sites were reported to update robots.txt files to block Anthropic bots, while Anthropic was said to introduce newly named bots, escalating a technical access-control conflict publishers’ attempts to block Anthropic’s bots.

The stakes extend beyond any one crawler: later coverage describes Reddit alleging continued Anthropic access after it said scraping had stopped, against a backdrop of licensing arrangements with OpenAI and Google Reddit’s later scraping lawsuit and licensing contrast. That makes the treatment of publisher terms and stop requests consequential for how training data is sourced.

First-order effects

  • Freelancer.com, iFixit and the other complainants must absorb crawler traffic and decide whether to tighten technical blocks or pursue enforcement over alleged violations of their site terms.
  • Anthropic’s data-collection practices face immediate scrutiny from affected publishers, particularly around whether its bots honor requested access limits.

Second-order effects

  • Other publishers have an incentive to revise robots.txt policies and deploy more active anti-scraping measures; subsequent coverage of bot-name changes shows why static blocks can become an operational contest the shift from robots.txt blocks to newly named bots.
  • The dispute strengthens publishers’ leverage to seek explicit permissions or paid data arrangements rather than rely on informal crawler conventions.

Third-order effects

  • If crawler identification and robots.txt compliance remain contested, web data for model training is likely to move toward a mix of licensed sources, hardened access controls, and disputes over enforceable consent.
  • The durable industry question is whether publicly reachable content remains a broadly usable model input or becomes governed by publisher-controlled technical and contractual gates.

The trend: AI training is turning open-web crawling from a largely implicit practice into a negotiated market for content access and control.

Discussion

  • @kwiens Kyle Wiens on x
    Hey @AnthropicAI: I get you're hungry for data. Claude is really smart! But do you really need to hit our servers a million times in 24 hours? You're not only taking our content without paying, you're tying up our devops resources. Not cool.
  • @kwiens Kyle Wiens on x
    If any of those requests accessed our terms of service, they would have told you that use of our content expressly forbidden. But don't ask me, ask Claude! If you want to have a conversation about licensing our content for commercial use, we're right here. [image]
  • @bobkitten @bobkitten on x
    @kwiens @AnthropicAI Post a fake fixit for a fake device so you'll be able to track who misappropriated your data. Mapmakers insert fake towns and roads to detect copying, Trivial Pursuit had 2 fake Q&As out of 6k in the Genus Edition. “How to replace the battery in the Dipsogeni…
  • @jason_koebler Jason Koebler on x
    Anthropic's AI scraper bot hit iFixit's website a million times in a day in violation of its terms of service ... even if you specifically spell out that you don't want to have your stuff stolen it is often stolen anyway https://www.404media.co/...
  • @jordancmeyer Jordan Meyer on x
    A single scraper cost readthedocs over $5k in May by downloading 73TB of data. Most image, music, and video datasets for AI only contain links to the media files. Downloading them externalizes the network costs to the hosts... https://about.readthedocs.com/ ...
  • @glenngabe Glenn Gabe on x
    ifixit wasn't blocking via robots.txt though. Now they are, so it *should* stop :) -> Anthropic's crawler is ignoring websites' anti-AI scraping policies “ClaudeBot, the web crawler that Anthropic uses to scrape training data for AI models like Claude, has hammered iFixit's [imag…
  • @rahll Reid Southen on x
    Anthropic AI is ripping off @iFixit en masse against their ToS and tying up their resources. These thieving AI companies need to be held accountable, this shit is out of control.
  • @barelylingual @barelylingual on x
    @kwiens @AnthropicAI We've had to block what we assume are ai companies spamming our APIs instead of downloading our freely available database dumps over here at Wikipedia
  • @readthedocs @readthedocs on x
    AI crawlers have downloaded over 100 TB in the last couple months, partially from bugs and abusive behavior. We ask AI crawlers to include rate and bandwidth limits, and ask them to partner with us to mitigate impact and support our work. https://about.readthedocs.com/ ...
  • @kwiens Kyle Wiens on x
    @BobKitten @AnthropicAI Oh we're way past that https://chatgpt.com/...
  • @ericholscher Eric Holscher on x
    @kwiens @AnthropicAI Yea, they were hammering us over at @readthedocs as well. We were planning to write a blog post on it, since this behavior is definitely gonna get all AI crawlers blocked because of abuse, not even because of the copyright issues.
  • @matt_barrie Matt Barrie on x
    @mysticaltech ... Dude 80,000 requests every 5 minutes
  • @matt_barrie Matt Barrie on x
    @kwiens @AnthropicAI We have the same issue @AnthropicAI scrapers absolutely smashing our servers, feel like sending them a bill.
  • @kwiens Kyle Wiens on x
    @jrhunt @AnthropicAI Our TOS banned ML training before their crawl, afterwards we added them to robots.txt.
  • r/technology r on reddit
    AI start-up Anthropic accused of ‘egregious’ data scraping