/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

Developers say aggressive AI crawlers are overwhelming open-source infrastructure; LibreNews: up to 97% of some projects' traffic comes from AI companies' bots

Software developer Xe Iaso reached a breaking point earlier this year when aggressive AI crawler traffic from Amazon overwhelmed …

Ars Technica Benj Edwards

Context & Ripple Effects

This report extends an established pattern of AI-crawler load causing operational harm beyond publishers: a relentless crawl that took down an e-commerce site had already shown how mismanaged bot access can exhaust a smaller operator's capacity.

The pressure is especially consequential for open source because its public-facing infrastructure is often maintained with limited operational slack. It also sits alongside a widening move to restrict crawlers, including news outlets blocking AI web crawlers.

First-order effects

  • Open-source projects receiving bot-dominated traffic must absorb higher infrastructure load and operational attention, with maintainers such as Xe Iaso facing disruption to resources intended for human users.
  • Amazon and other AI companies named by developers become immediate targets for demands that their crawlers reduce load or better respect access controls.

Second-order effects

  • Project operators are likely to add friction or filtering for automated visitors, drawing attention to defenses such as open-source anti-scraping tools and crawler traps and potentially making legitimate automated access harder.
  • Hosting and edge-security providers gain a larger role in separating AI collection traffic from normal project use, shifting an infrastructure-management burden onto maintainers and their service vendors.

Third-order effects

  • If AI data collection continues to impose costs on volunteer-run projects, open-source availability may increasingly depend on paid traffic controls and formal bot-access policies rather than default openness.
  • The pattern points toward a broader negotiation over who pays for access to public technical knowledge: AI firms, infrastructure intermediaries, or the communities maintaining the underlying resources.

The trend: AI training-data collection is turning open web and open-source access from a default public good into a contested infrastructure and cost-allocation problem.

Discussion

  • @kirb.me Adam Demasi on bluesky
    Exact same story for us.  I run a wiki at a loss and don't want to accept donations, but AI bots who know nothing about MediaWiki and just blindly follow every link forced me to upgrade servers, buy Cloudflare Pro, and spend weeks writing firewall rules to block the bad guys and …
  • @cyberciti.biz @cyberciti.biz on bluesky
    FOSS infrastructure is under attack by AI companies thelibre.news/foss-infrast... Please share for awareness, reach and to public shame Microsoft, Meta, OpenAI, Perplexity and other such AI companies.  It is madness out there.
  • @fastly.com @fastly.com on bluesky
    If you're a FOSS project dealing with overwhelming AI scraper bots, we will provide free security services for your project at no cost to you ❤️ #TeamSleep thelibre.news/foss-infrast...
  • @iethics @iethics on bluesky
    “It remains unclear why these companies don't adopt more collaborative approaches and... rate-limit their #data harvesting runs so they don't overwhelm source websites.  #Amazon, #OpenAI, #Anthropic, and #Meta did not immediately respond to requests for comment”: arstechnica.com/…
  • @alexmontgomery Alexander H. Montgomery on bluesky
    If your project requires doing something like this (immoral if not illegal) to function, your project is not sustainable, and deserves to be stuck in a Nepenthes-style trap.
  • @xeiaso.net @xeiaso.net on bluesky
    Don't put in the news that I got mad, because they put in the news that I got even: arstechnica.com/ai/2025/03/d...
  • @ilpeach Mr Peach on bluesky
    This pisses me off so much.  Effectively burning to the ground all the good will behind this efforts.  “FOSS infrastructure is under attack by AI companies”.  Another reason for hating this BS fad.
  • @kristi.nikolla.me Kristi Nikolla on bluesky
    AI crawlers are costing open source projects thousands of dollars in infrastructure costs, and frequently bringing everything offline in one massive constant never-ending DDoS.  This is not OK!
  • @akkartik.name Kartik Agaram on bluesky
    I don't really understand why AI crawlers have never respected robots.txt.  It's been several years at this point.  Anyone have perspectives to explain?  (Not justify; goes without saying that this is terrible for the open web and also for their own self-interest.)  —  thelibre.n…
  • @nishtahir.com Nish Tahir on bluesky
    If you run AI infrastructure be a good neighbor and remember that a lot of infrastructure out there is run by hobbyists and volunteers. thelibre.news/foss-infrast...
  • @stsquad@mastodon.org.uk Alex on mastodon
    Personally I don't mind my code being ingested to train #LLM models.  Freedoms 1 and 3 of the four essential software freedoms allow for study and redistribution of modified versions of code.  Of course those freedoms don't allow for stripping the license obligations from derivat…
  • @bagder@mastodon.social @bagder@mastodon.social on mastodon
    The AI bots that desperately need OSS for code training, are now slowly killing OSS by overloading every site.  —  The curl website is now at 77TB/month, or 8GB every five minutes.  —  https://arstechnica.com/...
  • @veronica@mastodon.online Veronica Olsen on mastodon
    Who could have guessed that an industry whose entire business model is based on theft would behave like malware attacks on the Internet?  🤔  —  https://arstechnica.com/...  #AI #DDoS #Crawlers
  • @frimelle Lucie-Aimée Kaffee on x
    I've been part of open source for years, it's worrying to see AI scrapers straining the infrastructure we all rely on. FOSS projects are seeing outages, fake bug reports, and mounting pressure. AI must give back, not just take. 📖 https://thelibre.news/...
  • @geeknik @geeknik on x
    Imagine your community garden being harvested by robots that then sell the vegetables back to you—with rootkit seasoning. https://thelibre.news/...
  • @bibryam Bilgin Ibryam on x
    “In 2.5 hours, we had 81k requests—97% were bots.” FOSS is under attack by AI companies https://thelibre.news/...
  • r/artificial r on reddit
    Open Source devs say AI crawlers dominate traffic, forcing blocks on entire countries