Developers say aggressive AI crawlers are overwhelming open-source infrastructure; LibreNews: up to 97% of some projects' traffic comes from AI companies' bots
Software developer Xe Iaso reached a breaking point earlier this year when aggressive AI crawler traffic from Amazon overwhelmed …
Ars TechnicaBenj Edwards
Context & Ripple Effects
This report extends an established pattern of AI-crawler load causing operational harm beyond publishers: a relentless crawl that took down an e-commerce site had already shown how mismanaged bot access can exhaust a smaller operator's capacity.
The pressure is especially consequential for open source because its public-facing infrastructure is often maintained with limited operational slack. It also sits alongside a widening move to restrict crawlers, including news outlets blocking AI web crawlers.
First-order effects
Open-source projects receiving bot-dominated traffic must absorb higher infrastructure load and operational attention, with maintainers such as Xe Iaso facing disruption to resources intended for human users.
Amazon and other AI companies named by developers become immediate targets for demands that their crawlers reduce load or better respect access controls.
Second-order effects
Project operators are likely to add friction or filtering for automated visitors, drawing attention to defenses such as open-source anti-scraping tools and crawler traps and potentially making legitimate automated access harder.
Hosting and edge-security providers gain a larger role in separating AI collection traffic from normal project use, shifting an infrastructure-management burden onto maintainers and their service vendors.
Third-order effects
If AI data collection continues to impose costs on volunteer-run projects, open-source availability may increasingly depend on paid traffic controls and formal bot-access policies rather than default openness.
The pattern points toward a broader negotiation over who pays for access to public technical knowledge: AI firms, infrastructure intermediaries, or the communities maintaining the underlying resources.
The trend: AI training-data collection is turning open web and open-source access from a default public good into a contested infrastructure and cost-allocation problem.
Exact same story for us. I run a wiki at a loss and don't want to accept donations, but AI bots who know nothing about MediaWiki and just blindly follow every link forced me to upgrade servers, buy Cloudflare Pro, and spend weeks writing firewall rules to block the bad guys and …
FOSS infrastructure is under attack by AI companies thelibre.news/foss-infrast... Please share for awareness, reach and to public shame Microsoft, Meta, OpenAI, Perplexity and other such AI companies. It is madness out there.
If you're a FOSS project dealing with overwhelming AI scraper bots, we will provide free security services for your project at no cost to you ❤️ #TeamSleep thelibre.news/foss-infrast...
“It remains unclear why these companies don't adopt more collaborative approaches and... rate-limit their #data harvesting runs so they don't overwhelm source websites. #Amazon, #OpenAI, #Anthropic, and #Meta did not immediately respond to requests for comment”: arstechnica.com/…
If your project requires doing something like this (immoral if not illegal) to function, your project is not sustainable, and deserves to be stuck in a Nepenthes-style trap.
This pisses me off so much. Effectively burning to the ground all the good will behind this efforts. “FOSS infrastructure is under attack by AI companies”. Another reason for hating this BS fad.
AI crawlers are costing open source projects thousands of dollars in infrastructure costs, and frequently bringing everything offline in one massive constant never-ending DDoS. This is not OK!
I don't really understand why AI crawlers have never respected robots.txt. It's been several years at this point. Anyone have perspectives to explain? (Not justify; goes without saying that this is terrible for the open web and also for their own self-interest.) — thelibre.n…
If you run AI infrastructure be a good neighbor and remember that a lot of infrastructure out there is run by hobbyists and volunteers. thelibre.news/foss-infrast...
Personally I don't mind my code being ingested to train #LLM models. Freedoms 1 and 3 of the four essential software freedoms allow for study and redistribution of modified versions of code. Of course those freedoms don't allow for stripping the license obligations from derivat…
The AI bots that desperately need OSS for code training, are now slowly killing OSS by overloading every site. — The curl website is now at 77TB/month, or 8GB every five minutes. — https://arstechnica.com/...
Who could have guessed that an industry whose entire business model is based on theft would behave like malware attacks on the Internet? 🤔 — https://arstechnica.com/... #AI #DDoS #Crawlers
I've been part of open source for years, it's worrying to see AI scrapers straining the infrastructure we all rely on. FOSS projects are seeing outages, fake bug reports, and mounting pressure. AI must give back, not just take. 📖 https://thelibre.news/...