A look at some options for fighting AI companies' scraping, including open-source Anubis' cryptographic JavaScript math challenges and Cloudflare's AI Labyrinth
If you are a website owner you should definitely check it out! [embedded post] Brewster Kahle / @brewster.kahle.org : Interview of the new librarian of the British Library : — www.bloomberg.com/features/202... Q&As on AI: [image] @mkultra.monster : I'm already blocking AI scraping with Anubis, and although this is good, and I hope it kills AI scraping completely, I also am not sure if you should trust Cloudflare with your data, as they are an American company dealing with the likes of Google and Amazon [embedded post] A. J. Hondo / @ajh1138 : Thanks to AI and social media scraping, you're going to start seeing your own likeness in ads. I'm surprised this hasn't happened already. Business opportunity for people who like to creep other people out! Paris Marx / @parismarx.com : love this story about a canadian protecting us from ai bots 🥰 [embedded post] Mastodon: @mttaggart@infosec.exchange : Cool as hell to see @cadey get the recognition she deserves. — https://www.404media.co/... X: Arthur Nix / @thearthurnix : @HarmlessYardDog AI polluting the internet with slop and then scraping and training on said slop is helping too @bigfundu : 🧵 1/ The AI free ride is over Cloudflare (hosting ~20% of the web) just flipped the script on internet scraping. CF now blocks AI crawlers unless they pay up. This is a seismic shift for “free lunch” AI that's been scooping up web content without giving back. Here's why it [image] @navigate_ai : The Danish parliament passed a law that basically says: your face and your voice are yours. No more scraping selfies to feed someone's AI without asking. 🧵 [image] Julius Ruechel / @juliusruechel : Before AI you had thousands upon thousands of small websites making money by teaching how-to skills online, all competing with each other to rank at the top of Google search results so they could get eyeballs and ad revenue. AI is destroying that business model by scraping their @alexissfallon : @DoeEyedGirlie Honestly I think it's because of the AI scraping, lots of people used the site and never made accounts but authors are locking their work to protect from the scraping. so now there's a huge group of people who suddenly HAVE to make accounts to interact with those works Forums: r/technology : The Open-Source Software Saving the Internet From AI Bot Scrapers r/antiai : The Open-Source Software Saving the Internet From AI Bot Scrapers Beehaw : The Open-Source Software Saving the Internet From AI Bot Scrapers BeauHD / Slashdot : The Open-Source Software Saving the Internet From AI Bot Scrapers See also Mediagazer
Context & Ripple Effects
Website owners are moving beyond robots.txt toward active defenses. Cloudflare first offered a crawler-blocking tool, while publishers had already found that bot-name blocks could be sidestepped by newly named Anthropic crawlers.
The latest options arrive just after Cloudflare introduced a pay-per-crawl model and default AI-crawler blocking for new sites, framing scraping as both a technical access-control problem and a potential commercial transaction.
First-order effects
- Site operators can deploy Anubis-style JavaScript challenges or Cloudflare’s AI Labyrinth to raise the cost of automated collection, rather than relying solely on crawler directives.
- AI crawlers face more denied requests, deceptive paths, or payment gates on participating sites; Cloudflare customers can centralize those controls through one intermediary.
Second-order effects
- Crawler operators will have stronger incentives to identify themselves, negotiate access, or adapt collection methods as publisher defenses become more active.
- The split between open-source defenses and Cloudflare-managed controls makes the choice of anti-scraping vendor a consequential trade-off between operational convenience and reliance on a major infrastructure provider.
Third-order effects
- If these tools spread, web content access is likely to shift from an open crawling default toward explicit permission, payment, or adversarial blocking—a core evolution of Cloudflare’s AI-crawler controls.
- That transition could concentrate bargaining power with large hosting and delivery platforms unless interoperable, self-hosted defenses such as Anubis remain practical for smaller publishers.
The trend: AI training-data collection is becoming a contested access market in which publishers combine technical barriers with licensing and payment controls.