Internet Archive says it is using Cloudflare's reconfigured Always On caching tools to improve its Wayback Machine archiving
Lily Hay Newman / Wired :
Context & Ripple Effects
By 2020 the Wayback Machine had become one of the web's core public records — the preservation project had grown from 2TB in 1997 to roughly 100PB — but it still captured pages by crawling, which put it at the mercy of whatever sat between the crawler and the page. Cloudflare's Always On tooling sits exactly there: when Cloudflare reconfigures how it caches and serves origin content, it changes what an outside crawler can see at all.
That dependency has only sharpened since. The Archive later weathered a [[a:877378|DDoS attack and data breach that forced the Wayback Machine back online only in provisional read-only form]], and by 2026 Originality AI counted 23 major news websites plus Reddit blocking the Archive's crawler outright, with journalists and advocacy groups signing a letter in its defense. This 2020 partnership reads, in hindsight, as the cooperative counterpoint to that adversarial drift.
First-order effects
- The Internet Archive gets better capture fidelity on the large share of the web served through Cloudflare, since Always On caching keeps content retrievable even when origins are slow or down.
- Cloudflare gains a marquee civic-credentials use case for Always On, positioning its network configuration choices as part of the public record rather than purely a commercial CDN feature.
Second-order effects
- Publishers' infrastructure decisions quietly become archival policy: a site's Cloudflare settings now help determine whether the Wayback Machine can preserve it — a lever some outlets later used to block the crawler entirely.
- Other edge and hosting providers face a soft precedent: cooperating with the Archive becomes the visible benchmark, making refusal to support archival access harder to defend without explanation.
Third-order effects
- If the pattern holds, preservation of the web stops being an independent-crawler project and becomes a negotiated arrangement with a handful of infrastructure intermediaries — raising questions about whether access for archives should be a default of the edge layer or something requiring policy guarantees.
- The Archive's growing reliance on external platforms compounds its structural fragility already exposed by the breach and the copyright fights with labels like UMG: a commons whose completeness depends on other companies' goodwill.
The trend: Web archiving is shifting from autonomous crawling toward negotiated cooperation with the infrastructure layer that controls what crawlers can reach — with blocking campaigns pushing the other way.