/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

Internet Archive says it is using Cloudflare's reconfigured Always On caching tools to improve its Wayback Machine archiving

Lily Hay Newman / Wired :

Wired Lily Hay Newman

Context & Ripple Effects

By 2020 the Wayback Machine had become one of the web's core public records — the preservation project had grown from 2TB in 1997 to roughly 100PB — but it still captured pages by crawling, which put it at the mercy of whatever sat between the crawler and the page. Cloudflare's Always On tooling sits exactly there: when Cloudflare reconfigures how it caches and serves origin content, it changes what an outside crawler can see at all.

That dependency has only sharpened since. The Archive later weathered a [[a:877378|DDoS attack and data breach that forced the Wayback Machine back online only in provisional read-only form]], and by 2026 Originality AI counted 23 major news websites plus Reddit blocking the Archive's crawler outright, with journalists and advocacy groups signing a letter in its defense. This 2020 partnership reads, in hindsight, as the cooperative counterpoint to that adversarial drift.

First-order effects

  • The Internet Archive gets better capture fidelity on the large share of the web served through Cloudflare, since Always On caching keeps content retrievable even when origins are slow or down.
  • Cloudflare gains a marquee civic-credentials use case for Always On, positioning its network configuration choices as part of the public record rather than purely a commercial CDN feature.

Second-order effects

  • Publishers' infrastructure decisions quietly become archival policy: a site's Cloudflare settings now help determine whether the Wayback Machine can preserve it — a lever some outlets later used to block the crawler entirely.
  • Other edge and hosting providers face a soft precedent: cooperating with the Archive becomes the visible benchmark, making refusal to support archival access harder to defend without explanation.

Third-order effects

  • If the pattern holds, preservation of the web stops being an independent-crawler project and becomes a negotiated arrangement with a handful of infrastructure intermediaries — raising questions about whether access for archives should be a default of the edge layer or something requiring policy guarantees.
  • The Archive's growing reliance on external platforms compounds its structural fragility already exposed by the breach and the copyright fights with labels like UMG: a commons whose completeness depends on other companies' goodwill.

The trend: Web archiving is shifting from autonomous crawling toward negotiated cooperation with the infrastructure layer that controls what crawlers can reach — with blocking campaigns pushing the other way.

Discussion

  • @wired @wired on x
    The Internet Archive's Wayback Machine has been invaluable for maintaining a history of long-forgotten pages. Now its deep memory will help make sure the sites you visit never go down, through a partnership with internet infrastructure company Cloudflare. https://www.wired.com/..…
  • @evanderburg Eric Vanderburg on x
    Internet Archive's way cool Wayback Machine gets way more websites in Cloudflare fail-over deal https://i.securitythinkingcap.com/ Rgqssf https://twitter.com/...
  • @brittanymbrown Brittany Brown on x
    Excited to be partnering with the @internetarchive. @Cloudflare's Always Online™ now uses the Wayback Machine's extensive library to ensure visitors have access to content when origin servers are down: https://www.wired.com/...