/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

The Internet Archive stores a total of 22 petabytes, adds four petabytes per year, and uses ~7K processes crawling the web to capture 1.5B items per week

Nathan Mattise / Ars Technica : Tweets: @louisgray Tweets: Louis Gray / @louisgray : The @internetarchive has 22 petabytes of archived information. With proper backups, that's 44 petabytes of data. Their mission is incredibly important. http://arstechnica.com/...

Ars Technica Nathan Mattise

Context & Ripple Effects

This 2018 snapshot captures the Internet Archive mid-growth: 22 petabytes stored, four more added each year, and roughly 7,000 crawling processes capturing 1.5 billion items weekly. The later coverage shows how conservative those numbers were — by its 25th anniversary profile the Archive held over 70PB, and a 2022 look put it near 100PB, up from just 2TB in 1997.

Scale is only half the story. The same corpus documents the legal counterpressure: existential copyright fights with labels like UMG, half a million books pulled from Open Library, a settlement with major music publishers, and a DDoS attack and breach that briefly took the Wayback Machine offline.

First-order effects

  • The ~7,000-process crawl operation defines what gets preserved at all — every paywall, robots exclusion, or site redesign in 2018 decides which of the 1.5B weekly items make it into the permanent record.
  • Storage economics compound directly: at four petabytes of new data per year, doubled by backup requirements, the nonprofit's infrastructure budget scales linearly with its mission.

Second-order effects

  • Copyright holders' lawsuits force collection shrinkage regardless of technical capacity — the removal of 500,000+ Open Library books and the Great 78 Project settlement show legal exposure, not storage, becoming the binding constraint.
  • As commercial platforms lock down or delete content, demand shifts toward the Archive as the fallback record, pulling in ~100TB of uploaded data daily by 2025 and raising the stakes of every outage like the Wayback Machine breach.

Third-order effects

  • The Archive is hardening into civic memory infrastructure: its cataloging of ~73,000 US government webpages expunged by the Trump administration positions it as the check against state-level erasure, a role no commercial actor is incentivized to fill.
  • If the trajectory from 22PB to ~100PB holds, web preservation becomes an exabyte-scale undertaking whose viability depends on resolving the copyright framework — either through settlements and licensing or through legal reform — rather than on storage costs alone.

The trend: Web archiving is scaling from a petabyte-era preservation hobby into critical civic infrastructure, with copyright litigation rather than storage capacity setting the pace of what survives.