The Internet Archive stores a total of 22 petabytes, adds four petabytes per year, and uses ~7K processes crawling the web to capture 1.5B items per week
Nathan Mattise / Ars Technica : Tweets: @louisgray Tweets: Louis Gray / @louisgray : The @internetarchive has 22 petabytes of archived information. With proper backups, that's 44 petabytes of data. Their mission is incredibly important. http://arstechnica.com/...
Context & Ripple Effects
This 2018 snapshot captures the Internet Archive mid-growth: 22 petabytes stored, four more added each year, and roughly 7,000 crawling processes capturing 1.5 billion items weekly. The later coverage shows how conservative those numbers were — by its 25th anniversary profile the Archive held over 70PB, and a 2022 look put it near 100PB, up from just 2TB in 1997.
Scale is only half the story. The same corpus documents the legal counterpressure: existential copyright fights with labels like UMG, half a million books pulled from Open Library, a settlement with major music publishers, and a DDoS attack and breach that briefly took the Wayback Machine offline.
First-order effects
- The ~7,000-process crawl operation defines what gets preserved at all — every paywall, robots exclusion, or site redesign in 2018 decides which of the 1.5B weekly items make it into the permanent record.
- Storage economics compound directly: at four petabytes of new data per year, doubled by backup requirements, the nonprofit's infrastructure budget scales linearly with its mission.
Second-order effects
- Copyright holders' lawsuits force collection shrinkage regardless of technical capacity — the removal of 500,000+ Open Library books and the Great 78 Project settlement show legal exposure, not storage, becoming the binding constraint.
- As commercial platforms lock down or delete content, demand shifts toward the Archive as the fallback record, pulling in ~100TB of uploaded data daily by 2025 and raising the stakes of every outage like the Wayback Machine breach.
Third-order effects
- The Archive is hardening into civic memory infrastructure: its cataloging of ~73,000 US government webpages expunged by the Trump administration positions it as the check against state-level erasure, a role no commercial actor is incentivized to fill.
- If the trajectory from 22PB to ~100PB holds, web preservation becomes an exabyte-scale undertaking whose viability depends on resolving the copyright framework — either through settlements and licensing or through legal reform — rather than on storage costs alone.
The trend: Web archiving is scaling from a petabyte-era preservation hobby into critical civic infrastructure, with copyright litigation rather than storage capacity setting the pace of what survives.