Profile of Brewster Kahle and the Internet Archive, which marked its 25th anniversary in May and stores 70PB+ of data, including 635B webpages and 34M books
Meet Brewster Kahle, the internet's chief librarian — Audio player loading... On the same day in 1996, Brewster Kahle founded …Tweets:@ron_miller,@lanceulanoff, and@jkfruit.Thanks:@techradarproTweets:Ron Miller /@ron_miller:Just a tremendously useful resource. Without it, so much content would be lost forever. As an example, I found an early Salesforce screenshot for an article I wrote a couple of years ago on their 20th anniversary. https://twitter.com/...Lance Ulanoff /@lanceulanoff:“As wel
Context & Ripple Effects
Twenty-five years after Brewster Kahle founded the Internet Archive in 1996, the nonprofit has become the web's working memory: over 70 petabytes spanning 635 billion captured webpages and 34 million books. The growth curve is steep — back in 2018 the collection stood at just 22 petabytes, adding four per year via roughly 7,000 crawling processes, so the corpus has more than tripled since.
The anniversary profile lands mid-arc: within a few years the Archive would reach ~100 petabytes while drawing existential copyright challenges from music labels like UMG, and Kahle would later describe lawsuits that threatened to bankrupt the organization entirely (his interview on fair use and the Archive's future). The same scale that makes it indispensable is what puts it in court.
First-order effects
- Researchers and journalists get a de facto primary-source library: Ron Miller's cited example — recovering an early Salesforce screenshot for an anniversary article — shows the Wayback Machine functioning as infrastructure for tech history that companies themselves no longer host.
- Kahle's institution now operates at a scale where storage growth (from four petabytes added annually in 2018 to a 100-petabyte corpus by 2022) becomes a permanent fundraising and engineering commitment rather than a side project.
Second-order effects
- Rights holders like UMG respond to unlicensed mass copying not with takedowns but with litigation aimed at the nonprofit itself, converting a preservation project into a legal counterparty and forcing the Archive to defend fair use as a business-survival question.
- Publishers and platforms gain a new calculus: content deleted from the live web may persist in the Archive, which raises the cost of quietly rewriting corporate history — the Salesforce screenshot case is the template.
Third-order effects
- If the pattern holds, the decisive constraint on independent digital libraries shifts from storage economics (solved: terabytes to hundreds of petabytes in a generation) to copyright law — court outcomes will effectively set the ceiling on how much of the public web one nonprofit may preserve.
- A loss or settlement that curtails the Archive's lending and copying practices would leave web-scale preservation without an institutional home, pushing the function toward fragmented national libraries and commercial crawlers.
The trend: Web archiving is scaling far faster than the copyright framework governing it, making litigation rather than storage capacity the binding limit on independent preservation.