/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

A look at the Internet Archive, as the preservation project grows from 2TB in 1997 to ~100PB now, and issues with scraping paywalled news sites and social media

Within the walls of a beautiful former church in San Francisco's Richmond district, racks of computer servers hum and blink with activity. Tweets: @brianroemmele and @ftmag Tweets: Brian Roemmele / @brianroemmele : The @internetarchive is perhaps one of the most important archives of this epoch. The 100 Petabytes of data it has allows for research into just about any subject that would have otherwise been impossible. Even with this great work, we lose more data than is created every day. https://twitter.com/... @ftmag : 'Consider how hollow coverage of Queen Elizabeth's death would have been had it not been illustrated with profound archival material' — @DaveLeeFT on preserving the internet's backpages https://www.ft.com/...

Financial Times Dave Lee

Context & Ripple Effects

The Internet Archive's storage curve tells the story of the modern web: from 2TB in 1997 to 22 petabytes by 2018, past 70PB+ and 635B webpages at its 25th anniversary, to roughly 100PB today. Each jump tracked the web's own expansion — and now tracks its closing.

That closing is the friction point in this FT piece: paywalled news sites and social media platforms are harder to scrape than the open web the Archive was built on, arriving just as founder Brewster Kahle is already fighting existential copyright battles with music labels like UMG. The NPR coverage of expunged government pages shows why the mission matters even as access narrows.

First-order effects

  • Paywalled news sites and social media platforms restricting scraping directly shrink what the Archive's crawlers can capture, so the historical record goes thin exactly where public discourse now happens.
  • Brewster Kahle's team faces a second front: on top of the label-driven copyright fights, preservation of news and social content now requires negotiating with publishers who treat their archives as revenue assets.

Second-order effects

  • Researchers who relied on the Wayback Machine for deleted tweets and vanished articles lose a substitute record, pushing academic and journalistic demand toward paid archival services or platform-run transparency tools.
  • If closed platforms keep blocking crawlers, other memory institutions face the same wall, raising the price of entry for any competitor trying to do what the Archive does at nonprofit scale.

Third-order effects

  • If the pattern holds, the default state of the web becomes ephemeral: content behind paywalls and login walls decays without an independent copy, making institutions like the Internet Archive de facto critical infrastructure whose survival depends on legal accommodation rather than technical capability alone.
  • Preservation rights may need formalizing — either through licensing deals between archives and platforms or regulatory carve-outs — because goodwill-based scraping cannot scale against walled gardens.

The trend: As the web closes into paywalls and walled platforms, independent archives are shifting from passive crawlers to negotiated custodians of the public record.