/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

Originality AI: 23 major news websites and Reddit currently block the Internet Archive's crawler; journalists and advocacy groups sign a letter backing the IA

As major news outlets cut off the Wayback Machine, journalists and advocacy groups are rallying to protect the Internet Archive's vast collection of web pages.

Wired Kate Knibbs

Context & Ripple Effects

This dispute extends a boundary that publishers had already drawn around AI crawlers: major outlets had restricted automated access for model training, while some publishers also blocked Common Crawl alongside OpenAI’s bot. The issue is now archival access, not only training-data access.

Reddit had already limited the Archive to its homepage after saying AI firms were scraping through the service; the wider set of blocks turns that earlier Reddit restriction into a broader test of whether publisher bot controls can also constrain the public web record. The support letter gives the Archive a visible constituency as it faces wider copyright pressure.

First-order effects

  • The Wayback Machine will be unable to add new captures from the 23 named news sites and Reddit through its crawler, creating gaps in its forward-looking record of those publishers’ pages.
  • The letter puts journalists and advocacy groups publicly behind the Internet Archive, strengthening its case that archival crawling serves uses distinct from AI data collection.

Second-order effects

  • Publishers’ anti-AI crawl policies can increasingly affect preservation tools when those tools share the same automated-access channel, forcing archives and site owners to negotiate bot-specific exceptions rather than rely on a single robots policy.
  • Researchers, reporters, and the public may have fewer independently preserved versions of blocked pages, while publishers gain more control over what remains accessible through their own sites.

Third-order effects

  • If blocking spreads, control of web history could shift from broad third-party archives toward publishers and licensed repositories, making the durability of online evidence more dependent on individual site policies.
  • The conflict points to a need for clearer distinctions between model-training crawlers and public-interest archival crawlers; without them, AI-era access controls may become a general-purpose enforcement surface for the web.

The trend: AI-driven publisher control over automated access is expanding into a broader contest over who may preserve, reuse, and verify the public web.

Discussion

  • @knibbs @knibbs on x
    Journalists know that losing the Wayback Machine would be a nightmare: https://www.wired.com/...
  • @stinasdemons.com @stinasdemons.com on bluesky
    It is imperative that this resource be saved  —  The short-sightedness of media saying their reticence to allow/support the Internet Archive is their fear of “AI” scraping would be funny if this wasn't so dire  —  Like, are you familiar w literally everything now?  This is not th…
  • r/DataHoarder r on reddit
    The Internet's Most Powerful Archiving Tool Is in Peril
  • r/technology r on reddit
    The Internet's Most Powerful Archiving Tool Is in Peril
  • r/internetarchive r on reddit
    The Internet's Most Powerful Archiving Tool Is in Peril
  • @fightforthefuture.org @fightforthefuture.org on bluesky
    With many newspapers closed & no clear path for local public libraries to preserve digital-only reporting, the work of safeguarding journalism's record increasingly falls to  —  @archive.org - and their ability to preserve knowledge for the public good is under attack  —  www.wir…