/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

← → days · ↑ ↓ browse · Enter similar · o open

Harvard releases a high-quality dataset of nearly 1M public-domain books, created with funding from Microsoft and OpenAI, that anyone can use to train AI tools

The project's leader says that allowing everyone to access the collection of public-domain books will help “level the playing field” in the AI industry.

Wired Kate Knibbs

Context & Ripple Effects

The release creates a permissioned alternative to the contested book corpora that powered earlier models, including Books3’s use in training prominent AI systems. It also arrives as publishers explore licensed training access, as reflected in Microsoft’s reported HarperCollins arrangement.

Harvard’s later Institutional Books 1.0 release, built from 983,000 public-domain books, shows this collection becoming a durable research asset rather than a one-off archive.

First-order effects

  • Researchers, startups, and other developers gain a broadly available book corpus for training and evaluating AI tools without relying on opaque or allegedly pirated libraries.
  • Harvard becomes a steward of a shared AI-data resource; Microsoft and OpenAI’s funding supports an asset available beyond their own model-development efforts.

Second-order effects

  • The dataset gives smaller AI teams a lawful baseline corpus, reducing one barrier to experimentation even though it does not erase larger firms’ compute and distribution advantages.
  • It raises the comparative value of licensed and clearly governed text sources as copyright pressure constrains datasets such as the contested Books3 collection.

Third-order effects

  • If institutions continue publishing well-documented training corpora, AI competition may increasingly separate into open data access on one side and scarce compute, proprietary data, and product distribution on the other.
  • Public-domain and licensed collections could become core AI research infrastructure, while disputed-source datasets face a less defensible role in the ecosystem.

The trend: AI training data is shifting toward governed, reusable public and licensed corpora as developers seek alternatives to legally contested archives.

Discussion

  • @knibbs Kate Knibbs on bluesky
    Exclusive: Another big public domain dataset is coming—and this one has big AI players backing it.  —  www.wired.com/story/harvar...
  • @tedunderwood.me Ted Underwood on bluesky
    So glad to see recognition that AI has increased the urgency of open culture initiatives.  [embedded post]
  • @clementdelangue Clem on x
    can't wait for it to be on HF for the community!