US court filings detail Anthropic's Project Panama, an effort to “destructively scan” up to 2M books with a hydraulic “cutting machine” led by an ex-Google exec
In early 2024, executives at artificial intelligence start-up Anthropic ramped up an ambitious project they sought to keep quiet.
Washington Post
Context & Ripple Effects
Anthropic's book-corpus practices were already under scrutiny through an authors' copyright case that ended in a $1.5 billion proposed resolution over downloaded books. The newly disclosed operational details add texture to that wider dispute: not just what material entered an AI corpus, but how a lab assembled it and documented its provenance.
First-order effects
The filings give authors, publishers and litigants a more concrete record to assess Anthropic's acquisition and digitization practices, while increasing the company's need to explain the controls surrounding Project Panama.
The disclosure may shape how beneficiaries of the authors' copyright class-action settlement evaluate the separation between pirated-source claims and books obtained through physical scanning.
Second-order effects
Publishers and other rightsholders gain a clearer basis to press AI developers for auditable sourcing, retention and compensation terms rather than relying on broad descriptions of training-data practices.
Rival labs face added pressure to document corpus provenance and scanning workflows, since physical-copy acquisition does not by itself settle questions about downstream AI use.
Third-order effects
If court disclosures continue to expose the operational chain behind model training, AI-data governance is likely to shift from abstract policy commitments toward traceable corpus records that can support licensing, litigation and compliance.
The durable divide will be between labs able to demonstrate governed acquisition paths and those relying on opaque or legally contested data sources.
The trend: This is one data point in the move toward governed AI corpora, where training-data provenance becomes a commercial and legal capability rather than a back-office detail.
In their lawsuit, the authors alleged that Meta higher-ups considered paying for books to train their AI models but opted to instead download millions of books free from “torrent” platforms that facilitate online piracy. www.washingtonpost.com/technology/ 2...
Slicing the spines off of millions of books, while downloading pirated versions of millions more: Inside one company's secret plan to ‘destructively scan every book in the world.’ By Aaron Schaffer, @WillOremus and @nitashatiku https://www.washingtonpost.com/ ...
This story led me to conclude that the rule of law is an illusion clung to only by those who lack sufficient lust for power & money. The method of buying used books, ripping their spines apart & scanning every page turned out to be the more legally sound method www.washingtonpos…
New: Unsealed court docs detail Big Tech's yearslong, secret race to ingest the collective works of humanity, including Anthropic's project to “destructively scan all the books in the world.” [image]
Anthropic's project to “destructively scan all the books in the world” was actually viewed as the *more* ethical and legally sound approach to training its AI. — Previously the industry standard had been simply to pirate vast “shadow libraries” of digitized books for free onlin…
Journalists have to eat, so the frame for this article plays up “slicing spines” of books &c. But imo this is actually (like Google Books) an inspiring story of people getting their acts together to organize and transform knowledge. I only wish universities could coordinate as …
“The humans cut apart all the books filled with their knowledge to teach the AI” is absolutely the beginning of a good sci-fi series — www.washingtonpost.com/technology/ 2...
New: Unsealed court docs detail Big Tech's yearslong, secret race to ingest the collective works of humanity, including Anthropic's project to “destructively scan all the books in the world.” — Gift link: wapo.st/4rjXAMQ
machines of spine-slicing grace? ... some of my favorite snippets from newly-released court docs in the Anthropic books copyright case let's start w/ Project Panama, their plan to “destructively scan all the books in the world” to train AI [image]
Centuries of human culture, stolen & fed to a machine so it can regurgitate AI slop for tech companies to force-feed the masses in an attempt rot our brains enough that we can be controlled via a surveillance techno-capitalist state.
@AnnahBackstrom ... Yes, great piece shining light on the unsavory practices of tech companies in training their AI models. One comment—while it's true that both Alsup & Chhabria ruled in favor of fair use, they were two quite different lower court decisions, & Chhabria left the …
Our story today on Anthropic's “Project Panama” — which was an effort to find a more legal/ethical approach to vacuuming up the world's books than the previous industry standard, which was simply to torrent them from online pirates. Gift link: https://www.washingtonpost.com/ ...
Inside a tech company's secretive plan to destroy millions of books | Court filings reveal how AI companies raced to obtain more books to feed chatbots …