A look at LibGen, one of the largest online pirate libraries, with 7.5M+ books and 81M+ research papers, allegedly used by Meta and OpenAI to train AI models
Meta pirated millions of books to train its AI. Search through them here. — When employees at meta started developing …
The Atlantic Alex Reisner
Context & Ripple Effects
The LibGen allegations extend earlier scrutiny of large book corpora: coverage had already identified Books3 as training material for Meta's Llama and other models. The significance here is LibGen's far larger apparent catalog and its alleged connection to both Meta and OpenAI.
The story also sits between reports of internal Meta discussions about LibGen use and the later Kadrey v. Meta case, which centers on LibGen's alleged role in Llama training. It turns a disputed sourcing practice into a concrete provenance and copyright issue.
First-order effects
- Meta and OpenAI face more pointed questions over the provenance of training texts, while LibGen becomes a central source repository in the allegations rather than a peripheral piracy site.
- Authors and publishers gain a clearer factual target for examining whether works from a library of this scale entered model-training pipelines.
Second-order effects
- The alleged sourcing strengthens the relevance of the Kadrey v. Meta legal test, raising the stakes for model developers that relied on similarly assembled text datasets.
- Developers seeking lower-risk training inputs have greater incentive to use documented alternatives, such as Harvard's public-domain book dataset, or to build auditable licensing and provenance processes.
Third-order effects
- If courts and rights holders continue to focus on corpus origins, competitive advantage in foundation models may increasingly depend on governed access to text, not simply the ability to collect it at scale.
- The likely structural shift is toward training-data supply chains that can document rights and source history; the extent and cost of that shift will depend on legal outcomes and workable licensed-data markets.
The trend: AI training is moving from an era of opaque web-scale collection toward contested, governed content supply chains.
Related: Governed AI Corpus · Content as Inference Input · LibGen · Meta · Kadrey v. Meta, centered on the use of LibGen to train Llama AI models · Harvard releases a high-quality dataset of nearly 1M public-domain books
Related Coverage
- Search LibGen, the Pirated-Books Database That Meta Used to Train AI The Atlantic · Alex Reisner
- Teslas are being traded in at a record clip Sherwood News · Rani Molla
- US music industry bodies make case to Trump over AI training Music Ally · Stuart Dredge
- Start Up No.2409: Meta's book piracy revealed, Apple to shuffle Siri execs, TV+ losing big (but who cares), let bird flu rip?, and more The Overspill · Charlesarthur
- Authors fume over claims stolen books used to fuel AI The Canberra Times · Jennifer Dudley-Nicholson
- Meta pirated at least 101 of my books and articles, and tens of millions of others Marcus on AI · Gary Marcus
- I can't believe I have to say this againSearch LibGen, the Pirated-Books Database That Meta Used to Train AI … Lou In Progress · Lou-Wilham
- Pirated-Books Database LibGen, Which Meta Used to Train AI, Includes Books By Artists, Architects, Galleries and Museums ARTnews · Karen K. Ho
- @_alexreisner: Search LibGen, the Pirated-Books Database That Meta Used to Train AI Music Technology Policy · Chris Castle
- Meta Stole BOOKS for their Ai I saw this on Bluesky and dropping the article link here. … Twisted Dreams · Dracota
- Passing this info along from Bluesky at the request of a tumblr mutual who'd like to reblog … south downs founding father … · Irisbleufic
- ... You know that nagging suspicion that AI models have been trained on copyrighted work? … Ellipsus
- Let's move away from ascribing agency to an artifact “AI's problem...” to naming who is pirating these books for profit. — https://www.theatlantic.com/ ... @timnitGebru@dair-community.social · Timnit Gebru
- Copyright is a huge societal negative externality. Access to Knowledge should be universal and not just restricted to companies big enough to escape the law. To book publishers: close Libgen? You will never see my money again, because I will never buy a book again. I will just use the next LibGen. … @remixtures@tldr.nettime.org · Miguel Afonso Caetano
- Found out Meta now has all five of my books - last time a searchable DB was posted, they only had two. — Not a lot of things in publishing that are more demoralizing than having your work wholesale stolen so you can be replaced, and having absolutely nobody care. … @lizmonster@universeodon.com
- Meta illegally used my Scars of the Sundering trilogy (Malediction, Lament, and Salvation) to train its plagiarism machine. — My books are NOT free. Meta and Libgen STOLE from me, costing me money, and affecting my livelihood. — Authors, search your names here: — https://www.theatlantic.com/ ... … @JediSoth@chirp.enworld.org · Hans Cummings
- https://www.theatlantic.com/ ... “Meta employees spoke with multiple companies about licensing books and research papers, but they weren't thrilled with their options.” — To be clear, the option that Meta didn't like, was paying creators for our work. … @jwilker@wandering.shop · John Wilker
- I knew I was extensively pirated & AI was theft *and* books were highly prized because they're premium content for language usage but FUCK ME it still raises my blood pressure to see actual numbers. — ""books are actually more important than web data. … @susankayequinn@wandering.shop
- delighted to read about this massive dataset of stolen books that LLMs hoovered up for training data, and subsequently learn that literally every book i've ever written or contributed to is in it — just delighted — https://www.theatlantic.com/ ... @beep@follow.ethanmarcotte.com · Ethan Marcotte
- GenAI is built upon labor stolen from writers without compensation or consent. There is no “ethical” use. … Melanie Dusseau
- When Google digitized millions of books in 2005, publishers feared the death of their industry. Two decades later, digital discovery drives book sales in ways nobody predicted. … Lorenzo Thione
- If you've published books or journal articles, you can check whether they appear in the Libgen database that Meta used for AI training. … Chris Chapman
- The Atlantic has released a tool that allows anyone to view authors and their publications that were used to train Meta's AI systems without consent [intellectual theft, basically]. … Zaki Khalid
- Google Scholar is your vanity mirror - your instragrammed, airbrushed, filtered academic citation record. — Scopus is what you look like first thing in the morning. … Eamon Costello
- Among literally millions of books, Meta has stolen 35 of my titles - novels, poems, creative nonfiction, a short story - to feed its AI. … Alison Croggon
- Academic and/or writer colleagues who think AI is inherently a good thing might want to enter their names into the search engine linked below … Fiona Moore
Discussion
-
The Daddy Complex
The Daddy Complex
on x
Search LibGen, the Pirated-Books Database That Meta Used to Train AI …
-
@joemenn
Joseph Menn
on bluesky
Really an honor just to be on this list with so many amazing authors. Though that doesn't mean I'm not suing. [embedded post]
-
@amitkatwala
Amit Katwala
on bluesky
Strange time to be a writer. Here's confirmation (via @theatlantic.com) that Meta stole content from all three of my books to train its AI models using a database of pirated books called LibGen — www.theatlantic.com/technology/ a... [image]
-
@heydebigale
Debbie Gale Mitchell
on bluesky
If you want to check if meta stole your papers: — www.theatlantic.com/technology/ a...
-
@mcusolito
Michelle Cusolito
on bluesky
Two of my books were used to train AI without my approval or compensation. I've spent decades honing my craft. And those specific books took years to write. To create one of them, I had to be away from my home and family for 5 1/2 weeks. This is WRONG. www.theatlantic.com/tec…
-
@dougjohnstone
Doug Johnstone
on bluesky
So Meta have stolen all eighteen of my novels (and all my friends' books too) to train their stupid AI, with no permission, copyright or compensation. I assume every other AI is exactly the same. It's fucking theft, plain and simple. — www.theatlantic.com/technology/ a...
-
@mattsteinglass
Matt Steinglass
on bluesky
I looked up which of my articles Meta used to train its AI without paying copyright fees. Appropriately, it seems it read my 2002 article in Transitions about...intellectual property theft (hawkers of pirate CDs in African street markets) www.theatlantic.com/technology/ a... [i…
-
@jdslack
Jennifer Slack
on bluesky
Meta used this database to scrape for AI development. Put your name in the searchbox to see if they might have used your published work. — www.theatlantic.com/technology/ a...
-
@benskipper
@benskipper
on bluesky
8 of my books by @penandswordbooks.bsky.social have been ripped off by Meta. Huge numbers of authors are affected. This is serious copyright theft. @peterstefanovic.bsky.social @mikegalsworthy.bsky.social @barristersecret.bsky.social are you aware of this? — www.theatlantic.c…
-
@www.johnbleasdale.com
John Bleasdale
on bluesky
Bertolt Brecht argued that bank robbery is a minor crime compared to the crime that brought the bank such wealth in the first place. AI is consistent as a late capitalism moment of piracy and theft — www.theatlantic.com/technology/ a...
-
@naomiclifford
Naomi Clifford
on bluesky
Meta considered licensing books to train AI — but opted instead to pirate LibGen, a database of >7.5m books and 81m research papers, says The Atlantic. — Authors - check your works here: www.theatlantic.com/technology/ a... Meta has thieved 3 of my own works. — Class action …
-
@ajwestauthor
@ajwestauthor
on bluesky
I'm horrified to discover Meta used pirated copies of my novels to train its AI. It's an assault and a particularly cruel one to use my work to train the monster that threatens the ruination of original literature. We deserve compensation. — www.theatlantic.com/technology/ a.…
-
@Firlefanz.writing.exchange …
Hannah Steenbock
on bluesky
Nine of my books are listed among those that were illegally used to train AIs. — And I'm basically a self-publishing nobody. — Go check if they scraped yours, as well. — https://www.theatlantic.com/ technology/archive/2025/03/search- libgen-data-set/682094/
-
@paulphillips44
Paul Phillips
on bluesky
I posted this on the other place and thought I'd better share it here. — Not playing #vss365 or anything else today, #WritingCommunity. Feeling pretty violated. Check this link and you may too... theatlantic.com/technology/a... [images]
-
@bethreadscrime.com
@bethreadscrime.com
on bluesky
As if the pirating wasn't awful enough, this is terrible. It is a shame there isn't more DRM, for one of our games if someone had pirated it - all the characters would be poisoned and wearing pirate hats. The support queries about it being broken were fun 😅 — www.theatlantic.…
-
@markriedl
Mark Riedl
on bluesky
Looks like more than 40 of my works are in the dataset that Meta AI pirated www.theatlantic.com/technology/ a...
-
@kristiankiehling
Kristian Kiehling
on bluesky
The pirating of books by Mark Zuckerberg's company META to train its A.I. model Llama 3 should be enough to get him indicted in a European court, as many European authors are affected by this blatant copyright theft. — www.theatlantic.com/technology/ a...
-
@annafeatherstone
Anna Featherstone
on bluesky
Just found out meta stole my memoir about organic farming/native bees via a pirate site to train its AI... a huge number of Australian authors are just discovering massive, unauthorised theft of their books. www.theatlantic.com/technology/ a...
-
@joachim123
Joachim Frank
on bluesky
Acc to an article in The Atlantic, Meta, using LibGen, appears to have pirated millions of articles and book chapters to train AI. The Atlantic provides a searchable link, and I find 199 of my books, book chapter, review article and major publications, all used without my permis…
-
@kint
Jason Kint
on bluesky
This thread held up well considering Atlantic report this morning (www.theatlantic.com/technology/ a...). [embedded post]
-
@garymarcus
Gary Marcus
on bluesky
Jason Kint called it. Looks like Meta management were absolutely mendacious thieves at immense scale. — www.theatlantic.com/technology/ a... [embedded post]
-
@raxkingisdead
Rax ‘Levon Honkers’ King
on bluesky
i don't mind if a human being pirates my book in order to read it. i do mind if mark zuckerberg pirates my book to make meta AI more profitable
-
@laurarbelin
Laura Belin
on bluesky
Thanks to @ulidig.bsky.social for flagging this article for me. Several of my own publications from my “past life” covering Russian politics are included here. I did not consent and would not have consented to their use for this purpose. [embedded post]
-
@boxbrown
Brian Box Brown
on bluesky
META stole your works to train AI. Here's the database. If they stole from me they probably stole from you too — www.theatlantic.com/technology/ a... [image]
-
@dlknowles
Daniel Knowles
on bluesky
Just playing with Meta AI on Whatsapp and it takes like three messages to get it to admit it has actually read my book [image]
-
@authormsbev
@authormsbev
on bluesky
Looks like 40 of my titles were ripped off - even the French and Brazilian translations. I hate these thieving muthafuckas so fucking much!! [embedded post]
-
@moonalice.com
Roger McNamee
on bluesky
It is ironic that Meta pirated my book, Zucked: Waking Up to the Facebook Catastrophe, to train its LLMs. — This also pisses me off. This company, like the rest of Big Tech, does not believe that laws apply to them. [embedded post]
-
@maris
Maris Kreizman
on bluesky
lol my new book isn't out until July but Meta already used it to train its AI [embedded post]
-
@melissagiragrant.com
Melissa Gira Grant
on bluesky
Among them, Meta used my review of Kate Losse's book about Facebook's culture of non-consent... ... ... [embedded post]
-
@katienotopoulos
Katie Notopoulos
on bluesky
Wow, the Atlantic looked at what Meta stole for AI training, and made a search tool where you can see what it sucked up via LibGen (a pirated data set of books). — Let's check one thing real quick.... ok yup — www.theatlantic.com/technology/ a.... [image]
-
@ketanjoshi.co
Ketan Joshi
on bluesky
*some* of my work that Meta and OpenAI used to train their sentence generator software: my actual book + an old (2014) paper I co-authored on wind farm media coverage and health fears — Fuck these damp, twitchy little thieves — www.theatlantic.com/technology/ a... www.theatl…
-
@joebankswriter
Joe Banks
on bluesky
OK, this is quite something. Meta used pirate library LibGen to train its AI. You can search LibGen's dataset here: www.theatlantic.com/technology/ a... I looked, and yes, Hawkwind: Radical Escapism In The Age Of Paranoia (as they slightly mistitle it) is there... @markopilkin…
-
@michaellivingston.com
Michael Livingston
on bluesky
Meta used at least 16 of my books, and numerous articles, to help train the AI it will use to make billions. — Authors, search your name here: — www.theatlantic.com/technology/ a...
-
@ErikJonker@mastodon.social
Erik Jonker
on mastodon
Search LibGen, the Pirated-Books Database That Meta Used to Train AI — Millions of books and scientific papers are captured in the collection's current iteration. — https://www.theatlantic.com/ ... #theatlantic #LibGen #Meta #AI #Copyright #IP
-
@ravenbait@mastodon.scot
Sam Fleming
on mastodon
The Atlantic has posted a tool you can use to see if Meta trained its AI on your work. https://www.theatlantic.com/ ...
-
@michael_w_busch@mastodon.online
Michael Busch
on mastodon
I did not agree to have my research papers fed into the automated plagiarism machines. — QT Damon Beres @damonberes.com — 2025 March 20 — NEW: LibGen contains millions of pirated books and research papers, built over nearly two decades. From court documents, we know that …
-
@harrymccracken@mastodon.social
Harry McCracken
on mastodon
Nothing of mine, apparently, but my sister, father, and grandfather's work is all part of this stew, sad to say. https://www.theatlantic.com/ ...
-
@rr4idic@mastodon.online
Dr. Rachel Reddick
on mastodon
Just learned over on BlueSky that a short story and letter-to-the-editor I wrote got scraped into this, which was used to train Meta's LLM ("AI"). — I do not approve. — https://www.theatlantic.com/ ...
-
@invicticide@mastodon.gamedev.place
Josh Sutphin
on mastodon
Amazing behavior from Meta 😂 — https://www.theatlantic.com/ ... Meanwhile, they're trying to sink Sarah Wynn-Williams' tell-all book “Careless People”. Gee, I can't imagine where she got the title. — Did you know that simply deleting the text “all rights reserved” from some…
-
@thecommongreen@mastodon.scot
@thecommongreen@mastodon.scot
on mastodon
The Atlantic has published a search engine that can look through one of the databases of pirated written works that Meta used to train its AI. — I can see Common Weal work in there. I can see some of my own pre-CW work in there. In fact, I can see some of my work in there tha…
-
@bostonjoan
Boston Joan
on threads
Meta stole hundreds of thousands of books and articles to train its AI, including 7 articles of mine and my MEME WARS book. Grand theft academia. https://www.theatlantic.com/ ... @zuck @mosseri @andymstone Did you all know about this? …
-
@wurdsmyth
Miranda Dickinson
on threads
All of my books stolen to train Meta's AI, in three languages. Sixteen years of work, stolen to make a rich company even richer, while authors struggle to keep going. This is wholesale theft by companies who don't believe they should pay for anything. …
-
@nwbrownboi
@nwbrownboi
on threads
OK I am about to drop a really hot, nuanced take on Meta torrenting LibGen. I recognise this is a sensitive subject and to be clear - I don't approve of the torrenting. But there is a much broader point about LibGen that you - yes, you, a Western reader are not seeing. I ask y…
-
@nwbrownboi
@nwbrownboi
on threads
The problem isn't AI - it's who controls it. AI could be an open-source library, crediting & compensating authors, making knowledge truly accessible instead of locked in corporate models. (Some models are working on citations.) But that requires breaking the cycle of extractio…
-
@nwbrownboi
@nwbrownboi
on threads
LibGen wasn't built for piracy - it was built for access. Created in 2008 (17 years ago, long before AI) by Russian scientists, it served students & researchers in India, Africa, Iran - places where Western paywalls kept knowledge locked away. You shouldn't need a shadow librar…
-
@nwbrownboi
@nwbrownboi
on threads
For decades, Western publishers profited off knowledge hoarding. Now AI is absorbing books, and suddenly the institutions that never cared about access are crying theft. The gatekeepers are losing power, but that doesn't mean the people are winning. You ignored the fight over …
-
@jscalzi
John Scalzi
on threads
I have no doubt Meta's lawyers and accountants figured it would be cheaper to pay any potential fine than it would be to license the works (and yes, my work is in there, across several languages).
-
@karaswisher
Kara Swisher
on threads
1. The greedy information thieves of Meta missed one, my first, AOL.com. — 2. At least their Llama LLM is being trained that its overlord is a greedy information thief. — [image]
-
@bradthor
Brad Thor
on x
The Unbelievable Scale of Meta's Pirated-Books. Meta pirated millions of books to train its AI. 126 versions of my books were stolen and used without my permission. Story here: https://www.theatlantic.com/ ... [image]
-
@jason_kint
Jason Kint
on x
Link in second post here otherwise Twitter will suppress it. Please share first tweet. Much of the report seems to come from lawsuit which posted summary judgment last night arguing torrenting of protected IP alone blows up Meta's BS fair use defense. 2/2 https://www.theatlantic.…
-
@jason_kint
Jason Kint
on x
This Atlantic investigation just hit and my eyes are popping at the alleged lawbreaking by Facebook. “Eventually, the team at Meta got permission from ‘MZ’ — an apparent reference to Meta CEO Mark Zuckerberg—to download and use the data set.” 1/2 [image]
-
@mininghistory
Duncan Money
on x
Meta used 3 of my books and 6 articles to train its AI model, unbeknownst to me. Anyway, Meta are planning to spend $65bn this year on AI development, so I look forward to receiving my modest share of that. https://www.theatlantic.com/ ...
-
@garymarcus
Gary Marcus
on x
Meta pirated at least 101 of my books and scientific articles, and in every single case used them without my permission. Many other authors are discovering the same thing. 🧵1/2
-
@garymarcus
Gary Marcus
on x
Must read on the utter lack of ethics at Meta.
-
@jason_kint
Jason Kint
on x
A counterpoint on why stealing of pirated material and copyright law matters to American, our economy and the future.
-
@neilturkewitz
Neil Turkewitz
on x
During discovery, this message from a Meta employee was produced: “The problem is that people don't realize that if we license one single book, we won't be able to lean into fair use strategy.” Fair use ≠ a business strategy. This is extremely damning—piracy was a choice! 🔥🔥
-
@elamin88
@elamin88
on x
I'd like to think that with all the millions of books going into this, the one tiny thing Meta's AI picked up from my book is: when in doubt, throw an extra em dash on it [image]
-
r/Piracy
r
on reddit
The Unbelievable Scale of AI's Pirated-Books Problem
-
r/books
r
on reddit
The Unbelievable Scale of AI's Pirated-Books Problem
-
r/technology
r
on reddit
The Unbelievable Scale of AI's Pirated-Books Problem. Meta pirated millions of books to train its AI. Search through them here.