/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

A look at LibGen, one of the largest online pirate libraries, with 7.5M+ books and 81M+ research papers, allegedly used by Meta and OpenAI to train AI models

Meta pirated millions of books to train its AI.  Search through them here.  —  When employees at meta started developing …

The Atlantic Alex Reisner

Context & Ripple Effects

The LibGen allegations extend earlier scrutiny of large book corpora: coverage had already identified Books3 as training material for Meta's Llama and other models. The significance here is LibGen's far larger apparent catalog and its alleged connection to both Meta and OpenAI.

The story also sits between reports of internal Meta discussions about LibGen use and the later Kadrey v. Meta case, which centers on LibGen's alleged role in Llama training. It turns a disputed sourcing practice into a concrete provenance and copyright issue.

First-order effects

  • Meta and OpenAI face more pointed questions over the provenance of training texts, while LibGen becomes a central source repository in the allegations rather than a peripheral piracy site.
  • Authors and publishers gain a clearer factual target for examining whether works from a library of this scale entered model-training pipelines.

Second-order effects

  • The alleged sourcing strengthens the relevance of the Kadrey v. Meta legal test, raising the stakes for model developers that relied on similarly assembled text datasets.
  • Developers seeking lower-risk training inputs have greater incentive to use documented alternatives, such as Harvard's public-domain book dataset, or to build auditable licensing and provenance processes.

Third-order effects

  • If courts and rights holders continue to focus on corpus origins, competitive advantage in foundation models may increasingly depend on governed access to text, not simply the ability to collect it at scale.
  • The likely structural shift is toward training-data supply chains that can document rights and source history; the extent and cost of that shift will depend on legal outcomes and workable licensed-data markets.

The trend: AI training is moving from an era of opaque web-scale collection toward contested, governed content supply chains.

Discussion

  • The Daddy Complex The Daddy Complex on x
    Search LibGen, the Pirated-Books Database That Meta Used to Train AI …
  • @joemenn Joseph Menn on bluesky
    Really an honor just to be on this list with so many amazing authors.  Though that doesn't mean I'm not suing.  [embedded post]
  • @amitkatwala Amit Katwala on bluesky
    Strange time to be a writer.  Here's confirmation (via @theatlantic.com) that Meta stole content from all three of my books to train its AI models using a database of pirated books called LibGen  —  www.theatlantic.com/technology/ a...  [image]
  • @heydebigale Debbie Gale Mitchell on bluesky
    If you want to check if meta stole your papers:  —  www.theatlantic.com/technology/ a...
  • @mcusolito Michelle Cusolito on bluesky
    Two of my books were used to train AI without my approval or compensation.  I've spent decades honing my craft.  And those specific books took years to write.  To create one of them, I had to be away from my home and family for 5 1/2 weeks.  This is WRONG. www.theatlantic.com/tec…
  • @dougjohnstone Doug Johnstone on bluesky
    So Meta have stolen all eighteen of my novels (and all my friends' books too) to train their stupid AI, with no permission, copyright or compensation.  I assume every other AI is exactly the same.  It's fucking theft, plain and simple.  —  www.theatlantic.com/technology/ a...
  • @mattsteinglass Matt Steinglass on bluesky
    I looked up which of my articles Meta used to train its AI without paying copyright fees.  Appropriately, it seems it read my 2002 article in Transitions about...intellectual property theft (hawkers of pirate CDs in African street markets) www.theatlantic.com/technology/ a...  [i…
  • @jdslack Jennifer Slack on bluesky
    Meta used this database to scrape for AI development.  Put your name in the searchbox to see if they might have used your published work.  —  www.theatlantic.com/technology/ a...
  • @benskipper @benskipper on bluesky
    8 of my books by @penandswordbooks.bsky.social have been ripped off by Meta.  Huge numbers of authors are affected.  This is serious copyright theft. @peterstefanovic.bsky.social @mikegalsworthy.bsky.social @barristersecret.bsky.social are you aware of this?  —  www.theatlantic.c…
  • @www.johnbleasdale.com John Bleasdale on bluesky
    Bertolt Brecht argued that bank robbery is a minor crime compared to the crime that brought the bank such wealth in the first place.  AI is consistent as a late capitalism moment of piracy and theft  —  www.theatlantic.com/technology/ a...
  • @naomiclifford Naomi Clifford on bluesky
    Meta considered licensing books to train AI — but opted instead to pirate LibGen, a database of >7.5m books and 81m research papers, says The Atlantic.  —  Authors - check your works here: www.theatlantic.com/technology/ a...  Meta has thieved 3 of my own works.  —  Class action …
  • @ajwestauthor @ajwestauthor on bluesky
    I'm horrified to discover Meta used pirated copies of my novels to train its AI.  It's an assault and a particularly cruel one to use my work to train the monster that threatens the ruination of original literature.  We deserve compensation.  —  www.theatlantic.com/technology/ a.…
  • @Firlefanz.writing.exchange … Hannah Steenbock on bluesky
    Nine of my books are listed among those that were illegally used to train AIs.  —  And I'm basically a self-publishing nobody.  —  Go check if they scraped yours, as well.  —  https://www.theatlantic.com/ technology/archive/2025/03/search- libgen-data-set/682094/
  • @paulphillips44 Paul Phillips on bluesky
    I posted this on the other place and thought I'd better share it here.  —  Not playing #vss365 or anything else today, #WritingCommunity.  Feeling pretty violated.  Check this link and you may too... theatlantic.com/technology/a...  [images]
  • @bethreadscrime.com @bethreadscrime.com on bluesky
    As if the pirating wasn't awful enough, this is terrible.  It is a shame there isn't more DRM, for one of our games if someone had pirated it - all the characters would be poisoned and wearing pirate hats.  The support queries about it being broken were fun 😅  —  www.theatlantic.…
  • @markriedl Mark Riedl on bluesky
    Looks like more than 40 of my works are in the dataset that Meta AI pirated www.theatlantic.com/technology/ a...
  • @kristiankiehling Kristian Kiehling on bluesky
    The pirating of books by Mark Zuckerberg's company META to train its A.I. model Llama 3 should be enough to get him indicted in a European court, as many European authors are affected by this blatant copyright theft.  —  www.theatlantic.com/technology/ a...
  • @annafeatherstone Anna Featherstone on bluesky
    Just found out meta stole my memoir about organic farming/native bees via a pirate site to train its AI... a huge number of Australian authors are just discovering massive, unauthorised theft of their books. www.theatlantic.com/technology/ a...
  • @joachim123 Joachim Frank on bluesky
    Acc to an article in The Atlantic, Meta, using LibGen, appears to have pirated millions of articles and book chapters to train AI.  The Atlantic provides a searchable link, and I find 199 of my books, book chapter, review article and major publications, all used without my permis…
  • @kint Jason Kint on bluesky
    This thread held up well considering Atlantic report this morning (www.theatlantic.com/technology/ a...).  [embedded post]
  • @garymarcus Gary Marcus on bluesky
    Jason Kint called it.  Looks like Meta management were absolutely mendacious thieves at immense scale.  —  www.theatlantic.com/technology/ a...  [embedded post]
  • @raxkingisdead Rax ‘Levon Honkers’ King on bluesky
    i don't mind if a human being pirates my book in order to read it.  i do mind if mark zuckerberg pirates my book to make meta AI more profitable
  • @laurarbelin Laura Belin on bluesky
    Thanks to @ulidig.bsky.social for flagging this article for me.  Several of my own publications from my “past life” covering Russian politics are included here.  I did not consent and would not have consented to their use for this purpose.  [embedded post]
  • @boxbrown Brian Box Brown on bluesky
    META stole your works to train AI.  Here's the database.  If they stole from me they probably stole from you too  —  www.theatlantic.com/technology/ a...  [image]
  • @dlknowles Daniel Knowles on bluesky
    Just playing with Meta AI on Whatsapp and it takes like three messages to get it to admit it has actually read my book [image]
  • @authormsbev @authormsbev on bluesky
    Looks like 40 of my titles were ripped off - even the French and Brazilian translations.  I hate these thieving muthafuckas so fucking much!! [embedded post]
  • @moonalice.com Roger McNamee on bluesky
    It is ironic that Meta pirated my book, Zucked: Waking Up to the Facebook Catastrophe, to train its LLMs.  —  This also pisses me off.  This company, like the rest of Big Tech, does not believe that laws apply to them.  [embedded post]
  • @maris Maris Kreizman on bluesky
    lol my new book isn't out until July but Meta already used it to train its AI [embedded post]
  • @melissagiragrant.com Melissa Gira Grant on bluesky
    Among them, Meta used my review of Kate Losse's book about Facebook's culture of non-consent... ... ...  [embedded post]
  • @katienotopoulos Katie Notopoulos on bluesky
    Wow, the Atlantic looked at what Meta stole for AI training, and made a search tool where you can see what it sucked up via LibGen (a pirated data set of books).  —  Let's check one thing real quick.... ok yup  —  www.theatlantic.com/technology/ a....  [image]
  • @ketanjoshi.co Ketan Joshi on bluesky
    *some* of my work that Meta and OpenAI used to train their sentence generator software: my actual book + an old (2014) paper I co-authored on wind farm media coverage and health fears  —  Fuck these damp, twitchy little thieves  —  www.theatlantic.com/technology/ a...  www.theatl…
  • @joebankswriter Joe Banks on bluesky
    OK, this is quite something.  Meta used pirate library LibGen to train its AI.  You can search LibGen's dataset here: www.theatlantic.com/technology/ a...  I looked, and yes, Hawkwind: Radical Escapism In The Age Of Paranoia (as they slightly mistitle it) is there... @markopilkin…
  • @michaellivingston.com Michael Livingston on bluesky
    Meta used at least 16 of my books, and numerous articles, to help train the AI it will use to make billions.  —  Authors, search your name here:  —  www.theatlantic.com/technology/ a...
  • @ErikJonker@mastodon.social Erik Jonker on mastodon
    Search LibGen, the Pirated-Books Database That Meta Used to Train AI  —  Millions of books and scientific papers are captured in the collection's current iteration.  —  https://www.theatlantic.com/ ...  #theatlantic #LibGen #Meta #AI #Copyright #IP
  • @ravenbait@mastodon.scot Sam Fleming on mastodon
    The Atlantic has posted a tool you can use to see if Meta trained its AI on your work. https://www.theatlantic.com/ ...
  • @michael_w_busch@mastodon.online Michael Busch on mastodon
    I did not agree to have my research papers fed into the automated plagiarism machines.  —  QT Damon Beres @damonberes.com‬  —  2025 March 20  —  NEW: LibGen contains millions of pirated books and research papers, built over nearly two decades.  From court documents, we know that …
  • @harrymccracken@mastodon.social Harry McCracken on mastodon
    Nothing of mine, apparently, but my sister, father, and grandfather's work is all part of this stew, sad to say. https://www.theatlantic.com/ ...
  • @rr4idic@mastodon.online Dr. Rachel Reddick on mastodon
    Just learned over on BlueSky that a short story and letter-to-the-editor I wrote got scraped into this, which was used to train Meta's LLM ("AI").  —  I do not approve.  —  https://www.theatlantic.com/ ...
  • @invicticide@mastodon.gamedev.place Josh Sutphin on mastodon
    Amazing behavior from Meta 😂  —  https://www.theatlantic.com/ ...  Meanwhile, they're trying to sink Sarah Wynn-Williams' tell-all book “Careless People”.  Gee, I can't imagine where she got the title.  —  Did you know that simply deleting the text “all rights reserved” from some…
  • @thecommongreen@mastodon.scot @thecommongreen@mastodon.scot on mastodon
    The Atlantic has published a search engine that can look through one of the databases of pirated written works that Meta used to train its AI.  —  I can see Common Weal work in there.  I can see some of my own pre-CW work in there.  In fact, I can see some of my work in there tha…
  • @bostonjoan Boston Joan on threads
    Meta stole hundreds of thousands of books and articles to train its AI, including 7 articles of mine and my MEME WARS book.  Grand theft academia. https://www.theatlantic.com/ ... @zuck @mosseri @andymstone Did you all know about this? …
  • @wurdsmyth Miranda Dickinson on threads
    All of my books stolen to train Meta's AI, in three languages.  Sixteen years of work, stolen to make a rich company even richer, while authors struggle to keep going.  This is wholesale theft by companies who don't believe they should pay for anything. …
  • @nwbrownboi @nwbrownboi on threads
    OK I am about to drop a really hot, nuanced take on Meta torrenting LibGen.  I recognise this is a sensitive subject and to be clear - I don't approve of the torrenting.  But there is a much broader point about LibGen that you - yes, you, a Western reader are not seeing.  I ask y…
  • @nwbrownboi @nwbrownboi on threads
    The problem isn't AI - it's who controls it.  AI could be an open-source library, crediting & compensating authors, making knowledge truly accessible instead of locked in corporate models.  (Some models are working on citations.)  But that requires breaking the cycle of extractio…
  • @nwbrownboi @nwbrownboi on threads
    LibGen wasn't built for piracy - it was built for access.  Created in 2008 (17 years ago, long before AI) by Russian scientists, it served students & researchers in India, Africa, Iran - places where Western paywalls kept knowledge locked away.  You shouldn't need a shadow librar…
  • @nwbrownboi @nwbrownboi on threads
    For decades, Western publishers profited off knowledge hoarding.  Now AI is absorbing books, and suddenly the institutions that never cared about access are crying theft.  The gatekeepers are losing power, but that doesn't mean the people are winning.  You ignored the fight over …
  • @jscalzi John Scalzi on threads
    I have no doubt Meta's lawyers and accountants figured it would be cheaper to pay any potential fine than it would be to license the works (and yes, my work is in there, across several languages).
  • @karaswisher Kara Swisher on threads
    1. The greedy information thieves of Meta missed one, my first, AOL.com.  —  2. At least their Llama LLM is being trained that its overlord is a greedy information thief.  —  [image]
  • @bradthor Brad Thor on x
    The Unbelievable Scale of Meta's Pirated-Books. Meta pirated millions of books to train its AI. 126 versions of my books were stolen and used without my permission. Story here: https://www.theatlantic.com/ ... [image]
  • @jason_kint Jason Kint on x
    Link in second post here otherwise Twitter will suppress it. Please share first tweet. Much of the report seems to come from lawsuit which posted summary judgment last night arguing torrenting of protected IP alone blows up Meta's BS fair use defense. 2/2 https://www.theatlantic.…
  • @jason_kint Jason Kint on x
    This Atlantic investigation just hit and my eyes are popping at the alleged lawbreaking by Facebook. “Eventually, the team at Meta got permission from ‘MZ’ — an apparent reference to Meta CEO Mark Zuckerberg—to download and use the data set.” 1/2 [image]
  • @mininghistory Duncan Money on x
    Meta used 3 of my books and 6 articles to train its AI model, unbeknownst to me. Anyway, Meta are planning to spend $65bn this year on AI development, so I look forward to receiving my modest share of that. https://www.theatlantic.com/ ...
  • @garymarcus Gary Marcus on x
    Meta pirated at least 101 of my books and scientific articles, and in every single case used them without my permission. Many other authors are discovering the same thing. 🧵1/2
  • @garymarcus Gary Marcus on x
    Must read on the utter lack of ethics at Meta.
  • @jason_kint Jason Kint on x
    A counterpoint on why stealing of pirated material and copyright law matters to American, our economy and the future.
  • @neilturkewitz Neil Turkewitz on x
    During discovery, this message from a Meta employee was produced: “The problem is that people don't realize that if we license one single book, we won't be able to lean into fair use strategy.” Fair use ≠ a business strategy. This is extremely damning—piracy was a choice! 🔥🔥
  • @elamin88 @elamin88 on x
    I'd like to think that with all the millions of books going into this, the one tiny thing Meta's AI picked up from my book is: when in doubt, throw an extra em dash on it [image]
  • r/Piracy r on reddit
    The Unbelievable Scale of AI's Pirated-Books Problem
  • r/books r on reddit
    The Unbelievable Scale of AI's Pirated-Books Problem
  • r/technology r on reddit
    The Unbelievable Scale of AI's Pirated-Books Problem.  Meta pirated millions of books to train its AI.  Search through them here.