Prosecraft, a site that compiled 27K+ books to compare and rank the “vividness” of their language, shuts down after writers' backlash over its possible AI uses
Today, I'm taking down the prosecraft.io website … Sam Barsanti / The A.V. Club : Explaining the Prosecraft drama: This is why every author is mad at a website that counted words Robert Andrew / CityLife : Prosecraft.io Shuts Down After Backlash from Authors Linda Codega / Gizmodo : Fiction Analytics Site Prosecraft Shut Down After Author Backlash Mastodon: @imtheq@realsocial.life : Maddening. They won't be happy until every bit of art, literature and music is devalued to the point that human creatives are incapable of earning a decent living from their work. — We need a popular movement committed to supporting independent creators outside these monopolistic platforms. … X: Celeste Ng / @pronounced_ing : I count 20 @StephenKing books, 20 @jodipicoult books, plus books by tons of other contemporary writers (@legroff, Madeline Miller, Angie Thomas... just search and they're pretty much in there). Did any of these authors/publishers consent to this use of their work? Maureen Johnson / @maureenjohnson : I join the chain of those who did not consent to this. [image] Benji Smith / @benji_smith : I took down the prosecraft website. I'm sorry, everyone. https://blog.shaxpir.com/... @kaylaancrum : Watching weird tech people get seduced by mega corp boldness into believing the lie that all art is “content” and like “content” produced by novices, is unprotected and available for use has been very 😃 This guy is about to get sued by all ten of the largest publishers at once. @zachrosewriter : How DARE you, @benji_smith I demand you take my book off your site immediately. I do not consent to this, and never did. And I know my publisher never would [image] @lindaholmes : @pronounced_ing ... Mine's in there, and I definitely didn't. Katherine Locke / @bibliogato : He took the website down but that doesn't account for the theft and the database he fed our books into. Publishers and authors need to be aggressive and not let this go. Guys like this will just let the fury die down and quietly restart. Celeste Ng / @pronounced_ing : Speaking of which, this might be of interest to writers out there (and anyone else concerned about AI being trained on writing by humans, without consent): https://authorsguild.org/... Celeste Ng / @pronounced_ing : I'm not worried about AI writing being *better* than works by actual humans. But I *am* worried about companies (wrongly) thinking they can replace human writers with AI. And companies should DEFINITELY not be able to train AI on other people's work w/o consent or compensation. Hari Kunzru / @harikunzru : This company Prosecraft appears to have stolen a lot of books, trained an AI, and are now offering a service based on that data https://blog.shaxpir.com/... Celeste Ng / @pronounced_ing : Adding insult to injury, the site doing the “analysis” doesn't seem to understand what passive voice is, and arbitrarily assigns each word a “vividness” score, and then scores the book's “vividness” on that.
Context & Ripple Effects
Prosecraft turned published books into a comparative language dataset, placing ordinary literary analysis inside a widening dispute over whether creative work can be repurposed as AI input. The concern landed as AI-generated text was beginning to compete for writing work, making the provenance and downstream use of book data newly consequential.
The shutdown is an early example of authors using public pressure to challenge unconsented corpus-building. Later coverage of efforts to remove the Books3 dataset and calls for collective lobbying by writers and publishers shows how such objections broadened from a single site to control and compensation questions.
First-order effects
- Prosecraft loses its public-facing product, and readers lose access to its rankings and comparisons across its 27,000-plus-book corpus.
- Authors’ backlash makes potential AI use—not just the site’s stated analytics function—a central reputational risk for services that collect books without clear creator consent.
Second-order effects
- Other literary-data and AI-adjacent services face pressure to explain what they collect, how it is retained, and whether it can be used for model development.
- Publishers and author groups gain a concrete organizing case for seeking permissions, compensation, or restrictions around book datasets; the debate later extended to AI-assisted author services such as Spines’ editing and distribution offering.
Third-order effects
- If creators continue to contest unconsented training and analysis, book corpora may shift from openly assembled inputs toward licensed, controlled, or legally contested assets.
- The episode points to a broader split between tools that extract value from cultural catalogs and creators seeking a say in those downstream uses; the eventual balance will depend on collective bargaining, platform policy, and legal outcomes.
The trend: Creative works are increasingly being treated as inference inputs, prompting authors and publishers to demand control, transparency, and a share of AI-derived value.