Q&A with Katherine Forrest, a former federal judge for the SDNY, on copyright and generative AI, the Copyright Office's guidance on AI-generated work, and more
A conversation with Katherine Forrest Before they gobbled up headlines everywhere, large language models ingested truly staggering amounts of data to train their models. Mastodon: @jon@henshaw.social and @alex@dair-community.social Tweets: @stanfordnlp , @alexhanna , @marika_writes_ , @mehtabkn , and @mausellmoli Mastodon: @jon@henshaw.social : “...Google Books was to allow you to search for the book, and it gave you a portion of the copy for you to access. The point was discovering the underlying work. … Alex Hanna / @alex@dair-community.social : This @themarkup interview of Katherine Forest is informative. Though I disagree with the term “foundation models” as a marketing term by Stanford, Forest is spot on that the fair use test will mostly hinge on the market implications of AI-generated art and music. … Tweets: @stanfordnlp : Katherine Forrest @PaulWeissLLP: “I would like to move away from “large language model” because that causes people to get stuck into a language-only space. These models are really foundation models that are not just language-based but also photographic-, video-, audio-based....” https://twitter.com/... @alexhanna : This @themarkup interview of Katherine Forest is informative. She is spot on that the fair use test will mostly hinge on the market implications of AI-generated art and music. https://themarkup.org/... >> @marika_writes_ : “For generative AI, the value is not necessarily the discovery of the underlying creative work but rather a replacement of the work with something else. It is 100 percent copy leading to 100 percent replacement. And it's absolutely commercial.” https://twitter.com/... Mehtab Khan / @mehtabkn : There is a tendency to lump together the “input” stage into one big step but it's not. There are distinct stages under the “input” umbrella, each raising specific copyright/fair use questions. @alexhanna and I break down these steps in our paper https://twitter.com/... Mauricio Sellmann Oliveira / @mausellmoli : Katherine Forrest, former federal judge for the Southern District of New York, on large language models, I mean, foundation models and copyrights: https://themarkup.org/... https://twitter.com/...
Context & Ripple Effects
This Q&A sits early in the arc that later produced The New York Times' lawsuit against OpenAI and Microsoft: months before those suits landed, Katherine Forrest — a former SDNY federal judge — was already laying out how fair use might apply to models trained on ingested works, invoking the Google Books precedent of discovery versus copying. The interview also lands alongside the Copyright Office's guidance on whether AI-generated work qualifies for copyright at all.
What followed validates why her framing mattered: legal experts now read earlier fair use cases as offering mixed lessons for the Times case, and OpenAI has staked its fair-use-plus-opt-out defense on exactly the questions Forrest previews — where training ends and infringing output begins.
First-order effects
- For plaintiffs like The New York Times and the artists behind Stable Diffusion complaints, Forrest's judge's-eye view sharpens the central question: does ingesting works to build a capability count as transformative use, or as preparing a commercial substitute for them?
Second-order effects
- Model developers facing these suits have two live exits — win the fair use argument, or convert the legal pressure into paid licensing deals, which is precisely how analysts frame the lawsuits' dual function as courtroom fight or negotiation leverage.
Third-order effects
- Because commentators treat each stage of the pipeline separately — data ingestion, training, generation, output — an eventual judicial resolution will likely define per-stage rules rather than one blanket verdict, structuring how every future foundation model sources its corpus.
The trend: Copyright law for generative AI is being settled court by court, with the outcome either redrawing fair use around training corpora or forcing a standing licensing market between rights holders and model builders.