/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

Q&A with Katherine Forrest, a former federal judge for the SDNY, on copyright and generative AI, the Copyright Office's guidance on AI-generated work, and more

A conversation with Katherine Forrest Before they gobbled up headlines everywhere, large language models ingested truly staggering amounts of data to train their models. Mastodon: @jon@henshaw.social and @alex@dair-community.social Tweets: @stanfordnlp , @alexhanna , @marika_writes_ , @mehtabkn , and @mausellmoli Mastodon: @jon@henshaw.social : “...Google Books was to allow you to search for the book, and it gave you a portion of the copy for you to access.  The point was discovering the underlying work. … Alex Hanna / @alex@dair-community.social : This @themarkup interview of Katherine Forest is informative.  Though I disagree with the term “foundation models” as a marketing term by Stanford, Forest is spot on that the fair use test will mostly hinge on the market implications of AI-generated art and music. … Tweets: @stanfordnlp : Katherine Forrest @PaulWeissLLP: “I would like to move away from “large language model” because that causes people to get stuck into a language-only space. These models are really foundation models that are not just language-based but also photographic-, video-, audio-based....” https://twitter.com/... @alexhanna : This @themarkup interview of Katherine Forest is informative. She is spot on that the fair use test will mostly hinge on the market implications of AI-generated art and music. https://themarkup.org/... >> @marika_writes_ : “For generative AI, the value is not necessarily the discovery of the underlying creative work but rather a replacement of the work with something else. It is 100 percent copy leading to 100 percent replacement. And it's absolutely commercial.” https://twitter.com/... Mehtab Khan / @mehtabkn : There is a tendency to lump together the “input” stage into one big step but it's not. There are distinct stages under the “input” umbrella, each raising specific copyright/fair use questions. @alexhanna and I break down these steps in our paper https://twitter.com/... Mauricio Sellmann Oliveira / @mausellmoli : Katherine Forrest, former federal judge for the Southern District of New York, on large language models, I mean, foundation models and copyrights: https://themarkup.org/... https://twitter.com/...

The Markup Nabiha Syed

Context & Ripple Effects

This Q&A sits early in the arc that later produced The New York Times' lawsuit against OpenAI and Microsoft: months before those suits landed, Katherine Forrest — a former SDNY federal judge — was already laying out how fair use might apply to models trained on ingested works, invoking the Google Books precedent of discovery versus copying. The interview also lands alongside the Copyright Office's guidance on whether AI-generated work qualifies for copyright at all.

What followed validates why her framing mattered: legal experts now read earlier fair use cases as offering mixed lessons for the Times case, and OpenAI has staked its fair-use-plus-opt-out defense on exactly the questions Forrest previews — where training ends and infringing output begins.

First-order effects

  • For plaintiffs like The New York Times and the artists behind Stable Diffusion complaints, Forrest's judge's-eye view sharpens the central question: does ingesting works to build a capability count as transformative use, or as preparing a commercial substitute for them?

Second-order effects

Third-order effects

  • Because commentators treat each stage of the pipeline separately — data ingestion, training, generation, output — an eventual judicial resolution will likely define per-stage rules rather than one blanket verdict, structuring how every future foundation model sources its corpus.

The trend: Copyright law for generative AI is being settled court by court, with the outcome either redrawing fair use around training corpora or forcing a standing licensing market between rights holders and model builders.

Discussion

  • @stanfordnlp @stanfordnlp on x
    Katherine Forrest @PaulWeissLLP: “I would like to move away from “large language model” because that causes people to get stuck into a language-only space. These models are really foundation models that are not just language-based but also photographic-, video-, audio-based....” …
  • @alexhanna @alexhanna on x
    This @themarkup interview of Katherine Forest is informative. She is spot on that the fair use test will mostly hinge on the market implications of AI-generated art and music. https://themarkup.org/... >>
  • @marika_writes_ @marika_writes_ on x
    “For generative AI, the value is not necessarily the discovery of the underlying creative work but rather a replacement of the work with something else. It is 100 percent copy leading to 100 percent replacement. And it's absolutely commercial.” https://twitter.com/...
  • @mehtabkn Mehtab Khan on x
    There is a tendency to lump together the “input” stage into one big step but it's not. There are distinct stages under the “input” umbrella, each raising specific copyright/fair use questions. @alexhanna and I break down these steps in our paper https://twitter.com/...
  • @mausellmoli Mauricio Sellmann Oliveira on x
    Katherine Forrest, former federal judge for the Southern District of New York, on large language models, I mean, foundation models and copyrights: https://themarkup.org/... https://twitter.com/...