Microsoft, GitHub, and OpenAI told a San Francisco federal court to dismiss a proposed class action lawsuit against Copilot, citing fair use of open-source code
Blake Brittain / Reuters :
Context & Ripple Effects
The Copilot class action has been building since November, when the filing lawyers laid out their theory in a Q&A on challenging the legality of GitHub Copilot: that training on public open-source code without attribution breaks license terms at scale. With today's motion, Microsoft, GitHub, and OpenAI go on offense, asking the San Francisco federal court to kill the suit before it starts by framing code ingestion as fair use.
It matters because this is the earliest live test of the defense every AI developer will reuse: the same fair-use argument now anchors OpenAI's push to dismiss parts of the NYT's copyright suit, and the related coverage shows judges already pruning these cases — a US judge later dismissed most claims in a parallel LLM-training suit by Raw Story and AlterNet for lack of shown harm.
First-order effects
- Microsoft, GitHub, and OpenAI avoid class-wide discovery into Copilot's training and outputs while forcing the developers to defend each of their claims against a fair-use standard rather than a license-violation theory.
- The proposing developers must now show concrete harm from individual code reproduction — the same evidentiary bar that sank the Raw Story and AlterNet copyright claims against OpenAI.
Second-order effects
- A dismissal on fair-use grounds becomes a citable template for OpenAI's other defenses, including its bid to dismiss parts of the NYT suit, letting one ruling shape the whole copyright-litigation front.
- Open-source licensors and code-hosting platforms face pressure to attach explicit training restrictions to licenses, since permissive licenses are exactly what the defendants argue makes ingestion lawful.
Third-order effects
- If courts accept fair use for code-trained assistants — the trajectory visible where a judge later trimmed the Copilot suit from 22 claims to just two surviving claims — training-data disputes shift from infringement claims toward licensing markets and attribution norms as the only lever rights holders retain.
- The pattern points toward a bifurcated ecosystem: model builders trained on permissively licensed corpora treated as low-risk, while restrictive-license content demands paid deals or stays out of training sets.
The trend: Generative-AI copyright fights are converging on fair use as the decisive doctrine, with early rulings on code and news training data setting the template for every model builder's legal exposure.