Experts say it is not clear whether generative art made by AI systems trained on copyrighted data, like OpenAI's DALL-E 2, can be considered as fair use
but not the most important ones relating to the use of existing © works to train AI. https://www.wired.com/... @maasgad : “Is it right that the AIs of the future are able to produce something magical on the backs of our labor, potentially without our consent or compensation?” https://www.engadget.com/... Andrew Tarantola / @terrortola : Is generative art truly unique if the AI that made it was trained using copyrighted data scraped from the internet? If seeing and doing is good enough for a monkey, why not Dall-E? @danielwcooper investigates https://www.engadget.com/... John Bergmayer / @bergmayer : if there is no “substantial similarity” between the source material and the result, the result is not a “copy” so you don't even need fair use https://www.engadget.com/...
Context & Ripple Effects
This piece lands at the start of the arc: in mid-2022, experts including John Bergmayer were still debating whether outputs from models trained on scraped copyrighted data — OpenAI's DALL-E 2 among them — qualify as fair use, with some arguing that a lack of substantial similarity means no copying occurred at all.
Within months the abstract question turned concrete: [[a:984092|artists discovered their work in Stability AI's Stable Diffusion training set without consent or payment]], Japan's anime community pushed back against lenient data-scraping norms (Rest of World's reporting), and lawyers interviewed by The Verge flagged unresolved copyright questions as existential for startups. By late 2023, analysts were predicting a wave of copyright suits against OpenAI and peers as models produce infringing material without attribution.
First-order effects
- Artists whose work sits in scraped training datasets — the Stability AI group among them — face use of their labor with no consent mechanism, no attribution, and no compensation path while the fair-use question stays open.
- OpenAI and DALL-E 2 users operate under legal ambiguity: if courts later rule the training or outputs infringing, both the vendor and the commercial users of generated art carry exposure.
Second-order effects
- Rivals building image generators must now treat dataset provenance as a product decision — opt-in licensing, disclosure, or consent flows — because the Stability AI artist backlash showed that silent scraping carries reputational cost even before any court rules.
- The unresolved doctrine invites litigation as the default resolution path; coverage by year-end 2023 already anticipated more copyright suits against OpenAI and similar firms as attribution-free outputs proliferate.
Third-order effects
- If courts side against fair use for training on copyrighted data, the industry restructures around licensed corpora and paid rights holders, turning today's free-scraping assumption into a procurement line item; if they side for it, the Copyright Office's refusal to protect AI-made images still leaves generated output itself outside copyright, unsettling who owns what on both ends of the pipeline.
- Either outcome pushes toward an explicit public-data permission boundary — a norm, or regulation, defining what 'scraped from the internet' legally entitles model builders to take.
The trend: Generative AI is dragging copyright law from a fair-use gray zone toward an enforced permission regime for training data, with litigation and artist pressure forcing the reckoning.