/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

← → days · ↑ ↓ browse · Enter similar · o open

Experts say it is not clear whether generative art made by AI systems trained on copyrighted data, like OpenAI's DALL-E 2, can be considered as fair use

but not the most important ones relating to the use of existing © works to train AI. https://www.wired.com/... @maasgad : “Is it right that the AIs of the future are able to produce something magical on the backs of our labor, potentially without our consent or compensation?” https://www.engadget.com/... Andrew Tarantola / @terrortola : Is generative art truly unique if the AI that made it was trained using copyrighted data scraped from the internet? If seeing and doing is good enough for a monkey, why not Dall-E? @danielwcooper investigates https://www.engadget.com/... John Bergmayer / @bergmayer : if there is no “substantial similarity” between the source material and the result, the result is not a “copy” so you don't even need fair use https://www.engadget.com/...

Engadget Daniel Cooper

Context & Ripple Effects

This piece lands at the start of the arc: in mid-2022, experts including John Bergmayer were still debating whether outputs from models trained on scraped copyrighted data — OpenAI's DALL-E 2 among them — qualify as fair use, with some arguing that a lack of substantial similarity means no copying occurred at all.

Within months the abstract question turned concrete: [[a:984092|artists discovered their work in Stability AI's Stable Diffusion training set without consent or payment]], Japan's anime community pushed back against lenient data-scraping norms (Rest of World's reporting), and lawyers interviewed by The Verge flagged unresolved copyright questions as existential for startups. By late 2023, analysts were predicting a wave of copyright suits against OpenAI and peers as models produce infringing material without attribution.

First-order effects

  • Artists whose work sits in scraped training datasets — the Stability AI group among them — face use of their labor with no consent mechanism, no attribution, and no compensation path while the fair-use question stays open.
  • OpenAI and DALL-E 2 users operate under legal ambiguity: if courts later rule the training or outputs infringing, both the vendor and the commercial users of generated art carry exposure.

Second-order effects

  • Rivals building image generators must now treat dataset provenance as a product decision — opt-in licensing, disclosure, or consent flows — because the Stability AI artist backlash showed that silent scraping carries reputational cost even before any court rules.
  • The unresolved doctrine invites litigation as the default resolution path; coverage by year-end 2023 already anticipated more copyright suits against OpenAI and similar firms as attribution-free outputs proliferate.

Third-order effects

  • If courts side against fair use for training on copyrighted data, the industry restructures around licensed corpora and paid rights holders, turning today's free-scraping assumption into a procurement line item; if they side for it, the Copyright Office's refusal to protect AI-made images still leaves generated output itself outside copyright, unsettling who owns what on both ends of the pipeline.
  • Either outcome pushes toward an explicit public-data permission boundary — a norm, or regulation, defining what 'scraped from the internet' legally entitles model builders to take.

The trend: Generative AI is dragging copyright law from a fair-use gray zone toward an enforced permission regime for training data, with litigation and artist pressure forcing the reckoning.

Discussion

  • @bergmayer John Bergmayer on x
    if there is no “substantial similarity” between the source material and the result, the result is not a “copy” so you don't even need fair use https://www.engadget.com/...
  • @neilturkewitz Neil Turkewitz on x
    “By giving users commercial usage rights, OpenAI is sidestepping some of the tricky IP questions raised by this technology.” @JessicaRizzo19 ⁦@WIRED⁩ Some, perhaps—but not the most important ones relating to the use of existing © works to train AI. https://www.wired.com/...
  • @maasgad @maasgad on x
    “Is it right that the AIs of the future are able to produce something magical on the backs of our labor, potentially without our consent or compensation?” https://www.engadget.com/...
  • @terrortola Andrew Tarantola on x
    Is generative art truly unique if the AI that made it was trained using copyrighted data scraped from the internet? If seeing and doing is good enough for a monkey, why not Dall-E? @danielwcooper investigates https://www.engadget.com/...