/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

Adobe used images created by tools like Midjourney and uploaded to its stock marketplace by users, to train Firefly; Adobe says ~5% of images were AI-generated

Bloomberg

Context & Ripple Effects

Adobe had already chosen to accept and label generative-AI submissions on its stock marketplace, unlike Getty Images at the time, through its policy for labeled AI-made stock images. That marketplace decision created a pathway by which synthetic imagery could enter a library later used in model development.

The disclosure adds a concrete provenance question to Adobe’s cautious effort to embed generative AI in Creative Cloud amid uncertainty over creative professionals’ acceptance of the tools, as described in its early Creative Cloud AI rollout.

First-order effects

  • Adobe must account for the presence of user-uploaded, AI-generated Adobe Stock images in Firefly’s training set; it says those images represented about 5% of the corpus.
  • Contributors who supplied Stock images, including AI-made work, become directly relevant to how Adobe defines the boundary between marketplace licensing, labeling, and model-training use.

Second-order effects

  • Adobe Stock’s labeling and submission rules become more consequential: provenance signals on marketplace content can affect not only customer disclosure but the composition of downstream AI training data.
  • Rival creative-software and stock platforms face greater pressure to explain whether and how customer- or contributor-supplied assets can be reused for model development, particularly when those assets are themselves synthetic.

Third-order effects

  • If marketplaces increasingly double as training-data reservoirs, content provenance will become a core product and governance capability rather than a disclosure feature alone.
  • The episode points toward a more complex commercial chain in which AI-generated content can be sold, licensed, and then reused to improve new generation models; durable trust will depend on clear permissions and traceability.

The trend: Creative platforms are turning their content marketplaces into AI inputs, making provenance and reuse terms central to workflow-native AI commercialization.

Discussion

  • @thebrianpenny Brian Penny on threads
    Adobe says its Firefly dataset is ethical, but it's really not.  About 14% of its training dataset is AI outputs mostly from Midjourney.  And by 2030, the largely unethical AI in Adobe Stock will eclipse human contributions.
  • @louisanslow Louis Anslow on x
    Expect more of this. Big incumbent corporations using ‘ethics’ as a way to tacitly demonize smaller players.
  • @garymarcus Gary Marcus on x
    This is really sad. Have to scratch Adobe off the really short list of big tech companies that I thought were ethically sourcing their training data.
  • @abebab Abeba Birhane on x
    so they lied
  • @brodyford_ Brody Ford on x
    Scoop: Adobe's been keeping a little secret about its AI. Its Firefly model was branded as more ethical than rivals like Midjourney. But it was actually trained on images from them. With @rachelmetz https://www.bloomberg.com/...
  • @woke8yearold Aleph on x
    1. subscribe to Midjourney 2. sell images you made to Boomers on Adobe Stock 3. profit this guy knows what's up [image]
  • @mattcameronlane Matthew Lane on x
    Adobe has already hit copyright escape velocity. That's the theory that models can be trained on uncopyrightable images/text AI output. There never was a sustainable market for artists to make money licensing training data. https://finance.yahoo.com/...
  • @javilopen Javi Lopez on x
    Holy shit 🤯 [image]
  • @emmanuel_2m Emm on x
    Link - https://www.bloomberg.com/... And a few extracts. [image]
  • @rachelmetz Rachel Metz on x
    🚨Scoop from me and @BrodyFord_: Adobe has touted Firefly as a safe and ethical alternative to competitors like Midjourney... but it also quietly trained Firefly on Midjourney images. Adobe's ‘Ethical’ Firefly AI Was Trained on Midjourney Images https://www.bloomberg.com/...
  • @rachelmetz Rachel Metz on x
    @mobileraj @BrodyFord_ This is not about data contamination— it's about the company making a deliberate choice to include AI-generated images from Adobe Stock in its dataset while publicly positioning itself as very different from the companies whose tools were used to make those…
  • r/graphic_design r on reddit
    Turns out Adobe's AI was also trained on output from Midjourney and OpenAI
  • r/ArtistHate r on reddit
    Adobe's “Ethical” Firefly ML Model Was Trained on Midjourney Images