Adobe used images created by tools like Midjourney and uploaded to its stock marketplace by users, to train Firefly; Adobe says ~5% of images were AI-generated
Bloomberg
Context & Ripple Effects
Adobe had already chosen to accept and label generative-AI submissions on its stock marketplace, unlike Getty Images at the time, through its policy for labeled AI-made stock images. That marketplace decision created a pathway by which synthetic imagery could enter a library later used in model development.
The disclosure adds a concrete provenance question to Adobe’s cautious effort to embed generative AI in Creative Cloud amid uncertainty over creative professionals’ acceptance of the tools, as described in its early Creative Cloud AI rollout.
First-order effects
Adobe must account for the presence of user-uploaded, AI-generated Adobe Stock images in Firefly’s training set; it says those images represented about 5% of the corpus.
Contributors who supplied Stock images, including AI-made work, become directly relevant to how Adobe defines the boundary between marketplace licensing, labeling, and model-training use.
Second-order effects
Adobe Stock’s labeling and submission rules become more consequential: provenance signals on marketplace content can affect not only customer disclosure but the composition of downstream AI training data.
Rival creative-software and stock platforms face greater pressure to explain whether and how customer- or contributor-supplied assets can be reused for model development, particularly when those assets are themselves synthetic.
Third-order effects
If marketplaces increasingly double as training-data reservoirs, content provenance will become a core product and governance capability rather than a disclosure feature alone.
The episode points toward a more complex commercial chain in which AI-generated content can be sold, licensed, and then reused to improve new generation models; durable trust will depend on clear permissions and traceability.
The trend: Creative platforms are turning their content marketplaces into AI inputs, making provenance and reuse terms central to workflow-native AI commercialization.
Adobe says its Firefly dataset is ethical, but it's really not. About 14% of its training dataset is AI outputs mostly from Midjourney. And by 2030, the largely unethical AI in Adobe Stock will eclipse human contributions.
Scoop: Adobe's been keeping a little secret about its AI. Its Firefly model was branded as more ethical than rivals like Midjourney. But it was actually trained on images from them. With @rachelmetz https://www.bloomberg.com/...
Adobe has already hit copyright escape velocity. That's the theory that models can be trained on uncopyrightable images/text AI output. There never was a sustainable market for artists to make money licensing training data. https://finance.yahoo.com/...
🚨Scoop from me and @BrodyFord_: Adobe has touted Firefly as a safe and ethical alternative to competitors like Midjourney... but it also quietly trained Firefly on Midjourney images. Adobe's ‘Ethical’ Firefly AI Was Trained on Midjourney Images https://www.bloomberg.com/...
@mobileraj @BrodyFord_ This is not about data contamination— it's about the company making a deliberate choice to include AI-generated images from Adobe Stock in its dataset while publicly positioning itself as very different from the companies whose tools were used to make those…