/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

After criticism from users, Adobe Chief Product Officer Scott Belsky denies using customer projects to train the company's generative AI services

Bloomberg : Tweets: @tcpeter Tweets: Tim Peter / @tcpeter : This does lead to questions about where many generative AI projects (or *any* LLM) will get high quality training data. “The internet” is not necessarily a good answer. Plenty of examples exist where using internet data made AI's dumber https://twitter.com/...

Bloomberg

Context & Ripple Effects

Scott Belsky's denial lands mid-criticism: users had flagged that Adobe's services could touch their work, and his response — that customer projects are not used to train Adobe's generative AI — is an attempt to close the question at the product-leadership level rather than through legal text. The exchange also surfaces the harder problem Tim Peter raises: where any LLM or generative model gets high-quality training data, since 'the internet' has repeatedly made models worse.

The denial did not end the scrutiny. Later coverage shows the same tension recurring: Adobe was found to have trained Firefly partly on images generated with rivals' tools like Midjourney that users uploaded to its stock marketplace, then faced a backlash over terms of service allowing automated and manual access to user content, before finally issuing an [[a:867066|explicit clarification that it does not train Firefly on customer content and will never claim ownership of customer work]].

First-order effects

  • Adobe's creative users get a direct assurance from its Chief Product Officer that their project files are off-limits for generative AI training — but the assurance is verbal, not yet written into terms they can rely on.
  • Belsky's statement puts Adobe's own training-data pipeline under immediate examination, since the company must now reconcile the denial with what actually feeds Firefly.

Second-order effects

  • Rivals in creative software gain a positioning wedge: any competitor that contractually guarantees customer content stays out of training data can attack Adobe on trust rather than features.
  • Adobe is pushed toward codifying data-use promises in its terms of service — which is exactly where the later backlash and clarification cycle played out.

Third-order effects

  • If the pattern holds, training-data permissions become a standard contractual clause across SaaS: vendors that monetize user content for model training face churn pressure, forcing an industry-wide split between platforms that train on customer data and those that explicitly don't.
  • The deeper constraint Belsky's denial points to is supply: as clean internet-scale data proves unreliable, proprietary and licensed corpora become the contested resource, and who holds rights to high-quality creative work becomes a structural advantage.

The trend: Generative AI vendors are being forced from informal assurances toward explicit, enforceable guarantees about whether customer content trains their models, as trust over data use turns into a competitive fault line.

Discussion

  • @tcpeter Tim Peter on x
    This does lead to questions about where many generative AI projects (or *any* LLM) will get high quality training data. “The internet” is not necessarily a good answer. Plenty of examples exist where using internet data made AI's dumber https://twitter.com/...