/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

← → days · ↑ ↓ browse · Enter similar · o open

Tech companies working with AI are shielding themselves from accountability by outsourcing data collection and model training to academic and non-profit groups

Yesterday, Meta's AI Research Team announced Make-A-Video, a “state-of-the-art AI system that generates videos from text.”

Waxy.org Andy Baio

Context & Ripple Effects

Days after Meta detailed Make-A-Video, its text-to-video generator that produces five-second clips but stays closed to outside access, this piece argues the announcement's structure matters as much as the model itself: by routing data collection and training through academic and non-profit groups, corporate labs get a research veneer that blurs who is accountable when the training data is contested.

The pattern the article flags has since hardened rather than faded — later reporting found the same intermediary structure at work at other major labs, and content owners have begun pushing back with demands for paid licensing.

First-order effects

  • Meta gets to present Make-A-Video as an academic research artifact while withholding model access, so criticism of its data practices lands on university and non-profit partners rather than on Meta's own product pipeline.

Second-order effects

  • The same outsourcing pattern scaled across the industry: a Proof investigation later found Apple, Nvidia, Anthropic, and others had trained on datasets built from YouTube transcripts without permission — exactly the accountability gap the academic-shield structure creates.
  • Rights holders responded by forcing the issue into the open: Alphabet, Meta, and OpenAI began negotiating content licenses with Hollywood studios, converting what was quietly scraped into a priced negotiation.

Third-order effects

  • If the pattern holds, AI training data splits into two regimes — laundered-through-intermediary collection for labs that want plausible deniability, and formal licensing markets for owners with leverage — pushing regulators toward provenance and disclosure requirements for training corpora.
  • Provenance tooling becomes the industry's self-defense: Meta's own release of Video Seal watermarking shows labs building origin-tracking infrastructure to manage the reputational risk their data practices created.

The trend: AI training-data practices are migrating from deniable scraping through academic intermediaries toward licensed, provenance-tracked pipelines as rights holders and investigators force accountability into the open.

Discussion

  • @flyingtrilobite Glendon Mellow on x
    “Asking for permission slows technological progress, but it's hard to take back something you've unconditionally released into the world.” -@waxpancake gets it. #artcred #WPAII #wordpromptaiimages #aiart https://waxy.org/...
  • @waxpancake Andy Baio on x
    Tech companies working with AI — from big names like Google and Meta to upstarts like Stability AI — are outsourcing data collection to academic/nonprofit research groups, shielding them from potential accountability and legal liability. https://waxy.org/...
  • @simonw Simon Willison on x
    Most interestingly: All 10,727,582 videos in that particular training set are from Shutterstock! You can read more about WebVid-10M here: https://m-bain.github.io/...