Tech companies working with AI are shielding themselves from accountability by outsourcing data collection and model training to academic and non-profit groups
Yesterday, Meta's AI Research Team announced Make-A-Video, a “state-of-the-art AI system that generates videos from text.”
Context & Ripple Effects
Days after Meta detailed Make-A-Video, its text-to-video generator that produces five-second clips but stays closed to outside access, this piece argues the announcement's structure matters as much as the model itself: by routing data collection and training through academic and non-profit groups, corporate labs get a research veneer that blurs who is accountable when the training data is contested.
The pattern the article flags has since hardened rather than faded — later reporting found the same intermediary structure at work at other major labs, and content owners have begun pushing back with demands for paid licensing.
First-order effects
- Meta gets to present Make-A-Video as an academic research artifact while withholding model access, so criticism of its data practices lands on university and non-profit partners rather than on Meta's own product pipeline.
Second-order effects
- The same outsourcing pattern scaled across the industry: a Proof investigation later found Apple, Nvidia, Anthropic, and others had trained on datasets built from YouTube transcripts without permission — exactly the accountability gap the academic-shield structure creates.
- Rights holders responded by forcing the issue into the open: Alphabet, Meta, and OpenAI began negotiating content licenses with Hollywood studios, converting what was quietly scraped into a priced negotiation.
Third-order effects
- If the pattern holds, AI training data splits into two regimes — laundered-through-intermediary collection for labs that want plausible deniability, and formal licensing markets for owners with leverage — pushing regulators toward provenance and disclosure requirements for training corpora.
- Provenance tooling becomes the industry's self-defense: Meta's own release of Video Seal watermarking shows labs building origin-tracking infrastructure to manage the reputational risk their data practices created.
The trend: AI training-data practices are migrating from deniable scraping through academic intermediaries toward licensed, provenance-tracked pipelines as rights holders and investigators force accountability into the open.