AI labs are buying Slack, Jira, and email archives from defunct startups to build “reinforcement learning gyms” and train AI agents in simulated workplaces
Defunct startups are being liquidated for their Slack archives, Jira tickets, and email threads—operational exhaust that AI labs now treat as premium training data.
ForbesAnna Tong
Context & Ripple Effects
The coverage traces a widening search for training inputs beyond conventional labeling: OpenAI previously expanded contractor work for data labeling and software-engineering instruction, while Mercor was reported to be soliciting professionals’ prior work materials for AI training.
This purchase of failed companies’ operational records shifts that search toward end-to-end workplace traces. It also aligns with labs’ broader push for durable B2B adoption, including incentives for startups and increased hiring of forward-deployed engineers.
First-order effects
AI labs gain access to linked workplace artifacts—messages, tasks, and email threads—that can be used to construct simulated environments for training agents on multi-step work.
Defunct startups’ archives become liquidation assets, while the labs can train against operational context rather than isolated prompts or labels.
Second-order effects
Labs competing for enterprise-agent revenue have another route to improve workplace-task performance, complementing customer-adoption efforts and forward-deployed engineering teams.
The reported market for prior work materials makes ownership, consent, and employer-IP boundaries more consequential for data brokers, liquidators, and companies whose records may be resold.
Third-order effects
If this practice expands, operational exhaust could become a distinct input market for agent development, alongside human labeling and voluntarily contributed data.
Enterprise AI differentiation may increasingly depend on access to realistic, connected work environments—not just model capability—raising sustained pressure for clearer data provenance and reuse controls.
The trend: Agent builders are moving from training on discrete examples toward acquiring the contextual records needed to simulate and automate real workplace workflows.
AI labs are buying internal communications of defunct startups to train their agents. Emails, Slack archives, etc. Personally identifiable info is removed by data resellers. But how would you feel knowing your former board/CEO is selling your comms to recover losses/pay debts?
now's the right time to announce that I'm starting the Nth RL env company, but only focused on successful company data if you're the founder of a successful startup (1b+ arr only), please DM me with a zip of your git repo and HRIS data
AI labs are paying hundreds of thousands of dollars to buy email, Slack and Jira threads from dead startups as feedstock for ‘reinforcement learning gyms,’ which specialize in using defunct company data to build simulated work environments https://www.forbes.com/...