Internal docs: Scale AI's efforts to train Google's Gemini were flooded with “spammy behavior” from unqualified independent contractors submitting shoddy work
How the startup that just scored a $14 billion investment from Meta struggled to contain ‘spammy behavior’ from unqualified contributors as it trained Gemini. Bluesky: @glinden , @darthbluesky , and @matt-dragon Bluesky: Greg Linden / @glinden : It doesn't bode well for ScaleAI or their customers that they didn't seem to know this. It's well known in the ML community that it's very hard to get high quality output from very cheap human judges, that merely combining the votes of many absurdly low pay workers is just garbage in garbage out. @darthbluesky : were the unqualified independent contractors using AI [embedded post] Matt Dragon / @matt-dragon : Destroying AI from the inside [embedded post]
Context & Ripple Effects
Scale AI’s role in model training has long depended on a large contractor base: earlier coverage described contractors auditing Bard responses across technical and legal subjects under tight deadlines, while Scale was also seeking to expand from labeling into higher-margin AI tools.
This quality-control report lands immediately after disclosures that Scale’s customer-training materials were broadly accessible through shared Google Docs, compounding scrutiny of the vendor’s operational controls for major AI customers.
First-order effects
- Google’s Gemini training work handled through Scale is exposed to a direct data-quality risk: unqualified contractor submissions can contaminate the human-feedback pipeline unless identified and removed.
- Scale faces immediate pressure to strengthen contributor screening, task validation, and audit trails for customer projects, particularly after its customer training documents were reportedly left accessible by link.
Second-order effects
- Customers that outsource training and evaluation work may demand more granular quality reporting and tighter access controls from Scale and comparable vendors.
- The finding reinforces the trade-off visible in time-constrained contractor audits of Bard answers: faster, lower-cost human evaluation can require more expensive review layers to make outputs usable.
Third-order effects
- If repeated across training vendors, frontier-model developers may shift more of high-stakes evaluation toward smaller, vetted expert pools or build internal oversight systems, rather than treating large contractor marketplaces as interchangeable supply.
- The sector’s differentiation will increasingly rest on provenance, validation, and secure handling of human-generated training data—not simply access to a large labeling workforce.
The trend: AI training-data providers are being judged less as labor marketplaces and more as quality-assurance and governance infrastructure for frontier-model development.