/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

Internal docs: Scale AI's efforts to train Google's Gemini were flooded with “spammy behavior” from unqualified independent contractors submitting shoddy work

How the startup that just scored a $14 billion investment from Meta struggled to contain ‘spammy behavior’ from unqualified contributors as it trained Gemini. Bluesky: @glinden , @darthbluesky , and @matt-dragon Bluesky: Greg Linden / @glinden : It doesn't bode well for ScaleAI or their customers that they didn't seem to know this.  It's well known in the ML community that it's very hard to get high quality output from very cheap human judges, that merely combining the votes of many absurdly low pay workers is just garbage in garbage out. @darthbluesky : were the unqualified independent contractors using AI [embedded post] Matt Dragon / @matt-dragon : Destroying AI from the inside [embedded post]

Inc Sam Blum

Context & Ripple Effects

Scale AI’s role in model training has long depended on a large contractor base: earlier coverage described contractors auditing Bard responses across technical and legal subjects under tight deadlines, while Scale was also seeking to expand from labeling into higher-margin AI tools.

This quality-control report lands immediately after disclosures that Scale’s customer-training materials were broadly accessible through shared Google Docs, compounding scrutiny of the vendor’s operational controls for major AI customers.

First-order effects

  • Google’s Gemini training work handled through Scale is exposed to a direct data-quality risk: unqualified contractor submissions can contaminate the human-feedback pipeline unless identified and removed.
  • Scale faces immediate pressure to strengthen contributor screening, task validation, and audit trails for customer projects, particularly after its customer training documents were reportedly left accessible by link.

Second-order effects

  • Customers that outsource training and evaluation work may demand more granular quality reporting and tighter access controls from Scale and comparable vendors.
  • The finding reinforces the trade-off visible in time-constrained contractor audits of Bard answers: faster, lower-cost human evaluation can require more expensive review layers to make outputs usable.

Third-order effects

  • If repeated across training vendors, frontier-model developers may shift more of high-stakes evaluation toward smaller, vetted expert pools or build internal oversight systems, rather than treating large contractor marketplaces as interchangeable supply.
  • The sector’s differentiation will increasingly rest on provenance, validation, and secure handling of human-generated training data—not simply access to a large labeling workforce.

The trend: AI training-data providers are being judged less as labor marketplaces and more as quality-assurance and governance infrastructure for frontier-model development.

Discussion

  • @glinden Greg Linden on bluesky
    It doesn't bode well for ScaleAI or their customers that they didn't seem to know this.  It's well known in the ML community that it's very hard to get high quality output from very cheap human judges, that merely combining the votes of many absurdly low pay workers is just garba…
  • @darthbluesky @darthbluesky on bluesky
    were the unqualified independent contractors using AI [embedded post]
  • @matt-dragon Matt Dragon on bluesky
    Destroying AI from the inside [embedded post]