Leaked docs show Outlier and Scale AI use freelancers to write prompts about suicide, abuse, and terrorism to stress-test AI, urging creativity but banning CSAM
Effie Webb / Business Insider : Bluesky: @theopriestley.com and @hypervisible Bluesky: Theo / @theopriestley.com : Check LinkedIn and other job boards, they're awash with these types of freelance gigs now paying min. wage to spend hours writing prompts and assessing the outputs to “improve” AI. [embedded post] @hypervisible : “Freelancers are encouraged to ‘stay creative’ as they test AI with prompts about torture or animal cruelty, leaked training documents obtained by Business Insider show.”
Context & Ripple Effects
This reporting adds detail to the expanding human labor layer behind model safety work. Scale AI had already recruited subject-matter writers, including writers and poets for model-improvement tasks, while more recent coverage showed journalists taking data-training assignments such as fact-checking and prompt drafting.
The new documents distinguish adversarial testing from ordinary content generation: freelancers are asked to devise and assess extreme prompts, with an explicit CSAM boundary. It follows earlier reporting on contractors labeling violent and toxic material for ChatGPT, underscoring that harmful-content safeguards depend on people encountering that material.
First-order effects
- Outlier and Scale AI’s freelance workforce is directly exposed to suicide, abuse, terrorism, torture, and animal-cruelty scenarios as part of prompt creation and output assessment, while CSAM is excluded from the assignment scope.
- The companies gain a scalable source of adversarial test cases and human judgments that can be used to identify failures in models’ harmful-content responses.
Second-order effects
- The work expands the market for specialized AI-training contractors beyond labeling into red-teaming and evaluation; applicants with relevant editorial or analytical skills may be drawn into this segment, as seen in journalists taking AI-training work.
- Model developers and their vendors face greater pressure to define task boundaries, worker guidance, and escalation processes, since the quality of safety testing depends on both prompt coverage and consistent human evaluation.
Third-order effects
- If this outsourcing pattern persists, AI safety evaluation will become a larger, more formalized labor supply chain rather than a function performed solely inside model labs.
- The human cost and governance of repeated exposure to harmful material could become a differentiator among AI-data vendors, strengthening demand for auditable worker protections and evaluation standards.
The trend: AI development is moving toward labor-intensive, outsourced adversarial evaluation as models require broader testing against harmful and policy-sensitive behavior.