On September 30, 2026, AWS will shut down Amazon Mechanical Turk, a marketplace that once routed penny-paid AI-training tasks to more than 500,000 people. On May 7, Scale AI won a $500 million Defense Department contract to sift data and assist decisions. Both sit in the market for human judgment, but the gap between a penny task and a defense contract is larger than price.
Mechanical Turk made judgment cheap—and hard to verify
On September 2, AWS said it would close Mechanical Turk “following an assessment”, after stopping new customer registrations and placing the service in maintenance. The company offered no further explanation.
Amazon launched Mechanical Turk in 2005 to answer a specific question: how could a company buy a few seconds of human effort without hiring a person, managing an offshore vendor, or knowing who performed the work? The marketplace bypassed traditional outsourcing by turning manual, monotonous tasks into inventory.
Requesters could isolate work, describe it in advance, check it cheaply, and purchase it one click at a time. Workers did not need the customer’s context because the task specification was supposed to contain everything that mattered. Requesters could sample quality after delivery, and another contributor could take the next task if one disappeared.
By 2016, Mechanical Turk had become a marketplace of more than 500,000 workers training AI for pennies per task. A vast crowd was decomposing intelligence into pieces small enough to price.
The same design weakened the signal that buyers thought they were purchasing. By 2018, psychology researchers warned that bots might be degrading survey data gathered through Mechanical Turk. A later case study estimated that 33% to 46% of workers used large language models for a text-summarization assignment.
AWS has not tied those findings to its closure decision. They expose the marketplace’s structural weakness: once requesters optimized for accepted tasks, a human account stopped proving that a person had supplied the judgment. Mechanical Turk made labor liquid by making worker identity incidental, then reached a market in which identity and provenance became part of the product.
Agents move human work to the guardrail
AI agents change the unit of work. A model that answers a question produces an output to inspect. A model that uses tools, follows a multistep plan, and enters organizational systems produces a chain of actions whose errors can survive long enough to become someone else’s operating condition.
AI coding agents can now complete complex projects with minimal oversight. In May, the United States, United Kingdom, Australia, Canada, and New Zealand warned that organizations were giving agents more access than they could safely monitor. More capable agents need fewer human touches during execution; broader permissions raise the price of one missed failure.
A customer-service system can pass routine queries among specialized agents and escalate when it detects frustration. The person receives the case because the standard answer has stopped working. That reviewer must understand why, decide what authority is required, and leave enough evidence for the system to learn from the exception.
As routine interventions disappear, the remaining reviewers need more context, stronger judgment, and clearer authority over consequential actions.
Frontier evaluation demands skill without guaranteeing its price
Model evaluators make the shift clearest. Commodity labelers apply a known category, while frontier evaluators hunt for the category a system does not handle: the ambiguous instruction, adversarial prompt, domain-specific mistake, or plausible answer whose harm becomes visible only in context.
Leaked documents showed Outlier and Scale AI freelancers writing and assessing prompts involving suicide, abuse, terrorism, torture, and animal cruelty. The instructions reportedly encouraged creativity while prohibiting child sexual abuse material. Workers were improvising inside an explicitly managed risk perimeter.
Companies have not consistently paid for that difficulty. The same reporting described some stress-testing assignments as minimum-wage gigs, even though workers faced disturbing material and exercised judgment that a simple instruction could not capture. AI companies can recognize scarce expertise at the customer interface while treating the person supplying it as replaceable.
Some data companies are now replacing low-cost labelers with highly paid specialists in fields such as finance as reasoning models demand harder examples. The shift has not reached every worker, leaving job quality behind the work itself.
Models can judge other models, validate answers against knowledge bases, and absorb routine checks through automatic metrics. People increasingly choose the rubric, resolve conflicts among signals, construct adversarial cases, and decide whether a measured improvement is safe enough to deploy.
Scale’s contract puts a price on accountability
Customers pay providers to recruit qualified people, calibrate their judgments, protect sensitive work, verify provenance, manage disagreement, and deliver a result that a demanding organization can use. Crowd size remains one input into that managed trust.
Scale AI completed a $100 million Defense Department deal in 2025, then moved into a larger institutional role the next year.
Awarded through the US Chief Digital and AI Office, the later contract covers helping the department sift through data and assist decision-making. At that scale, the department also needs a vendor to coordinate work, document standards, resolve disagreement, and answer for delivery.
After Meta’s $14.3 billion investment in Scale AI, rival providers reported increased client interest from customers concerned about Scale’s independence. Customers had widened the quality calculation to include the supplier’s ownership and incentives.
The exception layer belongs in the operating design
Companies often describe humans in the loop as a temporary patch: the model handles most cases while a person catches residual errors until the next release. Calling reviewers temporary preserves the old automation blueprint, where progress means removing the person from one more box.
Federal agencies already have to plan around durable accountability. The US Office of Management and Budget requires every agency to file an annual AI report and place a senior leader over its AI systems, even as existing law remains unprepared to assign liability when agents act beyond their intended bounds.
Teams deploying agents must decide which actions require approval, which anomalies trigger escalation, what evidence they retain, and how an exception becomes training material instead of an isolated rescue. A useful system separates low-risk repetition from high-consequence ambiguity, routes each case to the appropriate machine or person, records why the handoff occurred, and feeds the resolution back into the model and operating procedure.
Model developers already rely on people to write ideal responses for supervised fine-tuning and rank competing outputs for reinforcement learning from human feedback. Adding contributors can increase volume, but their judgments determine what behavior the system learns. Developers often need a smaller group with better context, defined authority, and responsibility for consequences.
When OpenAI asked contractors to upload workplace documents, it left confidential-data removal to those contractors. A checkpoint can shift risk to a reviewer without giving that person broader authority. Durable review requires explicit rules for rejecting material, establishing ownership, and bearing the cost of a bad decision.
Mechanical Turk let a buyer purchase a result without knowing the worker, and AWS will switch it off on September 30. Scale AI’s $500 million contract sits at the other end of the same market, surrounded by standards, institutional controls, and a supplier responsible for delivery. The penny task counted a stranger’s click; the agent-era workflow needs a named reviewer with authority to stop the machine.