Psychology researchers claim bots might be responsible for uptick in low-quality data in recent surveys they conducted via Amazon's Mechanical Turk
Emily Dreyfuss / Wired :
Context & Ripple Effects
Mechanical Turk has spent years as the internet's invisible back office: coverage of how its workers earn pennies to train AI and of Expensify routing user-submitted documents through the platform established it as cheap, on-demand human labor. The new wrinkle is that some of that labor may not be human at all — psychology researchers report an uptick in low-quality survey responses they attribute to bots.
That lands on a workforce already strained: a later survey found Turkers feel they have little recourse for underpayment, technical errors, or abuse. If requesters start distrusting the pool itself, the people paying the price of tighter screening will likely be the legitimate workers, not the bots.
First-order effects
- Psychology researchers running surveys on Mechanical Turk can no longer take response validity for granted — their findings are contaminated at the source, forcing re-runs and stricter screening before results are published.
- Amazon faces a platform-integrity problem on Mechanical Turk: the marketplace's core promise to requesters, that tasks are completed by humans, is the thing being questioned.
Second-order effects
- Requesters respond by layering attention checks, qualification tests, and identity verification onto every task, raising the effective cost of crowdsourced work and squeezing the margins of legitimate workers who already report weak recourse when things go wrong.
- Rival crowdsourcing and data-labeling vendors get a selling point — verified-human workforces — while academic buyers begin weighing whether survey panels built on open marketplaces are worth the cleanup cost.
Third-order effects
- If bot contamination is structural rather than episodic, open micro-task marketplaces drift toward gated, verified-labor models, and the cheap-crowd substrate that trained early AI systems and powered consumer services like Expensify's document processing becomes harder to rely on.
- The episode foreshadows a broader measurement crisis: as automated agents get better at passing as humans online, any dataset gathered from anonymous crowd participants — surveys, labels, ratings — needs provenance guarantees it currently lacks.
The trend: Crowdsourced micro-work platforms are entering a trust crisis in which the line between human gig workers and automated agents blurs, pushing research and AI-data buyers toward verified-identity labor.