/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

In a case study, researchers estimate that between 33% and 46% of Mechanical Turk workers used LLMs when completing a text summarization task

File this one under inevitable, but hilarious.  Mechanical Turk is a service that from its earliest days seemed to invite shenanigans …

TechCrunch Devin Coldewey

Context & Ripple Effects

Mechanical Turk has a documented history of data-quality anxiety: psychology researchers flagged a possible bot-driven uptick in low-quality survey responses back in 2018, and the same year a UN survey found workers on microtask platforms averaging $6.54/hr in the US. The platform's economics have always rested on cheap human effort — workers earning pennies to train AI — with little recourse when underpaid or shortchanged.

The new wrinkle is that the cheap human effort may increasingly be LLM-mediated. If a third to nearly half of workers on a summarization task outsource the actual cognition to a model, the buyer is paying human-task prices for machine output — and the 2018 bot scare looks like a preview of a much larger contamination problem.

First-order effects

  • Researchers and any buyer of crowd-work output face immediate data-integrity risk: a summarization study's results are now partly machine-generated, so findings built on Turk responses need LLM-detection screening before they can be trusted.
  • Workers who do the task themselves now compete against peers who delegate to LLMs — the honest worker's per-task time cost rises relative to cheaters at the same piece rate.

Second-order effects

  • Amazon is pushed toward verification tooling — attention checks, style fingerprinting, LLM-output detectors — adding overhead to a marketplace whose margins depend on frictionless microtasks, echoing the quality controls the 2018 bot scare first demanded.
  • Task pricing comes under pressure: if a model can do summarization at near-zero marginal cost, requesters have a reference price far below the human piece rate, squeezing the already thin earnings documented in the UN survey.

Third-order effects

  • Crowd work may bifurcate: tasks where a human's identity, judgment, or lived experience is the product survive, while pure cognitive microtasks migrate to LLM pipelines — with the marketplace's value shifting from supplying labor to certifying that labor is human.
  • If contamination is this widespread in research samples, academic and industry norms for crowd-sourced data may harden around provenance verification, making 'verified human' the scarce commodity Turk sells.

The trend: Crowd-labor platforms are absorbing LLMs faster than their quality controls can adapt, turning 'is this output actually human-made?' into the core verification problem of the microtask economy.

Discussion

  • @manoelribeiro Manoel on x
    One of our key sources of human data is no longer fully “human”! We estimate that 33-46% of crowd workers on MTurk used large language models (LLMs) in a text production task - which may increase as ChatGPT and the like become more popular and powerful. https://arxiv.org/... [ima…
  • @maggiekb1 Maggie Koerth on x
    So a major place scientists go to run experiments is about 1/3 ai now. And those results are still getting treated as human, including in tasks used to train ai. https://twitter.com/...
  • @jeremybowers Jeremy Bowers on x
    It's too late, we poisoned MTurk https://twitter.com/...
  • @noamchompers Noam Chompers on x
    In addition to being a disaster for a lot of empirical work, this would also be very funny https://twitter.com/...
  • @emollick Ethan Mollick on x
    In a great (& destructive) irony, Mechanical Turk is just AI, now. The MTurk crowdworking platform is a major place for researchers (and companies) to get humans to do small tasks & experiments. But this paper finds 33-46% of Turkers use LLMs to do tasks. https://arxiv.org/... [i…
  • @rdbinns @rdbinns on x
    Our AI couldn't really do the job so we outsourced it to a human who could, but didn't, who outsourced it to another AI that couldn't really do the job either. https://twitter.com/...
  • @atabarrok Alex Tabarrok on x
    Wait, so the mechanical Turk now actually is a mechanical Turk?! What a time to be alive! https://twitter.com/...
  • @perttu_h @perttu_h on x
    Large Language Models are making MTurk etc. fundamentally unreliable in collecting text data, as predicted in our CHI paper (https://dl.acm.org/...) https://twitter.com/...
  • @veredshwartz Vered Shwartz on x
    This seems consistent with my experience lately. We had to manually verify text written by annotators and filtered out quite a bit of text that looked LM-generated. In addition, the annotations for verification/ranking tasks were so noisy that we decided to not do them in mturk. …
  • @rstephens Robert Stephens on x
    “Rechanical Turk” https://www.techmeme.com/...
  • @noahpinion Noah Smith on x
    There goes every MTurk study... https://twitter.com/...