Scale AI and CAIS' Remote Labor Index benchmark, which tests AI agents on freelance tasks, finds the best AI could perform just <3% of the work, earning $1,810
A new benchmark measures how well AI agents can automate economically valuable chores. Human-level AI is still some ways off. Bluesky: @carlzimmer.com . Forums: r/artificial and r/technews Bluesky: Carl Zimmer / @carlzimmer.com : AI agents tried to do graphic design, video editing, game development, and administrative chores like scraping data. “Even the best could perform less than 3 percent of the work, earning $1,810 out of a possible $143,991,” writes @willknight.bsky.social www.wired.com/story/ai-age... Forums: r/artificial : AI Agents Are Terrible Freelance Workers r/technews : AI Agents Are Terrible Freelance Workers
Context & Ripple Effects
Freelance markets had already shown exposure to generative AI: jobs in categories where the technology excels fell on platforms such as Upwork in an earlier study of AI-exposed freelance categories. The Remote Labor Index adds a harder constraint: broad, paid freelance workflows remain difficult for current agents to execute end to end.
The result also fits prior evidence that technical capability alone does not make automation economical; an MIT assessment of vision-task automation found that deployment costs limited the set of wages worth automating. The relevant question for agents is therefore reliable output per useful task, not task demonstrations alone.
First-order effects
- Scale AI and CAIS now have a benchmark result showing that leading agents capture only a small share of the tested freelance-work value, limiting claims that these agents can immediately replace broad independent-worker roles.
- Teams considering agents for design, editing, game-development, and administrative workflows will need to retain human execution or review for most of the assessed work.
Second-order effects
- Agent vendors will face pressure to focus on bounded workflows with measurable completion and quality criteria, rather than sell general-purpose freelance automation.
- Buyers are likely to scrutinize agent economics—completion rates, rework, and supervision—more closely when comparing tools or deciding whether to shift work away from freelancers.
Third-order effects
- If benchmarks like this become widely used, agent competition could move from model-level capability claims toward audited, task-level reliability and cost comparisons.
- The longer-term labor effect may be uneven: AI can still reshape demand for selected tasks even while end-to-end automation remains uneconomic or unreliable across whole freelance jobs.
The trend: AI adoption is moving toward task-by-task evaluation, where commercially useful autonomy depends on reliability, oversight costs, and measurable economic output rather than headline model capability.