How Amsterdam's experiment to create a fair welfare AI model, which considered 15 characteristics to evaluate welfare applicants for potential fraud, failed
This story is a partnership between MIT Technology Review, Lighthouse Reports, and Trouw, and was supported by the Pulitzer Center.
Context & Ripple Effects
Amsterdam’s failed welfare-model experiment lands in a Dutch policy history already marked by a court rejection of the secretive SyRI welfare-risk system on human-rights grounds. It also follows scrutiny of Rotterdam’s welfare-fraud tooling, where a reconstruction found discrimination tied to ethnicity and gender in an Accenture-made system’s data and design.
The arc matters because Amsterdam’s effort appears to have aimed at fairness rather than merely deploying an opaque vendor system. Its failure suggests that expanding or refining the characteristics used in a welfare-risk model does not, by itself, resolve the accountability and discrimination problems associated with automated benefit enforcement.
First-order effects
- Amsterdam’s welfare-AI experiment fails to provide a workable model for screening applicants for potential fraud, limiting its immediate value as a public-service decision tool.
- Welfare applicants remain exposed to the consequences of risk-based scrutiny, while officials must rely more heavily on non-automated processes or reconsider how any future model is governed.
Second-order effects
- Other public bodies considering fraud-detection systems face stronger pressure to demonstrate that fairness claims hold in practice, not simply that a model excludes or balances selected characteristics.
- The result reinforces scrutiny of outside suppliers: prior reporting found automated-fraud vendors could be overpaid and under-supervised, making procurement oversight and independent evaluation more consequential.
Third-order effects
- If similar projects continue to fail, welfare automation may shift from a question of model tuning toward whether high-impact eligibility and fraud decisions can be made accountable enough to justify algorithmic risk scoring.
- The broader direction is toward operational assurance—testing, transparency, recourse, and human-rights safeguards—as prerequisites for public-sector AI, rather than optional checks after deployment.
The trend: This is one data point in the move from experimental public-sector risk scoring toward stricter governance of AI used to investigate or penalize residents.