Amazon's Rekognition mistakes women as men 19% of the time, and darker-skinned women as men 31% of the time, more than similar services from IBM and Microsoft
In new tests, Amazon's system had more difficulty identifying the gender of female and darker-skinned faces than similar services from IBM and Microsoft.
Context & Ripple Effects
This test extends a documented pattern: a 2018 study already found 21%-35% error rates for dark-skinned women across Microsoft, IBM, and China's Megvii, and this new round adds Amazon — the one major vendor whose Rekognition performs worst of the three on both female faces overall (19% misgendered) and darker-skinned women (31%). The comparison against IBM and Microsoft is the story: same task, same test, materially different demographic error rates.
The finding lands mid-arc rather than at the end of it. Months later, an investigation found Rekognition misgenders trans, queer, and nonbinary individuals by design, and a 2020 speech-recognition study showed the same vendors' audio systems erring 35% with black speakers versus 19% with white — evidence that the gap is structural across modalities, not a one-off Rekognition flaw.
First-order effects
- Amazon's Rekognition enters vendor comparisons as the demographic-accuracy laggard, giving IBM and Microsoft a concrete, testable differentiator when public agencies and enterprises evaluate face-analysis APIs.
- Any customer using Rekognition for gender classification inherits a 19% error rate on women and 31% on darker-skinned women — a direct accuracy and liability exposure that lighter-error competitors do not carry.
Second-order effects
- Procurement for face-recognition contracts shifts toward published demographic error benchmarks, pressuring Amazon to retrain or disclose — and forcing all three vendors to treat per-demographic accuracy as a competitive metric rather than an internal QA detail.
- The documented misgendering of trans and nonbinary users compounds the commercial risk: buyers in identity-verification and access-control markets face reputational exposure that favors vendors with lower error rates on the same populations.
Third-order effects
- If the pattern holds across face and speech systems from the same five vendors, demographic error gaps become the trigger for external auditing and regulation of biometric AI — with binary gender classification itself, as the Jezebel investigation argued, an unreliable category to build products on.
- Vendors that cannot demonstrate uniform accuracy across demographics face a structural split in their customer base: consumer-scale deployments absorb the reputational cost while accuracy-sensitive buyers consolidate around the few providers that pass third-party tests.
The trend: Commercial AI systems are being held to per-demographic accuracy benchmarks by third-party tests, and vendors that fail them — Amazon most visibly here — are turning bias measurement into a competitive and regulatory fault line.