Study: by the start of 2019, over 26M consumers have added their DNA to four leading commercial ancestry and health databases, primarily AncestryDNA
The genetic genie is out of the bottle. And it's not going back. — As many people purchased consumer DNA tests in 2018 …
Context & Ripple Effects
This study quantifies how fast consumer genomics scaled: from Apple's early ResearchKit DNA-testing plans in 2015 to over 26 million consumers in four leading databases by early 2019, with AncestryDNA holding most of them. The scale matters because of what researchers had already shown — that 60% of Americans of European descent can be identified through public DNA databases even if they never took a test themselves.
The database growth also landed amid a governance vacuum: firms including Ancestry and 23andMe had only just signed voluntary transparency pledges on third-party sharing, and FamilyTreeDNA's decision to let the FBI search its 1M+ profiles showed how quickly commercial databases could become law-enforcement infrastructure without user consent.
First-order effects
- AncestryDNA's dominant share means a single company now holds genetic data for a large fraction of the identifiable US population of European descent, making its terms-of-service and sharing decisions consequential for millions who never opted into anything.
- FamilyTreeDNA's FBI cooperation sets an immediate precedent the larger databases must respond to — users of all four services now face the question of whether their data is searchable by police.
Second-order effects
- Voluntary industry pledges like the one Ancestry and 23andMe signed come under pressure to harden into enforceable rules, since the 26M-user scale makes 'be upfront' commitments inadequate against both law-enforcement access and breach risk.
- Relatives of test-takers — the majority who never submitted DNA but are identifiable through relatives' uploads — become stakeholders in companies' privacy policies they never agreed to, raising exposure for any firm marketing family-based testing.
Third-order effects
- If the pattern holds, consumer genetics converges toward de facto national identification infrastructure built by private companies, where each breach or law-enforcement partnership — as later seen in the 23andMe hack that exposed 6.9M customers' ancestry data — becomes a systemic event rather than a single-company incident.
- The gap between voluntary guidelines and the sensitivity of pooled genetic data points toward regulatory intervention defining consent not just for users but for genetically identifiable non-users.
The trend: Consumer DNA databases are scaling from genealogy novelty to population-scale identification infrastructure faster than consent frameworks and regulation can keep up.