/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

Researchers find unprotected database owned by an email validation company with 150 GB of plaintext marketing data, including 763M unique email addresses

Lily Hay Newman / Wired :

Wired Lily Hay Newman

Context & Ripple Effects

This lands two months after researchers surfaced Collection #1, a database claiming roughly 773 million unique email addresses alongside 21 million passwords — and the scale here is nearly identical, at 763 million addresses. The difference is provenance: where Collection #1 looked like an aggregation of prior breaches, this trove sat on the servers of an email validation company, a business whose entire function is to confirm which addresses are real and deliverable.

That makes it a working snapshot of who can actually be reached by email, held in plaintext by a firm most consumers have never heard of. It also fits a run of misconfigured-database discoveries through 2019, from the 274M-record MongoDB exposure of Indian citizens' PII to an ad agency leaving 150K+ form submissions open — a pattern Wired's Lily Hay Newman has tracked repeatedly.

First-order effects

  • The email validation company faces immediate remediation and reputational damage, while every marketer whose campaign data sat in those 150 GB now has to assess whether their customer lists were exposed.
  • The 763 million address owners move from generic spam risk to confirmed-target risk: validation data tells attackers which addresses are active, making phishing and targeted campaigns more efficient.

Second-order effects

  • Rival email verification vendors inherit a trust problem — enterprise buyers of list-hygiene services will demand proof of security posture, raising sales friction across the category.
  • Spam and phishing operators gain a pre-validated corpus that compounds collections like Collection #1, improving hit rates for credential-stuffing and mule-account recruitment without any new breach required.

Third-order effects

  • If the 2019 pattern — MongoDB, ad agencies, validation firms — plus later finds like the 184M-record Elastic database of unknown ownership holds, unsecured third-party aggregators become a standing exposure channel independent of hacking, shifting regulatory attention from breached companies to the brokers and processors nobody audits.
  • Cloud storage defaults and shared-responsibility models come under pressure as the recurring failure mode: the structural fix being argued is secure-by-default object storage and liability rules that make 'we forgot to lock the bucket' untenable for data intermediaries.

The trend: The biggest email exposures are migrating from hacked enterprises to unsecured data brokers and aggregators, whose plaintext troves of validated contact data are becoming the raw supply chain for spam and phishing at internet scale.