/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

Over 274M records with personal identifiable info of Indian citizens exposed by unsecured MongoDB database, in what may be part of a massive scraping operation

On May 1st, I have discovered an unprotected and publicly indexed MongoDB database which contained 275,265,298 records …

Security Discovery Bob Diachenko

Context & Ripple Effects

This is the third population-scale open-database find in as many months of 2019, and the second on MongoDB specifically: in January researchers found the resumes of 202M+ Chinese users sitting on an unsecured MongoDB server, and in March an email validation firm left 763M unique email addresses in plaintext exposed. The suspected scraping operation behind the Indian dataset suggests these are not one-off misconfigurations but a supply chain for bulk personal data.

The scale matters: at roughly 275M records, this single index approaches a meaningful fraction of India's adult population, making it less a corporate breach than a national-level exposure of names, contact details, and identity attributes.

First-order effects

  • Indian citizens whose records sit in the database have no notification channel — unlike a breached company's customers, subjects of scraped aggregates typically never learn their data was exposed, and the operator of record is unknown.

Second-order effects

  • The find adds pressure on India's then-pending data protection legislation, giving lawmakers a concrete domestic example to cite, while cloud database vendors face renewed scrutiny over permissive default configurations that keep producing these exposures.

Third-order effects

  • If scraping-as-source becomes the norm, breach liability frameworks built around a single responsible company break down: when no one admits ownership of a 275M-record dataset, there is no obvious party to regulate, notify victims, or sue — a gap the later 800M-record Chinese database exposed for months in 2022 shows was never closed.

The trend: Unsecured cloud databases are becoming the dominant leakage channel for population-scale personal data, with scraping operations assembling national datasets faster than breach-notification regimes can assign responsibility for them.