The resumes of 202M+ Chinese users, with personal data including home addresses and mobile numbers, were exposed online on an unsecured MongoDB database server
Data appears to have originated from a data scraping app that collected resumes from Chinese job portals.
Context & Ripple Effects
This exposure fits a pattern that kept repeating through 2019: researchers had already flagged over 590M resumes leaked by Chinese HR-focused companies via exposed databases within three months of this find, and a headhunting firm's misconfigured Elasticsearch cluster later spilled 20M+ more resumes. The common thread is that resume data scraped from job portals ends up pooled in cloud databases that are never locked down.
The scale here — 202M+ users with home addresses and mobile numbers on an open MongoDB server — also foreshadowed bigger finds, including [[a:982335|a Chinese database of up to 800M records with resident IDs and face images exposed for months in 2022]], and a parallel 274M-record MongoDB exposure of Indian citizens' data that same spring.
First-order effects
- Over 202 million Chinese job seekers now have home addresses and mobile numbers exposed to anyone who found the server before it was secured, with no way to recall data already copied.
- The scraping app that harvested resumes from Chinese job portals becomes the likely focal point for questions about how non-public personal details left those platforms at all.
Second-order effects
- Chinese job portals face pressure to restrict bulk scraping and API access, since their users' data is being siphoned into third-party databases they don't control.
- HR and recruiting firms holding aggregated resume troves — the segment behind the 590M-resume leak findings — face scrutiny over whether their own storage follows the same insecure defaults.
Third-order effects
- If scraped-data pooling continues unchecked, exposed databases become supply chains for dark-web resale, as later seen with reports of 538M Weibo user records offered for sale.
- A steady drumbeat of misconfigured MongoDB and Elasticsearch incidents points toward default-on security configurations and enforcement against companies that aggregate personal data without securing it.
The trend: Cloud databases left on insecure defaults are turning scraped personal data into recurring, industrial-scale breach waves across China's HR and social platforms.