Social Science One and Facebook say they are now making a dataset with ~38M URLs available to academic researchers, two years after announcing the initiative
and acted with such hesitation — before releasing a massive batch of data. Beats the alternative. https://twitter.com/... Gary King / @kinggary : Today, we released a privacy-protective Facebook dataset with 10 Trillion cell values https://socialscience.one/... Tomorrow, I'll give a webcast for TIER on differential privacy, a version of which we used with these data. Here're my slides & how to register: https://gking.harvard.edu/... Issie Lapowsky / @issielapowsky : @persily “Lots of politicians who have talked about privacy in a negative way at Facebook are also saying we want academic research to happen. They say those two things and don't realize they're not compatible.” - @alexstamos http://protocol.com/... Ramya Krishnan / @2ramyakrishnan : 1. Social Science One is a worthy effort, but more than two years after it was first announced it has failed to deliver. As we head into the 2020 elections, Facebook's platform remains alarmingly opaque. The public deserves more. https://www.wsj.com/... Alberto Alemanno / @alemannoeu : #Facebook releases huge dataset enabling social scientists to study some of the most important questions of our time about the effects of social media on democracy and elections with information to which they have never before had access #disinformation https://socialscience.one/... Jacob Ward / @byjacobward : The data is vast — comprising 38 million URLs shared more than 100 times each, and tons of associated data about how they moved across Facebook — but significantly limited. https://socialscience.one/... Carl Miller / @carljackmiller : This is actually a really big deal. After two years of pressure, Facebook has begun to release a huge dataset of URLs shared across their platform. It's explicitly intended to help with research on misinformation. Paper: https://socialscience.one/... https://socialscience.one/... Jason Kint / @jason_kint : ...always reminds me of 2016 timeline for Facebook's acquisition of Crowdtangle which had a dataset: 11/8 - 2016 election 11/10 - Zuck stupid statement 11/11 - Crowdtangle acquisition 11/19 - Zuckerberg tail between legs https://twitter.com/... Jason Kint / @jason_kint : Good news all around. Although I am amused at the night and day between Facebook's approach to GDPR when mining commercially valuable personal data for its own interests vs protecting deidenitifed data from academics. https://twitter.com/... https://twitter.com/... Davey Alba / @daveyalba : @random_walker The data contains ~38 million URLs; Social Science One said it “processed about an exabyte of raw data” from the Facebook platform which “produced a dataset that has about 10 trillion numbers in it.” SS1 has corrected the wording in their announcement https://socialscience.one/... Jesse Blumenthal / @jessekblum : Big news from @SocSciOne: academic researchers now have access to the URL full dataset (detailed info on ~38 million urls shared on Facebook) — one of the largest datasets ever available for social science research https://www.wsj.com/... @profcarroll : Good news in finding a solution balancing privacy and research but article doesn't point out that the dataset excludes 2016, which is arguably still a crime scene as Cambridge Analytica is probably still an active investigation per FOIA'd Bannon 302s. https://www.wsj.com/...
Context & Ripple Effects
The release closes a loop opened in mid-2018, when Social Science One was announced as a vehicle for giving scholars access to a petabyte of anonymized user data to study misinformation in elections. That promise came straight out of the post-Cambridge Analytica reckoning, when Facebook suspended quiz-app firms like CubeYou and needed a credible channel for independent research.
First-order effects
- Academic researchers studying elections and misinformation finally get their hands on a privacy-protective dataset of ~38 million URLs — two years later than planned, with differential privacy applied to the 10 trillion cell values Gary King describes.
- Facebook gains a concrete counterexample to its researcher-access criticism, delivered through a named Harvard-led partnership rather than ad-hoc data sharing.
Second-order effects
- The two-year lag and heavy privacy engineering set the de facto terms for future platform research deals: scholars get curated, anonymized snapshots instead of live API access, a shift already visible in reports that Facebook weakened or disabled the tools researchers used to track political ads ahead of 2020.
- Rival platforms now face pressure to match this model or be cast as less transparent — but the same structure invites the friction documented later, when researchers described Meta's pre-publication review and restrictive contracts as obstacles to unflattering findings.
Third-order effects
- If the pattern holds, independent study of major platforms runs through company-negotiated datasets rather than public endpoints, making research access itself a governed permission boundary — with the 2023 Meta-collaborated studies on algorithmic influence showing both what that access enables and how much of the framing sits with the platform.
The trend: Platform research is moving from open APIs toward curated, privacy-engineered datasets released on the company's timeline and terms.