/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

Social Science One and Facebook say they are now making a dataset with ~38M URLs available to academic researchers, two years after announcing the initiative

and acted with such hesitation — before releasing a massive batch of data. Beats the alternative. https://twitter.com/... Gary King / @kinggary : Today, we released a privacy-protective Facebook dataset with 10 Trillion cell values https://socialscience.one/... Tomorrow, I'll give a webcast for TIER on differential privacy, a version of which we used with these data. Here're my slides & how to register: https://gking.harvard.edu/... Issie Lapowsky / @issielapowsky : @persily “Lots of politicians who have talked about privacy in a negative way at Facebook are also saying we want academic research to happen. They say those two things and don't realize they're not compatible.” - @alexstamos http://protocol.com/... Ramya Krishnan / @2ramyakrishnan : 1. Social Science One is a worthy effort, but more than two years after it was first announced it has failed to deliver. As we head into the 2020 elections, Facebook's platform remains alarmingly opaque. The public deserves more. https://www.wsj.com/... Alberto Alemanno / @alemannoeu : #Facebook releases huge dataset enabling social scientists to study some of the most important questions of our time about the effects of social media on democracy and elections with information to which they have never before had access #disinformation https://socialscience.one/... Jacob Ward / @byjacobward : The data is vast — comprising 38 million URLs shared more than 100 times each, and tons of associated data about how they moved across Facebook — but significantly limited. https://socialscience.one/... Carl Miller / @carljackmiller : This is actually a really big deal. After two years of pressure, Facebook has begun to release a huge dataset of URLs shared across their platform. It's explicitly intended to help with research on misinformation. Paper: https://socialscience.one/... https://socialscience.one/... Jason Kint / @jason_kint : ...always reminds me of 2016 timeline for Facebook's acquisition of Crowdtangle which had a dataset: 11/8 - 2016 election 11/10 - Zuck stupid statement 11/11 - Crowdtangle acquisition 11/19 - Zuckerberg tail between legs https://twitter.com/... Jason Kint / @jason_kint : Good news all around. Although I am amused at the night and day between Facebook's approach to GDPR when mining commercially valuable personal data for its own interests vs protecting deidenitifed data from academics. https://twitter.com/... https://twitter.com/... Davey Alba / @daveyalba : @random_walker The data contains ~38 million URLs; Social Science One said it “processed about an exabyte of raw data” from the Facebook platform which “produced a dataset that has about 10 trillion numbers in it.” SS1 has corrected the wording in their announcement https://socialscience.one/... Jesse Blumenthal / @jessekblum : Big news from @SocSciOne: academic researchers now have access to the URL full dataset (detailed info on ~38 million urls shared on Facebook) — one of the largest datasets ever available for social science research https://www.wsj.com/... @profcarroll : Good news in finding a solution balancing privacy and research but article doesn't point out that the dataset excludes 2016, which is arguably still a crime scene as Cambridge Analytica is probably still an active investigation per FOIA'd Bannon 302s. https://www.wsj.com/...

Social Science One

Context & Ripple Effects

The release closes a loop opened in mid-2018, when Social Science One was announced as a vehicle for giving scholars access to a petabyte of anonymized user data to study misinformation in elections. That promise came straight out of the post-Cambridge Analytica reckoning, when Facebook suspended quiz-app firms like CubeYou and needed a credible channel for independent research.

First-order effects

  • Academic researchers studying elections and misinformation finally get their hands on a privacy-protective dataset of ~38 million URLs — two years later than planned, with differential privacy applied to the 10 trillion cell values Gary King describes.
  • Facebook gains a concrete counterexample to its researcher-access criticism, delivered through a named Harvard-led partnership rather than ad-hoc data sharing.

Second-order effects

  • The two-year lag and heavy privacy engineering set the de facto terms for future platform research deals: scholars get curated, anonymized snapshots instead of live API access, a shift already visible in reports that Facebook weakened or disabled the tools researchers used to track political ads ahead of 2020.
  • Rival platforms now face pressure to match this model or be cast as less transparent — but the same structure invites the friction documented later, when researchers described Meta's pre-publication review and restrictive contracts as obstacles to unflattering findings.

Third-order effects

  • If the pattern holds, independent study of major platforms runs through company-negotiated datasets rather than public endpoints, making research access itself a governed permission boundary — with the 2023 Meta-collaborated studies on algorithmic influence showing both what that access enables and how much of the framing sits with the platform.

The trend: Platform research is moving from open APIs toward curated, privacy-engineered datasets released on the company's timeline and terms.

Discussion

  • @random_walker Arvind Narayanan on x
    This is one of the most high-profile applications of differential privacy to date. Two documents by researchers who worked on the data release help understand why it took so long: Paper: https://arxiv.org/... Announcement: https://socialscience.one/...
  • @jason_kint Jason Kint on x
    What's incompatible to me is Facebook's former chief security officer being an expert able to position CA, literally a crime scene which Facebook helped cover-up and obfuscate, as an “overreaction.” I assume all press ask if he's still under a nondisparagement clause? https://twi…
  • @carljackmiller Carl Miller on x
    I forgot to say that they are now accepting Requests for Proposals for Facebook data for misinformation research. If you have a good use case (and I can think of a bazillion) submit a form here: https://docs.google.com/... https://twitter.com/...
  • @kantrowitz Alex Kantrowitz on x
    I'm sort of glad Facebook took so long — and acted with such hesitation — before releasing a massive batch of data. Beats the alternative. https://twitter.com/...
  • @kinggary Gary King on x
    Today, we released a privacy-protective Facebook dataset with 10 Trillion cell values https://socialscience.one/... Tomorrow, I'll give a webcast for TIER on differential privacy, a version of which we used with these data. Here're my slides & how to register: https://gking.harva…
  • @issielapowsky Issie Lapowsky on x
    @persily “Lots of politicians who have talked about privacy in a negative way at Facebook are also saying we want academic research to happen. They say those two things and don't realize they're not compatible.” - @alexstamos http://protocol.com/...
  • @2ramyakrishnan Ramya Krishnan on x
    1. Social Science One is a worthy effort, but more than two years after it was first announced it has failed to deliver. As we head into the 2020 elections, Facebook's platform remains alarmingly opaque. The public deserves more. https://www.wsj.com/...
  • @alemannoeu Alberto Alemanno on x
    #Facebook releases huge dataset enabling social scientists to study some of the most important questions of our time about the effects of social media on democracy and elections with information to which they have never before had access #disinformation https://socialscience.one/…
  • @byjacobward Jacob Ward on x
    The data is vast — comprising 38 million URLs shared more than 100 times each, and tons of associated data about how they moved across Facebook — but significantly limited. https://socialscience.one/...
  • @carljackmiller Carl Miller on x
    This is actually a really big deal. After two years of pressure, Facebook has begun to release a huge dataset of URLs shared across their platform. It's explicitly intended to help with research on misinformation. Paper: https://socialscience.one/... https://socialscience.one/...
  • @jason_kint Jason Kint on x
    ...always reminds me of 2016 timeline for Facebook's acquisition of Crowdtangle which had a dataset: 11/8 - 2016 election 11/10 - Zuck stupid statement 11/11 - Crowdtangle acquisition 11/19 - Zuckerberg tail between legs https://twitter.com/...
  • @jason_kint Jason Kint on x
    Good news all around. Although I am amused at the night and day between Facebook's approach to GDPR when mining commercially valuable personal data for its own interests vs protecting deidenitifed data from academics. https://twitter.com/... https://twitter.com/...
  • @daveyalba Davey Alba on x
    @random_walker The data contains ~38 million URLs; Social Science One said it “processed about an exabyte of raw data” from the Facebook platform which “produced a dataset that has about 10 trillion numbers in it.” SS1 has corrected the wording in their announcement https://socia…
  • @jessekblum Jesse Blumenthal on x
    Big news from @SocSciOne: academic researchers now have access to the URL full dataset (detailed info on ~38 million urls shared on Facebook) — one of the largest datasets ever available for social science research https://www.wsj.com/...
  • @profcarroll @profcarroll on x
    Good news in finding a solution balancing privacy and research but article doesn't point out that the dataset excludes 2016, which is arguably still a crime scene as Cambridge Analytica is probably still an active investigation per FOIA'd Bannon 302s. https://www.wsj.com/...