Google to Host Terabytes of Open-Source Scientific Data
Sources at Google have disclosed that the humble domain, http://research.google.com, will soon provide a home for terabytes of open-source scientific datasets. The storage will be free to scientists and access to the data will be free for all.
Context & Ripple Effects
This lands mid-arc: in November 2007 the Wall Street Journal reported Google's plans to build a service for storing users' personal data, framing storage itself as a product surface. Wired's report extends that logic into science — unnamed sources say the modest research.google.com domain is about to absorb terabytes of open-source scientific datasets, with storage free to scientists and access free to everyone.
The claim is sourced rather than officially announced (Mashable picked it up the same day as a 'social science data network'), so nothing here is confirmed pricing or a shipped product — but the direction fits Google's pattern of converting spare infrastructure capacity into free services that pull new workloads onto its platform.
First-order effects
- If the reported plan ships, working scientists get terabyte-scale dataset hosting at zero cost, removing the per-project server budget line that typically gates how much raw data a lab can publish.
- research.google.com stops being a static publications page and becomes a data distribution point, making Google the delivery layer between funded research projects and the public.
Second-order effects
- A credible free offer at this scale sets a zero-price benchmark that paid scientific data-hosting vendors and university IT departments provisioning their own repositories would have to match or justify against.
- Datasets sitting on Google's infrastructure become reachable through the same crawl-and-index machinery that powers web search, tilting discovery traffic — and the citation attention that follows — toward whatever is hosted there.
Third-order effects
- If storage for scientific data trends toward free, the scarce layer moves up the stack to curation and discovery — whoever indexes the world's datasets controls how science finds them, a position Google is structurally best-placed to occupy.
- Companies providing public-goods infrastructure for publicly funded research raises a governance question the field hasn't answered: who underwrites and maintains open scientific data when a single platform operator holds it.
The trend: Cloud-scale operators are absorbing scientific data hosting as a free utility, shifting the value in research from storing datasets to indexing and curating them.