Google processes over 20 petabytes of data per day
Google currently processes over 20 petabytes of data per day through an average of 100,000 MapReduce jobs spread across its massive computing clusters. The average MapReduce job ran across approximately 400 machines in September 2007 …
Context & Ripple Effects
This lands three days after Geeking with Greg first flagged the figure and one day after Jeff Dean and Sanjay Ghemawat's MapReduce paper ran in the ACM's Communications, so the number arrives with its architecture already documented: over 20 petabytes a day through roughly 100,000 jobs, each spanning about 400 machines as of September 2007.
What makes it more than a vanity stat is that Google has published the recipe alongside the measurement — rivals can read exactly how the scale is achieved, which converts an internal capability into an industry template.
First-order effects
- Google's cluster economics are now public: any competitor benchmarking against search-scale indexing knows the job size (400 machines), the throughput (20 PB/day), and the software layer that delivers it.
Second-order effects
- With the framework peer-reviewed and freely described, other web companies face pressure either to adopt MapReduce-style batch processing on commodity clusters or to explain why their own pipelines don't scale the same way — and hardware vendors gain a demand signal for dense, fault-tolerant commodity servers rather than exotic machines.
Third-order effects
- If the pattern holds, large-scale data processing migrates from specialized supercomputing to fleets of cheap machines coordinated by software — making the framework, not the iron, the scarce asset, and pushing the bottleneck toward moving data between nodes.
The trend: Web-scale computing is consolidating around published distributed-processing frameworks running on commodity clusters, with Google's disclosed numbers setting the reference point others measure against.