Database gods bitch about mapreduce
This is what disruption sounds like. — This rant by major database guys against mapreduce is pretty telling. — (You can read a good rebuttal here, and discussion on ycomb.) — The thing that disrupts you is always uglier and worse in some way.
Context & Ripple Effects
The rant lands days after two data points made MapReduce impossible for the database establishment to ignore: a petabyte-per-day production run reported earlier this month, and the ACM's formal publication of the MapReduce paper on January 7, which gave an academic stamp of legitimacy to what had been an internal web-company technique.
It also extends a fight that has been running since at least 2006, when Business Week profiled startups taking on the database giants. Skrenta's framing — that disrupted incumbents always describe their disruptor as ugly and worse — turns the experts' technical complaints into a case study in incumbent psychology.
First-order effects
- Prominent database authorities have staked their public credibility against MapReduce, forcing every practitioner evaluating large-cluster data processing to weigh relational orthodoxy against a technique already running at petabyte scale inside web companies.
Second-order effects
- The rebuttal and the Y Combinator discussion show the builder community rallying around MapReduce rather than the critics, pressuring database vendors to answer scale-out workloads on their merits instead of dismissing them.
Third-order effects
- If the pattern holds, data processing splits into two defended camps — systems optimized for structured queries versus brute-force cluster scans — and the eventual winners are likely hybrids that absorb the other side's criticisms, much as the incumbents' own tools once absorbed features from the systems they displaced.
The trend: Large-scale data processing is polarizing into relational-database defenders and MapReduce-style cluster computing, with each camp's critiques shaping the other's roadmap.