Parts of Amazon Web Services suffer an outage
Amazon web services are having trouble this evening and in the process are taking down some major sites. Among sites being impacted are Quora and HipChat. In addition, the Amazon outage has had an impact on Heroku, a division of Salesforce.
Context & Ripple Effects
This is the second widely felt AWS disruption in under a year, following the August 2011 outage that Amazon said it had resolved — a pattern rather than a one-off. The June 15 incident again shows how much of the consumer web sits on Amazon's infrastructure: Quora and HipChat go dark directly, while Heroku, which Salesforce acquired and runs atop AWS, passes the failure through to every app deployed on its platform.
The Heroku angle gives the story an enterprise edge: days earlier Salesforce was reported closing in on the Buddy Media acquisition and sharpening its social offerings, so its application-platform business is now visibly dependent on a competitor's uptime.
First-order effects
- Quora and HipChat are offline this evening with no fallback of their own, and Heroku-hosted applications lose service because their underlying AWS capacity is impaired.
- Salesforce inherits the reputational hit through Heroku — customers paying for a managed platform experience an outage branded Salesforce but caused by Amazon.
Second-order effects
- Startups building on Heroku or raw EC2 face pressure to engineer redundancy across regions or providers, raising their infrastructure costs to insure against a vendor they don't control.
- Rival platforms and managed-hosting competitors gain a sales argument built on Amazon's recurring failures, just as enterprises weigh commitments to AWS-based stacks.
Third-order effects
- If major AWS incidents keep recurring roughly annually, concentration risk becomes a board-level procurement question: the industry drifts toward multi-region and multi-provider architectures, and uptime SLAs start pricing in the platform layer beneath the platform.
- Regulators and large enterprise buyers begin treating a handful of cloud operators as systemic infrastructure, where one provider's fault domain spans thousands of unrelated sites at once.
The trend: Cloud computing's economics are concentrating the web onto a few providers, so each AWS fault now cascades into an internet-wide event — making resilience architecture, not capacity, the binding constraint.