AWS blames its hours-long Tuesday outage on network devices overloading, plans to revamp its status page to address complaints about updates and support cases
CNBC
Context & Ripple Effects
AWS had already experienced a North American outage that disrupted multiple online services in 2020, making incident communications part of the practical reliability experience for customers. The current response pairs a specific network-device diagnosis with a commitment to change how outage updates and support cases are handled.
First-order effects
AWS must revise its status-page and support-update process in response to complaints, while customers affected by the outage receive a stated technical explanation for the disruption.
Second-order effects
For AWS customers, outage communications become a more explicit part of evaluating the provider: support-case handling and status updates matter alongside restoration of service.
Third-order effects
Repeated outages attributed to different infrastructure layers point toward cloud reliability being judged not only by prevention, but by the clarity and usefulness of an operator's incident response.
The trend: Cloud providers are being pushed to treat incident communication and support workflows as core reliability capabilities, not ancillary customer-service functions.
The more I read the @awscloud analysis of the us-east-1 outage, the less confident I find myself in my understanding of failure modes and blast radii. It's not at all clear that AWS is fully aware of them, either.
“The thing that broke was basically ‘the internal AWS network’ which is kind of important as it turns out. A lot of things fell over as a result. We have a lot of learning to do about this newly discovered behavior and its triggers. We are deeply sorry.”
“It's always DNS, so they started there. It was not DNS, which is one for the record books. They then focused on moving traffic service by service, which eventually cleared the congestion issues.”
“We made a change internally that caused a bunch of internal things to become extremely chatty, like AWS employees defending the company if someone says something even slightly unflattering on Twitter.”
“All of the chatty stuff made it really hard to understand what was going on because everything started behaving like CDK evangelists and never shutting the hell up for one goddamned second. As a result, engineers had to basically guess what was broken.”
The writeup from the AWS us-east-1 outage is a super interesting read in managing complex systems and sometimes emergent behavior - and what to do when your observability systems are part of the problem and you're flying blind! https://twitter.com/...
Amazon Web Services explains outage and will make it easier to track future ones “The company also said it plans to revamp its status page.” https://www.cnbc.com/... // A sign of a maturing infrastructure is when the status page is a failure point and a redesign is needed.
Summary of the AWS outage - decidedly lacking on details, as is typical for Amazon. Seems like a cascading failure along multiple vectors, with design assumptions about availability that didn't hold water. https://aws.amazon.com/...
AWS's Outage summary: “ As the impact to services during this event all stemmed from a single root cause...” https://aws.amazon.com/... https://twitter.com/...
AWS with their post-mortem on the US-East outage on Tuesday. Critical internal network was overwhelmed, which cascaded into dependent control plane, monitoring, and customer support systems, dominoeing into RDS & EC2 deployments, then all hell broke loose. https://aws.amazon.com/…
As @lizthegrey points out, “making changes to DNS to mitigate” appears to be homed through us-east-1; using @awscloud for DNS looks like it may be A Mistake You Should Avoid as a result. This is important, concerning, and more than a smidgen disappointing as a customer. https://t…
interesting nugget: if you were multi-region but needed to push DNS to swap regions, you were SOL. “Route 53 APIs were impaired from 7:30 AM PST until 2:30 PM PST preventing customers from making changes to their DNS entries, but existing DNS entries were not impacted” https://tw…
The AWS outage, explained: “At 7:30 AM PST (on Tuesday), an automated activity to scale capacity of one of the AWS services hosted in the main AWS network triggered an unexpected behavior from a large number of clients inside the internal network.” https://twitter.com/...