AWS blames its hours-long Tuesday outage on network devices overloading, plans to revamp its status page to address complaints about updates and support cases
- A major Amazon Web Services outage on Tuesday started after network devices got overloaded, the company said on Friday.
CNBC
Related Coverage
- Summary of the AWS Service Event in the Northern Virginia (US-EAST-1) Region Amazon Web Services, Inc.
- Amazon explains the cause behind Tuesday's massive AWS outage BleepingComputer · Sergiu Gatlan
- Why Did Amazon's AWS Crash and Can It Happen Again? TheStreet · Kirk O'Neil
- Amazon explains outage that took out a large chunk of the internet Engadget · Jon Fingas
- Here's Why a Vital Amazon Web Services Region Went Down on Dec. 7 PCMag · Nathaniel Mott
- Amazon Web Services says overwhelmed network devices triggered outage The Verge · Emma Roth
- Why everything from Netflix to Nintendo goes offline when Amazon's servers have issues Insider · Ben Gilbert
- Amazon Says ‘Unexpected Behavior’ Caused Huge Cloud Outage Bloomberg
Discussion
-
@timperrett
Timothy Perrett
on x
Summary of the AWS outage - decidedly lacking on details, as is typical for Amazon. Seems like a cascading failure along multiple vectors, with design assumptions about availability that didn't hold water. https://aws.amazon.com/...
-
@stevesi
Steven Sinofsky
on x
Amazon Web Services explains outage and will make it easier to track future ones “The company also said it plans to revamp its status page.” https://www.cnbc.com/... // A sign of a maturing infrastructure is when the status page is a failure point and a redesign is needed.
-
@queenofcode
Melissa Benua
on x
The writeup from the AWS us-east-1 outage is a super interesting read in managing complex systems and sometimes emergent behavior - and what to do when your observability systems are part of the problem and you're flying blind! https://twitter.com/...
-
@wcgallego
Will Gallego
on x
AWS's Outage summary: “ As the impact to services during this event all stemmed from a single root cause...” https://aws.amazon.com/... https://twitter.com/...
-
@kennwhite
Kenn White
on x
AWS with their post-mortem on the US-East outage on Tuesday. Critical internal network was overwhelmed, which cascaded into dependent control plane, monitoring, and customer support systems, dominoeing into RDS & EC2 deployments, then all hell broke loose. https://aws.amazon.com/…
-
@quinnypig
Corey Quinn
on x
As @lizthegrey points out, “making changes to DNS to mitigate” appears to be homed through us-east-1; using @awscloud for DNS looks like it may be A Mistake You Should Avoid as a result. This is important, concerning, and more than a smidgen disappointing as a customer. https://t…
-
@lizthegrey
Liz Fong-Jones
on x
interesting nugget: if you were multi-region but needed to push DNS to swap regions, you were SOL. “Route 53 APIs were impaired from 7:30 AM PST until 2:30 PM PST preventing customers from making changes to their DNS entries, but existing DNS entries were not impacted” https://tw…
-
@quinnypig
Corey Quinn
on x
The @awscloud explanation of their outage earlier this week has been posted. https://aws.amazon.com/...
-
@tomkrazit
Tom Krazit
on x
The AWS outage, explained: “At 7:30 AM PST (on Tuesday), an automated activity to scale capacity of one of the AWS services hosted in the main AWS network triggered an unexpected behavior from a large number of clients inside the internal network.” https://twitter.com/...