/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

← → days · ↑ ↓ browse · Enter similar · o open

AWS blames its hours-long Tuesday outage on network devices overloading, plans to revamp its status page to address complaints about updates and support cases

CNBC

Context & Ripple Effects

AWS had already experienced a North American outage that disrupted multiple online services in 2020, making incident communications part of the practical reliability experience for customers. The current response pairs a specific network-device diagnosis with a commitment to change how outage updates and support cases are handled.

First-order effects

  • AWS must revise its status-page and support-update process in response to complaints, while customers affected by the outage receive a stated technical explanation for the disruption.

Second-order effects

  • For AWS customers, outage communications become a more explicit part of evaluating the provider: support-case handling and status updates matter alongside restoration of service.

Third-order effects

  • Repeated outages attributed to different infrastructure layers point toward cloud reliability being judged not only by prevention, but by the clarity and usefulness of an operator's incident response.

The trend: Cloud providers are being pushed to treat incident communication and support workflows as core reliability capabilities, not ancillary customer-service functions.

Discussion

  • @quinnypig Corey Quinn on x
    The more I read the @awscloud analysis of the us-east-1 outage, the less confident I find myself in my understanding of failure modes and blast radii. It's not at all clear that AWS is fully aware of them, either.
  • @quinnypig Corey Quinn on x
    “The thing that broke was basically ‘the internal AWS network’ which is kind of important as it turns out. A lot of things fell over as a result. We have a lot of learning to do about this newly discovered behavior and its triggers. We are deeply sorry.”
  • @quinnypig Corey Quinn on x
    “It's always DNS, so they started there. It was not DNS, which is one for the record books. They then focused on moving traffic service by service, which eventually cleared the congestion issues.”
  • @quinnypig Corey Quinn on x
    “We made a change internally that caused a bunch of internal things to become extremely chatty, like AWS employees defending the company if someone says something even slightly unflattering on Twitter.”
  • @quinnypig Corey Quinn on x
    Because this is incredibly dense and technical, let me try to simplify it. I'm sure I will be condescendingly corrected if I get this wrong...
  • @quinnypig Corey Quinn on x
    “All of the chatty stuff made it really hard to understand what was going on because everything started behaving like CDK evangelists and never shutting the hell up for one goddamned second. As a result, engineers had to basically guess what was broken.”
  • @queenofcode Melissa Benua on x
    The writeup from the AWS us-east-1 outage is a super interesting read in managing complex systems and sometimes emergent behavior - and what to do when your observability systems are part of the problem and you're flying blind! https://twitter.com/...
  • @stevesi Steven Sinofsky on x
    Amazon Web Services explains outage and will make it easier to track future ones “The company also said it plans to revamp its status page.” https://www.cnbc.com/... // A sign of a maturing infrastructure is when the status page is a failure point and a redesign is needed.
  • @timperrett Timothy Perrett on x
    Summary of the AWS outage - decidedly lacking on details, as is typical for Amazon. Seems like a cascading failure along multiple vectors, with design assumptions about availability that didn't hold water. https://aws.amazon.com/...
  • @wcgallego Will Gallego on x
    AWS's Outage summary: “ As the impact to services during this event all stemmed from a single root cause...” https://aws.amazon.com/... https://twitter.com/...
  • @kennwhite Kenn White on x
    AWS with their post-mortem on the US-East outage on Tuesday. Critical internal network was overwhelmed, which cascaded into dependent control plane, monitoring, and customer support systems, dominoeing into RDS & EC2 deployments, then all hell broke loose. https://aws.amazon.com/…
  • @quinnypig Corey Quinn on x
    As @lizthegrey points out, “making changes to DNS to mitigate” appears to be homed through us-east-1; using @awscloud for DNS looks like it may be A Mistake You Should Avoid as a result. This is important, concerning, and more than a smidgen disappointing as a customer. https://t…
  • @lizthegrey Liz Fong-Jones on x
    interesting nugget: if you were multi-region but needed to push DNS to swap regions, you were SOL. “Route 53 APIs were impaired from 7:30 AM PST until 2:30 PM PST preventing customers from making changes to their DNS entries, but existing DNS entries were not impacted” https://tw…
  • @quinnypig Corey Quinn on x
    The @awscloud explanation of their outage earlier this week has been posted. https://aws.amazon.com/...
  • @tomkrazit Tom Krazit on x
    The AWS outage, explained: “At 7:30 AM PST (on Tuesday), an automated activity to scale capacity of one of the AWS services hosted in the main AWS network triggered an unexpected behavior from a large number of clients inside the internal network.” https://twitter.com/...