/
Navigation
Chronicles
Browse all articles
Explore
Semantic exploration
Research
Entity momentum
Nexus
Correlations & relationships
Story Arc
Topic evolution
Drift Map
Semantic trajectory animation
Posts
Analysis & commentary
Pulse API
Tech news intelligence API
Browse
Entities
Companies, people, products, technologies
Domains
Browse by publication source
Handles
Browse by social media handle
Detection
Concept Search
Semantic similarity search
High Impact Stories
Top coverage by position
Sentiment Analysis
Positive/negative coverage
Anomaly Detection
Unusual coverage patterns
Analysis
Rivalry Report
Compare two entities head-to-head
Semantic Pivots
Narrative discontinuities
Crisis Response
Event recovery patterns
Connected
Search: /
Command: ⌘K
Embeddings: large
TEXXR

Chronicles

The story behind the story

days · browse · Enter similar · o open

AWS blames its hours-long Tuesday outage on network devices overloading, plans to revamp its status page to address complaints about updates and support cases

- A major Amazon Web Services outage on Tuesday started after network devices got overloaded, the company said on Friday.

CNBC

Discussion

  • @timperrett Timothy Perrett on x
    Summary of the AWS outage - decidedly lacking on details, as is typical for Amazon. Seems like a cascading failure along multiple vectors, with design assumptions about availability that didn't hold water. https://aws.amazon.com/...
  • @stevesi Steven Sinofsky on x
    Amazon Web Services explains outage and will make it easier to track future ones “The company also said it plans to revamp its status page.” https://www.cnbc.com/... // A sign of a maturing infrastructure is when the status page is a failure point and a redesign is needed.
  • @queenofcode Melissa Benua on x
    The writeup from the AWS us-east-1 outage is a super interesting read in managing complex systems and sometimes emergent behavior - and what to do when your observability systems are part of the problem and you're flying blind! https://twitter.com/...
  • @wcgallego Will Gallego on x
    AWS's Outage summary: “ As the impact to services during this event all stemmed from a single root cause...” https://aws.amazon.com/... https://twitter.com/...
  • @kennwhite Kenn White on x
    AWS with their post-mortem on the US-East outage on Tuesday. Critical internal network was overwhelmed, which cascaded into dependent control plane, monitoring, and customer support systems, dominoeing into RDS & EC2 deployments, then all hell broke loose. https://aws.amazon.com/…
  • @quinnypig Corey Quinn on x
    As @lizthegrey points out, “making changes to DNS to mitigate” appears to be homed through us-east-1; using @awscloud for DNS looks like it may be A Mistake You Should Avoid as a result. This is important, concerning, and more than a smidgen disappointing as a customer. https://t…
  • @lizthegrey Liz Fong-Jones on x
    interesting nugget: if you were multi-region but needed to push DNS to swap regions, you were SOL. “Route 53 APIs were impaired from 7:30 AM PST until 2:30 PM PST preventing customers from making changes to their DNS entries, but existing DNS entries were not impacted” https://tw…
  • @quinnypig Corey Quinn on x
    The @awscloud explanation of their outage earlier this week has been posted. https://aws.amazon.com/...
  • @tomkrazit Tom Krazit on x
    The AWS outage, explained: “At 7:30 AM PST (on Tuesday), an automated activity to scale capacity of one of the AWS services hosted in the main AWS network triggered an unexpected behavior from a large number of clients inside the internal network.” https://twitter.com/...