---
title: 'Every Major AWS Outage (And Why They Keep Happening)'
source: 'https://youtube.com/watch?v=6C14E9sQ_-w'
video_id: '6C14E9sQ_-w'
date: 2026-08-10
duration_sec: 972
---

# Every Major AWS Outage (And Why They Keep Happening)

> Source: [Every Major AWS Outage (And Why They Keep Happening)](https://youtube.com/watch?v=6C14E9sQ_-w)

## Summary

This video examines the six major AWS outages that occurred in the US-East-1 region over 15 years, detailing the root causes—from human error to hidden dependencies—and the cascading failures that followed. It highlights how the concentration of internet infrastructure in one region makes it a single point of failure for the global digital economy.

### Key Points

- **The 2021 Outage: A Global Disruption** [00:03] — A 14-hour outage in US-East-1 took down hospitals, airlines, Coinbase, Slack, Snapchat, Duolingo, Fortnite, and Alexa, affecting roughly a third of all websites. The cause was a timing issue between two software processes in a single data center in Ashburn, Virginia.
- **What is US-East-1?** [01:20] — AWS divides infrastructure into regions; US-East-1 is the first region in Northern Virginia. Foundational services like S3, EC2, DynamoDB, and Lambda were built there first, making it the default choice for developers, leading to 30-50% of internet traffic running through it.
- **2011 Outage: A Network Upgrade Mistake** [02:15] — During a routine network upgrade, an engineer routed traffic to the backup network instead of the primary, causing packet loss. This triggered a feedback loop where EBS volumes attempted to remirror, clogging the network further. Reddit, Foursquare, and Quora went down; full recovery took four days.
- **2012 Outage: Storm and Chaos Monkey** [04:20] — A thunderstorm caused power fluctuations and generator failures in the Ashburn data centers, affecting Netflix, Pinterest, and Instagram. Netflix stayed up because they had built Chaos Monkey to randomly terminate instances, forcing resilience.
- **2017 Outage: The S3 Typo** [06:03] — An S3 engineer made a typo while debugging a billing issue, removing the index subsystem that tracks object locations. This caused widespread failures (Trello, Slack) and the AWS status dashboard itself went down because it was hosted on S3.
- **2020 Thanksgiving Outage: Kinesis Cascade** [08:09] — A change to Kinesis caused it to consume excessive resources, leading to failures in IAM, Cognito, CloudWatch, and Route 53. Engineers couldn't log in to fix the problem because IAM was down, creating a cascade that took hours to resolve.
- **2021 Outage: Amazon's Own Cloud Fails** [10:06] — Lambda and other services in US-East-1 went down, taking Alexa, Ring, and iRobot Roomba offline. Amazon's own fulfillment centers were impacted as scanners and Flex driver apps failed, showing Amazon's dependence on its own cloud.
- **2023 Outage: A DNS Race Condition** [12:14] — A rare timing condition in automation deleted the DNS record for DynamoDB's regional endpoint, making it unreachable. This cascaded to EC2 and other services, taking down Netflix, Slack, Coinbase, and more for 14 hours.

### Conclusion

The video concludes that US-East-1's concentration of internet infrastructure makes it a critical single point of failure. Despite post-mortems and new safeguards, the complexity of the system means failures will continue to occur in unpredictable ways, affecting the entire digital world.

## Transcript

hospitals across the United States try to pull up patient records. Nothing loads. Airline gate agents try to issue boarding pass. The system is frozen. Coinbase goes down. Slack goes down. Snapchat goes down. Dualingo goes down.
Fortnite goes down. People try to ask Alexa what's happening. Alexa doesn't answer. For the next 14 hours, a significant portion of the modern &gt;&gt; Amazon, &gt;&gt; Amazon, Amazon.com suffered a major
outage today. The US, the major cloud computing services went offline. &gt;&gt; Roughly a third of all websites on the internet use it. It's difficult to believe how many use it. Uh so this outage immediately sent shock waves
around the world. &gt;&gt; There's no cyber attack, no natural disaster, no ransomware. The cause is a timing issue between two software processes inside a data center in Ashurn, Virginia. One data center, one
region, one company's infrastructure, and the world holds its breath. Here's the thing, though. This wasn't new. This was the sixth time. Before we get into it, what even is US- East1? AWS divides its infrastructure into regions,
geographical areas where they've built data centers. US East one is their Northern Virginia region. And it's not just another region. It was the first region. Every foundational AWS service like S3, EC2, Dynamo DB, and Lambda,
they were built and tested there first, which means it became the default. developers building in the early 2010s picked US- East1 because it was where everything was. And once enough people were on it, it became the region you had
to use. Today, estimates put somewhere between 30 and 50% of all internet between 30 and 50% of all internet traffic running through that one region, one data center cluster, one state, Virginia. Keep that in mind.
old. The cloud is still a new idea. Most companies still run their own servers, but an early generation of startups like Reddit, Forsquare, Quora, and Hootuite have gone allin on this new thing called the cloud. That morning, the AWS team
begins a routine network upgrade in US East1. The goal is to move traffic to higher capacity network connections, a standard procedure. The engineer standard procedure. The engineer executing the change makes a mistake.
Instead of routing traffic to the primary network and shifting the old connections to backup, they do it in reverse. The backup network gets the primary traffic, but the backup wasn't built for that load. It starts dropping
packets. And here's where things get interesting. EBS elastic block store the virtual hard drives that EC2 instances run on begins trying to remirror itself.
Thousands of volumes all detecting that their data might be at risk. All simultaneously trying to create backup copies of themselves. Normally this is fine. Normally the reiring takes spare network capacity. There was no spare
capacity. The network was already overwhelmed. So the re mirroring clogs the network further which causes more EBS volumes to think they're losing data which causes more reiring which clogs the network more a feedback loop a
spiral. Reddit goes down forsquare goes down. Quora goes down. &gt;&gt; AWS the major cloud computing services went offline early this morning which brought down major websites such as Reddit forsquare and giant bomb. AWS
engineers work around the clock. Full recovery takes four days. The cloud was three years old and it had already proven something important. When it proven something important. When it breaks in ways nobody predicted.
thunderstorms tears through Northern Virginia. The kind of storm with 80 mph winds and rapid fire lightning. &gt;&gt; Racho. It left a big mark on our region. The damage impacting thousands of people in our area. It
&gt;&gt; hits the Ashberry data centers. Emergency generators kick in. Then the transfer switches start malfunctioning. There are the hardware that manages the transition from grid power to generator power. Some of them fail on the way up.
Power fluctuate. Servers go down. Netflix, Pinterest, Instagram, Heroku, all affected. But here's the interesting part. Netflix stayed up. Their competitors were scrambling. Netflix engineers spent two minutes watching
their dashboards, saw that AWS was degraded, and their systems automatically shifted traffic away from the affected zone. They had spent the previous two years building something called chaos monkey software that
randomly terminated their own instances in production on purpose to force their systems to become resilient to exactly this kind of failure. They had literally
rehearsed for this disaster. So, while Instagram went dark and Pinterest went down, Netflix kept streaming and engineers across the internet learned a lesson about what it means to build for failure. The storm lasted a few hours.
failure. The storm lasted a few hours. The outage lasted about the same. says it's experiencing issues with its cloud-based computing service, which is
used by nearly a million customers. &gt;&gt; No storm this time. No network upgrade &gt;&gt; No storm this time. No network upgrade gone wrong. This time, it's a typo. An S3 engineer is debugging a billing issue. S3 or simple service is the
backbone of the internet. Websites, apps, databases, enormous amounts of the web either live on S3 or depend on it. The engineer follows an established playbook. They run a command to remove a small number of servers from one of S3's
internal subsystems, but they type in the wrong number. Instead of removing a small number of servers, they remove a large one. Specifically, they removed the index subsystem, the service that tracks the location of every single
object in S3. Every photo, every file, every database backup. Without the index, S3 can't find anything. Requests start failing. Trello goes down. Slack goes down. Thousands of apps relying on
goes down. Thousands of apps relying on S3 go down. The AWS team tries to restart the index subsystem. And here's the problem. S3 has grown enormously the problem. S3 has grown enormously since it launched in 2006. The index
subsystem hasn't needed a full restart in years. Nobody knows how long it's in years. Nobody knows how long it's going to take. It takes hours. The crawl bar moves slowly across the screen as engineers wait. At some point, someone
tries to check the AWS status dashboard to see if there's an update. The status dashboard is hosted on S3. It won't load. Amazon is using Amazon to check if
Amazon is down. The dashboard shows all green, everything is fine, services are operational because the dashboard itself can't update its own status. 4 hours
later, S3 recovers. The billing issue gets fixed. The engineer presumably has gets fixed. The engineer presumably has a very long walk home. Amazon quietly adds a note to their postmortem that they're going to move the status
they're going to move the status dashboard off of S3. Thanksgiving, AWS makes a change to Kinesis data streams. Kinesis is a
service for processing real-time data. It's not as well known as S3 or EC2, but it turns out it's quietly woven into the foundation of how AWS works internally. The change causes Kinesis to start consuming far more resources than
expected on the servers running the front end of the service. Kinesis starts to struggle and then things get strained. IM starts having problems. IM is AWS's authentication service. It's the thing that verifies who you are and
whether you're allowed to do something. Cognto starts having problems. That's the service that lets applications handle user signin. Cloudatch starts
having problems. That's the monitoring service, the one engineers use to figure out what's wrong. Route 53 health checks start failing. Autocaling stops working.
All of them had dependencies on Kinesis that nobody had fully mapped out. When Kinesis slowed down, it dragged everything connected to it. Engineers try to log into the AWS console to diagnose the situation. They can't log
diagnose the situation. They can't log in. AM is down. The tool you need to fix the problem is broken by the problem. It takes hours to untangle the cascade. services recover slowly, one by one through the night. On Thanksgiving,
people notice things are working again. But at the time, almost no one has any But at the time, almost no one has any idea what happened.
&gt;&gt; If you're having some trouble this morning using Amazon's web services, you &gt;&gt; Amazon Web Services suffered a major outage today, disrupting access to many popular websites. It took down services including Prime Music and video Alexa
and its Ring Smartome systems. &gt;&gt; Something is wrong with AWS Lambda and &gt;&gt; Something is wrong with AWS Lambda and several other services in US East1. Apps start throwing errors. &gt;&gt; People trying to use Instacart, Venmo,
Kindle, Roku, and Disney Plus have reported issues. &gt;&gt; Nothing catastrophic at first. Then it gets worse. But this outage is different from the others because of who goes down. Alexa stops responding. Tens of
millions of people say Alexa and get silence. Ring cameras go offline. Doorbell cameras, security cameras, all offline. Iroot Roomba vacuums scheduled
to clean while their owners are at work stop midfloor and refuse to move. And stop midfloor and refuse to move. And then it gets worse. Amazon itself starts to break. Amazon Flex drivers, who are the gig workers who deliver Amazon
packages in their own cars, open their apps to start their shifts. The apps won't load. They can't scan packages. They can't start routes. Inside Amazon fulfillment centers, the handheld scanners that workers use to track
inventory stop connecting. The company that owns the cloud is being taken down by its own cloud. For a few hours, Amazon's ability to fulfill orders, which you know is its entire reason for existing, is being impacted by AWS going
down. The same infrastructure Amazon built to sell to the world is the same infrastructure Amazon depends on to function. The outage eventually resolves, but Amazon doesn't go into great detail about the root cause. The
great detail about the root cause. The silence is its own kind of answer.
an automated system is doing what it always does, managing DNS records for &gt;&gt; This outage so massive, in fact, it's really hard to pinpoint an industry that wasn't affected by this. &gt;&gt; Dynamo DB is Amazon's flagship database
service. But it's not just a database people use. Internally, dozens of AWS services use Dynamob as their coordination layer, storing state, tracking configurations, managing metadata. It's the connective tissue of
the platform. The automation manages hundreds of thousands of DNS records, constantly updating them to point to the endpoint. Somewhere in that automation, two redundant components encounter a rare timing condition, erased. They
collide in a way that's never happened before. The automation deletes the DNS record for the Dynamo DB regional endpoint. Not the data, not the servers,
just the address that every AWS service uses to find Dynamo DB. In seconds, uses to find Dynamo DB. In seconds, Dynamob becomes unreachable. Not broken, just unfindable. And then the cascade begins. EC2 which are the virtual
machines that power the entire cloud stores its operational data in Dynamob. With Dynamo DB unreachable EC2's orchestration system stalled new instances can't launch. Existing
instances can't be managed. Other services that depend on those services start failing too. Netflix goes down. Slack goes down. Coinbase goes down. Expedia goes down. Hospitals can't access records. Airlines can't issue
boarding passes. Banks go dark. Engineering teams around the world open their AWS management consoles to diagnose the problem. The consoles are either unreachable or showing stale data. The fixed tools are broken by the
things being fixed. Again, AWS engineers work to recreate the missing DNS record, but the systems that should automatically recover are themselves dependent on Dynamo DB. They can't recover. Restoring everything carefully
in the right order, validating each step. This takes 14 hours. step. This takes 14 hours. 14 hours. Six outages. 15 years. One region. A typo, a storm, a mistyped command, a holiday eve configuration
command, a holiday eve configuration change, a dependency nobody mapped, a race condition in automation, six completely different causes, six completely different failure modes and yet same region, same impact, same
yet same region, same impact, same story. The uncomfortable truth about US East one is this. We have known for 15 years that it is too important to fail. years that it is too important to fail. And we have watched it fail anyway
repeatedly for reasons ranging from bad luck to human error to hidden complexity that nobody und understood until it broke. Every time Amazon publishes a post-mortem. Every time they promise new safeguards. Every time the next outage
finds a different crack to follow through. Because systems this complex don't fail in ways you predict. They fail in the ways you didn't think to protect against. And because 30 to 50% of the internet chose the same region,
the same data center cluster, the same slice of land in Northern Virginia. Every time a single thing goes wrong there, the whole world finds out about
there, the whole world finds out about it.
