AWS’s us-east-1 region went dark again on May 7, 2026. A cooling failure inside a single data-center hall cut power to servers in the use1-az4 availability zone, and full recovery stretched to roughly 28 hours. Coinbase, FanDuel, and CME Group all took direct hits. It was the third major us-east-1 failure since December 2021, and the second in just seven months.
The pattern is hard to ignore. AWS still runs the world’s largest cloud computing business, with 28% of global infrastructure spending as of the fourth quarter of 2025, according to Synergy Research Group. But its oldest, busiest region keeps breaking in new ways: a scaling bug in 2021, a DNS race condition in October 2025, and now a physical cooling failure. Each incident renews the same question for engineering teams and regulators alike. How much of the internet can safely depend on one region, one zone, one provider?
What Happened: Inside AWS’s 28-Hour us-east-1 Outage
AWS’s own status updates trace the incident to a physical failure, not a software bug. Multiple cooling units failed inside a single data-center hall serving use1-az4 late on May 7. As the room heated past safe operating thresholds, servers shut themselves down automatically to protect the hardware, a built-in safety response that cut power to racks across the affected zone. “EC2 instances and EBS volumes hosted on impacted hardware are affected by the loss of power during the thermal event,” AWS said in a status update reported by The Register.
Restoring service wasn’t a simple reboot. Cooling had to drop back below thermal thresholds before engineers could safely re-power the hardware, which is why the fix took most of a day rather than a few hours.
Hour by Hour
- 5:25 PM PDT, May 7 (00:25 UTC, May 8): AWS reports rising error rates as cooling units fail in use1-az4.
- Overnight: Automatic shutdowns cut power to affected racks. EC2 instances and EBS volumes go dark, and dependent services including ELB, EKS, Redshift, MSK, ElastiCache, OpenSearch, SageMaker, IoT Core, and NAT Gateway start degrading.
- 1:50 PM PDT, May 8 (about 20 hours in): Cooling stabilizes to pre-incident levels. Most EC2 instances and EBS volumes recover.
- Through Friday evening: Data-heavy services finish restoring, pushing full recovery to roughly 28 hours.
Which Services and Companies Went Dark
The direct damage stayed inside one availability zone, but use1-az4 hosts enough concentrated capacity that the outage still reached major consumer platforms. Coinbase’s trading systems went offline for about seven hours, with core exchange functions down for more than five, according to the company’s own postmortem. FanDuel couldn’t process cash-outs during a live Lakers-Thunder Western Conference playoff game. CME Group’s institutional trading tool, CME Direct, threw error screens at users. The humanitarian data-collection platform KoboToolbox went dark starting at 00:32 UTC, though its EU instance stayed up.
On the infrastructure side, AWS listed EC2 and EBS as directly affected, with ELB, EKS, Redshift, MSK, ElastiCache, OpenSearch, SageMaker, IoT Core, and NAT Gateway all degrading as a result. That’s a wide blast radius for a single-zone event, and it shows how many managed services sit on top of raw EC2 and EBS capacity without customers realizing it until something breaks.
Why AWS’s us-east-1 Region Keeps Failing
us-east-1, based in Northern Virginia, is AWS’s oldest and largest region. It’s also the default choice for new accounts, the control plane for several global services, and home to a huge share of startups that never migrate elsewhere once they launch there. That concentration makes every us-east-1 failure louder than an outage anywhere else. One LinkedIn post from IT professional Thennarasu Natesan, tracking AWS status history, put it in blunt numbers: “us-east-1: 3x more downtime than other AWS regions in 2025.”
The region has now failed in three structurally different ways since 2021: a scaling bug that overloaded internal networking gear in December 2021, a DNS automation race condition that erased a critical endpoint record in October 2025, and a physical cooling failure in May 2026. Different root causes, same result. Engineers who built redundancy for one failure mode got caught by another.
A Timeline of Major Cloud Outages, 2021-2026
The May outage didn’t happen in isolation. AWS, Azure, and Google Cloud have all logged multi-hour failures since 2021, though AWS’s incidents cluster more heavily in a single region than its rivals’ do. The comparison below draws on AWS and Azure status histories, Google Cloud’s incident reports, and outage-cost researcher OutageCost.com’s tracking database.
| Date | Provider | Region / Service | Duration | Root Cause | Estimated Cost |
|---|---|---|---|---|---|
| Dec 7, 2021 | AWS | us-east-1 (EC2, ECS, Lambda, SNS) | ~7 hours | Automated scaling activity overloaded internal network devices | $150M+ est. |
| Jan 8-11, 2025 | Azure | East US 2 (AZ01) | ~50h 13m | Networking configuration change | Not quantified |
| Jun 12, 2025 | Google Cloud | Service Control / IAM (multi-region) | ~3h (7h in us-central1) | Invalid automated quota policy update | $50M+ est. |
| Oct 20, 2025 | AWS | us-east-1 (DynamoDB DNS) | ~15 hours | DNS automation race condition | $38M-$581M insured (CyberCube) |
| Oct 29-30, 2025 | Azure | Front Door (global) | ~8.5 hours | Inadvertent config change bypassed validation | Not yet quantified |
| Feb 2026 | Azure | Storage / VM extension packages (multi-region) | ~12 hours | Remediation workflow disabled anonymous read access | Not quantified |
| May 7-8, 2026 | AWS | us-east-1 (use1-az4) | ~28 hours | Data-center cooling / thermal failure | $25M+ est. (modeled) |
Two older incidents anchor the far end of that record. AWS’s S3 outage in February 2017 ran about four hours and disrupted a large share of the internet, and the December 2021 us-east-1 failure carries a similar estimate. OutageCost.com models both at $150 million or more in aggregate customer impact. The May 2026 event hasn’t produced an independent estimate that detailed, mostly because its damage stayed confined to a single zone instead of an entire region.
The Dollar Cost of Cloud Downtime
Cloud providers don’t publish customer-impact figures, so most cost estimates come from third-party analysts modeling revenue at risk. For the October 2025 DynamoDB outage, insurance-analytics firm CyberCube put preliminary insured losses at $38 million to $581 million, in an event it internally nicknamed “Amazonk,” and estimated that roughly 70,000 organizations were affected, more than 2,000 of them large enterprises. CyberCube still rated the overall insurance impact as moderate: a low-to-mid single-digit hit to loss ratios industry-wide.
The May 2026 event is harder to price. Its single-zone blast radius was narrower than October’s region-wide failure, and no insurer has published a comparable figure yet. OutageCost.com’s own modeling places aggregate customer losses in the tens of millions of dollars, with a working estimate around $25 million, but frames that explicitly as a modeled figure rather than a disclosed one.
Why SLA Credits Don’t Cover the Real Loss
Whatever the true number turns out to be, AWS’s compensation won’t come close to covering it. AWS’s EC2 service-level agreement pays back 10% to 30% of the affected service’s monthly bill, never a share of the customer’s lost revenue. Run $100,000 a day of AWS-dependent revenue, and a 15-hour regional outage works out to roughly $62,500 of exposure before any failover kicks in, using OutageCost.com’s downtime-cost model. A credit worth 10% to 30% of a monthly AWS bill rarely gets close to covering that kind of loss.
AWS vs Azure vs Google Cloud: Market Share and Reliability Record
AWS still leads the cloud infrastructure market by a wide margin. It held 28% of worldwide cloud infrastructure spending in the fourth quarter of 2025, ahead of Microsoft Azure’s 21% and Google Cloud’s 14%, according to Synergy Research Group. Together, the big three control 63% of a market Synergy sized at $119 billion for the quarter, up roughly $29 billion year over year, a growth rate near 30%.
| Provider | Q4 2025 Market Share | Documented Major Incidents | Tracking Window | SLA Credit Range | Worst Documented Incident |
|---|---|---|---|---|---|
| AWS | 28% | 9 | Since 2012 | 10%-30% of monthly fee | Dec 2021 & Feb 2017 (~$150M+ each) |
| Microsoft Azure | 21% | 9 | Since 2018 | 25%-100% of monthly fee | Oct 2025 Front Door (~8.5h, global) |
| Google Cloud | 14% | 7 | Since 2016 | 10%-50% of monthly fee | Jun 2025 IAM outage (dozens of products) |
Market leadership and reliability leadership aren’t the same thing. Azure pays the most generous outage credits, up to 100% of a service’s monthly fee versus AWS’s 30% cap, while Google Cloud’s failures are rarer but spread faster once they start. A single control-plane bug can take down dozens of Google Cloud products worldwide within minutes, which is exactly what happened in June 2025.
What Engineers Are Saying About us-east-1
The gap between AWS’s marketing and its us-east-1 track record has become a running complaint among the engineers who operate on it daily. “AWS is generally considered to have 5 9s of reliability, meaning roughly 5 minutes of downtime per year,” wrote technology professional Elijah Mirecki, a benchmark us-east-1 hasn’t come close to meeting across three major incidents since 2021.
Others are less diplomatic about it. A commenter on Hacker News summed up the practitioner consensus during a discussion of a recent us-east-1 incident: “AWS us-east-1 fails constantly, it has terrible uptime, and you should expect it to go.” On LinkedIn, a poster using the handle Mavrukin framed the same complaint in historical terms: “us-east-1 was a liability 10+ years ago when I first used aws, it was a liability 5 years ago, and it’s a liability today.”
Aidan Steele, a cloud engineer who writes for the Last Week in AWS blog, frames the underlying trade-off more cynically: “AWS is going to milk customers like cows to achieve significant reliability.” His point is that real resilience means paying for multi-region architecture and premium support, not assuming the base product will hold up by default.
The Regulatory Angle: Why Cloud Concentration Risk Is Getting Attention
Regulators have started treating hyperscaler dependence as a systemic risk instead of a routine vendor-management detail. The EU’s Digital Operational Resilience Act, which has applied to financial firms and their tech suppliers since January 17, 2025, built an oversight framework specifically for third-party ICT concentration risk. The European Commission has set a broader target too: cut the bloc’s reliance on non-EU technology providers from more than 80% today to 40% by 2030, according to research published by the Cloud Security Alliance.
Canada’s government made a similar point more bluntly. Its National Cyber Threat Assessment for 2025-2026 states that vendor concentration is increasing cyber vulnerability, warning that a single incident at a dominant provider can hit an entire sector at once. A separate Cloud Security Alliance study on AI and cloud provider concentration risk found that 81% of surveyed organizations would face severe or critical disruption from a seven-day outage at a single vendor. Seven days is a long outage by any measure, but 28 hours is already enough to make that exposure feel real.
Competitive Comparison: How the Big Three Handle Failure
The three hyperscalers fail in genuinely different patterns. AWS incidents concentrate in us-east-1 and tend to hit the largest number of customers per event, largely because so much legacy capacity still sits there by default. Azure’s worst incidents, the October 2025 Front Door failure and a February 2026 storage and identity failure that ran roughly 12 hours, tend to surface through Microsoft 365 and Teams, giving them outsized visibility in the enterprise world even when the underlying cause is narrower. Google Cloud fails less often, with 7 documented major incidents since 2016 versus 9 each for AWS and Azure, but its global control plane means a single bad configuration change can ripple across dozens of products within minutes, which is what happened in June 2025.
None of the three make customers whole after an outage. AWS’s SLA credit tops out at 30% of a monthly service fee. Azure’s can reach 100% for its worst-affected services. Google Cloud caps out around 50% for Compute Engine. In every case, the payout is calculated against the provider’s invoice, not the customer’s actual loss, which is why analysts keep describing these credits as a rounding error next to real business impact.
Beyond the Big Three: CrowdStrike and Cloudflare Show the Risk Is Systemic
Hyperscalers aren’t the only single point of failure the internet leans on. CrowdStrike’s faulty Falcon sensor update in July 2024 wasn’t a cloud-provider failure at all, but it crashed Windows hosts worldwide, including Azure virtual machines, and analytics firm Parametrix estimated the damage to Fortune 500 companies alone at $5.4 billion. In November 2025, a configuration-propagation failure at Cloudflare degraded service for a large share of the roughly 2.4 billion internet users whose traffic routes through its network, an incident researchers estimated cost upward of $250 million. Shattered.io covered the scale of Cloudflare’s exposure in its 2026 threat report analysis.
The lesson compounds rather than cancels out. Enterprises that diversified away from a single cloud region still route through the same handful of CDNs, DNS providers, and endpoint security vendors. Concentration risk doesn’t disappear when a company adds a second cloud. It just moves up or down the stack.
Market Impact: Stocks, SLAs, and Enterprise Trust
None of this has slowed cloud spending. Gartner-reported SaaS spending grew nearly 12% in 2025 and is forecast to grow another 15% in 2026, according to Flexera’s 2026 State of the Cloud Report. Enterprises are spending more on cloud infrastructure even as outage headlines pile up, because running equivalent infrastructure in-house remains slower and more expensive for most workloads. That dynamic echoes what shattered.io found when Google Cloud posted 63% growth against AWS and Azure earlier this year: the market keeps rewarding scale even when reliability wobbles.
What’s changing is where the spending goes. Forrester’s 2026 predictions research flags cloud outages, private AI on private infrastructure, and the rise of smaller “neocloud” providers as defining 2026 themes. InfoWorld put the shift in a single headline: “2026: the year we stop trusting any single cloud.” Enterprise architects aren’t abandoning AWS. They’re building failover paths that assume it will fail again, the same instinct that shows up whenever a platform outage disrupts paying customers, as it did during Sony’s PSN outage that delayed a $1 million esports final and during Xbox’s second network outage in a single week.
5 Predictions for Cloud Reliability Through 2027
- Multi-AZ reviews become a board-level line item. After two major us-east-1 incidents in eight months, expect large AWS customers to demand documented, tested failover across availability zones, not architecture diagrams that merely assume it works.
- EU sovereignty rules start reshaping vendor selection. With DORA enforcement maturing and the EU’s 40%-by-2030 target for non-EU tech reliance, expect more regulated European workloads to shift toward EU-based options or explicit multi-cloud contracts.
- Cyber insurers start pricing hyperscaler concentration directly. CyberCube’s public modeling of the October 2025 event is a preview. Expect underwriters to ask which cloud, which region, and which availability zones a policyholder depends on before setting premiums.
- AWS puts real capital into us-east-1’s physical resilience. Two failures in one region within eight months, one software and one physical, is a pattern AWS can’t out-message. Expect visible investment in cooling redundancy and public commentary on data-center hardening.
- SLA credit structures face pressure to change. The gap between a 10%-to-30% service credit and losses in the tens or hundreds of millions is now documented across multiple incidents. Expect large customers to negotiate custom resilience clauses instead of relying on standard SLA terms.
Frequently Asked Questions
What caused the AWS us-east-1 outage in May 2026?
A physical cooling failure, not a software bug. Multiple cooling units failed in a single data-center hall serving the use1-az4 availability zone. As temperatures rose past safe limits, servers shut down automatically to protect the hardware, and the resulting power loss impaired EC2 instances and EBS volumes on the affected racks.
How long did the outage last?
AWS’s first status update went out around 5:25 PM PDT on May 7. Cooling stabilized roughly 20 hours later, around 1:50 PM PDT on May 8, and most services recovered then. Data-heavy services took longer to finish restoring, pushing full recovery to about 28 hours.
Which companies were affected?
Coinbase, FanDuel, CME Group, and the humanitarian platform KoboToolbox all reported disruptions. Coinbase’s trading systems were down for about seven hours. FanDuel couldn’t process cash-outs during a live NBA playoff game. CME Group’s CME Direct tool threw error screens at institutional traders.
Was this outage as severe as the October 2025 AWS outage?
No. The October 2025 incident was a region-wide DynamoDB DNS failure that affected an estimated 70,000 organizations and produced a CyberCube insured-loss estimate of $38 million to $581 million. May 2026’s failure stayed confined to a single availability zone, which limited the blast radius even though the outage itself ran longer: about 28 hours versus roughly 15.
How does AWS’s market share compare to Azure and Google Cloud?
AWS held 28% of worldwide cloud infrastructure spending in the fourth quarter of 2025, ahead of Azure’s 21% and Google Cloud’s 14%, according to Synergy Research Group. The three together control 63% of the market.
Does AWS compensate customers when us-east-1 goes down?
Only partially. AWS’s EC2 service-level agreement pays 10% to 30% of the affected service’s monthly fee, depending on how far uptime fell short of 99.99%. It never covers a percentage of the customer’s actual lost revenue.
Is us-east-1 less reliable than other AWS regions?
The available data says yes. us-east-1 has logged three major, structurally distinct failures since 2021, a scaling bug, a DNS race condition, and a cooling failure, more than any other AWS region over the same period. One tracking post put its 2025 downtime at three times the rate of other AWS regions.
What can businesses do to reduce exposure to AWS outages?
Spread workloads across multiple Availability Zones at minimum, and test that failover actually works rather than assuming it does. Multi-AZ architecture would have protected against the May 2026 single-zone event, though it wouldn’t have helped during October 2025’s region-wide control-plane failure, which is why serious resilience planning has to address both failure modes.
Related Coverage
- Google Cloud Hits 63% Growth, Outpaces AWS, Azure [2026]
- Cloudflare 2026 Threat Report: 47M Attacks, 31.4 Tbps Record [2026]
- Xbox Down 15 Hours: Second Outage in a Week [2026]
- PSN Outage Delays $1M Esports Final, Sony’s 7th [2026]
- WEF Cybersecurity Outlook 2026: Fraud Tops CEO Fears, 94% Cite AI Risk




