ChatGPT, Claude, Gemini and Grok all buckled within the same 90-minute window on Thursday morning, September 3, 2026, sending outage reports climbing into the tens of thousands and knocking millions of daily workflows offline at once. The trigger, according to multiple outlets tracking the event, was a regional failure inside Microsoft Azure’s East US infrastructure, the same cloud backbone that happens to host three of the four biggest AI chatbots on the market.
By the time engineers had mitigations in place, Downdetector had logged more than 37,000 reports for OpenAI’s ChatGPT, over 1,300 for Claude, and roughly 1,365 for Grok. Google’s Gemini, running on Google Cloud rather than Azure, stayed largely upright, with only about 500 reports at the peak. 9to5Google, Forbes and LADbible all confirmed the simultaneous disruption within minutes of the first reports surfacing.
Timeline: How the September 3 AI Outage Unfolded
The first cracks showed up early. Downdetector logged more than 5,000 ChatGPT-related reports by 7:53 a.m., a number that looked like routine morning noise until it kept climbing. Within a couple of hours that figure had jumped past 22,000, and by mid-morning combined reports for ChatGPT and OpenAI’s coding tool Codex topped 66,000, according to outage-tracking snapshots cited alongside the incident. A separate, independently corroborated tally put ChatGPT reports at more than 37,000 as of roughly 11 a.m. ET, the number most outlets settled on as the outage’s headline figure.
OpenAI’s own status page told a version of the same story. The company said it had detected “elevated errors across ChatGPT and Codex” at approximately 10:58 a.m. UTC, with the disruption affecting 15 separate components on ChatGPT and four on Codex. That timestamp translates to roughly 6:58 a.m. ET, meaning the underlying failure was already in motion well before Downdetector’s report volume made it obvious to the public. Engineers applied mitigations through the morning as report counts kept rising toward their late-morning peak.
Claude and Grok followed a tighter, later curve. Anthropic’s assistant peaked at 1,324 reports just before 11 a.m. ET, while xAI’s Grok peaked earlier, near 10 a.m. ET, at 1,365 reports. Microsoft’s own Copilot assistant was also flagged for “stability and downtime issues” as the incident played out, adding a fifth major assistant to the list of platforms touched by the same underlying infrastructure problem. By 12:42 p.m. ET (9:42 a.m. PT), most tracking services showed ChatGPT, Claude, Grok and Gemini returning to normal operation, closing out an incident that ran for roughly two hours from first detection to broad recovery.
Outage Reports by Platform: The Numbers
The scale gap between ChatGPT and everything else is the most striking number to come out of Thursday’s incident. Here is how the four major assistants compared at their respective peaks, based on Downdetector snapshots cited by multiple outlets covering the event.
| Platform | Operator | Peak reports | Approx. peak time (ET) | Primary cloud host |
|---|---|---|---|---|
| ChatGPT / Codex | OpenAI | 37,000+ (66,000+ combined with Codex, per some snapshots) | ~11:00 a.m. | Microsoft Azure |
| Claude | Anthropic | 1,324 | ~10:55 a.m. | Microsoft Azure |
| Grok | xAI / X | 1,365 | ~10:00 a.m. | Microsoft Azure-linked infrastructure |
| Gemini | ~500 | ~11:00 a.m. | Google Cloud | |
| Copilot | Microsoft | Not separately quantified | Morning hours | Microsoft Azure |
Read the table carefully and a pattern jumps out: every platform that stumbled hard runs on Azure infrastructure. Gemini, the outlier, was never really at risk of a full outage because it sits on Google’s own cloud stack. That single fact turned Thursday’s incident from an “AI is unreliable” story into a much narrower story about concentration risk in cloud infrastructure.
Root Cause: Azure East US Takes Down Three AI Giants at Once
According to reporting from Tech Times, the common thread linking ChatGPT, Claude and Grok’s simultaneous failure was a breakdown in Microsoft Azure’s East US region. All three services lean on Azure-hosted compute and networking for at least part of their production stack, and when that regional infrastructure degraded, the failure propagated to each of them at roughly the same time rather than one after another. Cloudflare, which routes and caches traffic for a large share of the internet, was also reported to be experiencing issues during the same window, and Amazon Web Services was named alongside it as showing symptoms, though neither was identified as the root trigger the way Azure East US was.
That combination matters. A single-region cloud failure is normally a contained event: it degrades service for the customers running workloads in that specific region, and traffic can often be rerouted. What made Thursday’s incident unusual is how many separate, competing AI companies had meaningful production dependencies sitting in the same region of the same cloud provider. OpenAI, Anthropic and xAI compete aggressively on model quality and pricing, but on infrastructure, at least on September 3, they turned out to share a surprising amount of exposure.
Why Gemini Kept Running
Gemini’s resilience wasn’t luck. Google runs Gemini on its own cloud, built and operated end-to-end by the same company that owns the model. That vertical integration meant Thursday’s Azure East US failure simply had no attack surface inside Google’s stack. The roughly 500 Gemini-related reports that did show up on Downdetector were most likely a mix of unrelated client-side issues, regional network hiccups, and users who assumed a broader internet problem was affecting every AI tool at once, whether or not it actually was.
OpenAI’s Response and Status Page Disclosure
OpenAI’s incident communication followed the pattern the company has used for prior disruptions this year: a status-page update acknowledging “elevated error rates across ChatGPT and Codex,” a component-level breakdown (15 ChatGPT components and four Codex components flagged as degraded), and a running set of updates as mitigations rolled out. The company did not immediately attribute the incident publicly to Azure by name in its own status messaging, according to the coverage reviewed, though the timing lines up precisely with the broader Azure East US narrative reported elsewhere.
For a company whose flagship product now sits at the center of hundreds of millions of daily interactions, a two-hour degradation window is not a minor blip. ChatGPT’s report volume alone, at 37,000-plus and climbing past 66,000 on a combined basis with Codex, made this one of the more heavily reported outages Downdetector has tracked for an AI platform in 2026.
Anthropic and xAI: Smaller Numbers, Same Root Cause
Claude’s 1,324-report peak and Grok’s 1,365-report peak are each a fraction of ChatGPT’s, but that gap mostly reflects relative user-base size rather than a difference in how badly each service was affected technically. Both Anthropic and xAI depend on Azure-linked infrastructure for meaningful parts of their production traffic, and both showed the classic signature of a shared-infrastructure incident: reports rising and falling on a similar curve to ChatGPT’s, offset by roughly the smaller scale of each platform’s user base. Neither Anthropic nor xAI was singled out in the reporting as having a unique, service-specific root cause distinct from the Azure East US failure.
Microsoft Copilot’s Rough Morning
Copilot, Microsoft’s own AI assistant, is arguably the most telling data point in the whole incident. It runs on Microsoft’s own cloud, and it still had a rough morning, described in coverage as suffering “stability and downtime issues.” If Azure’s own first-party AI product wasn’t immune to the East US disruption, that’s a strong signal the failure sat deep in shared infrastructure layers rather than in any one customer’s configuration, whether that customer is OpenAI, Anthropic, or Microsoft itself.
Historical Context: A Year Defined by Cloud Fragility
Thursday’s incident didn’t happen in a vacuum. 2026 has been a rough year for cloud reliability across every major provider, and this outage joins a growing list of high-profile incidents that have put a spotlight on how much of the internet, and now how much of the AI industry, rides on a handful of hyperscale regions.
| Incident | Provider | Approx. duration / scale | What it exposed |
|---|---|---|---|
| September 3 AI outage | Microsoft Azure (East US) | ~90 minutes to ~2 hours, 37,000+ reports | ChatGPT, Claude and Grok share Azure exposure |
| AWS us-east-1 outage | Amazon Web Services | 28 hours, third us-east-1 failure of the year | Repeat fragility in AWS’s flagship region |
| Google Cloud multi-service outage | Google Cloud | 2 hours 22 minutes, 33 services hit | Even vertically integrated stacks aren’t immune |
| Cloudflare August outage streak | Cloudflare | 13 outages logged in 8 days | Edge/CDN layer instability compounding cloud risk |
The pattern across these incidents is consistent: it is rarely the AI model itself that fails. It’s the cloud region underneath it, the edge network in front of it, or the load balancer routing traffic to it. Shattered.io has covered each of these in detail, including the 28-hour AWS us-east-1 outage, the Google Cloud outage that hit 33 services, and Cloudflare’s 13-outage stretch in August. Thursday’s event is the first of these in 2026 to visibly take down three separate, competing AI labs at the exact same time.
Competitive Comparison: Single-Cloud vs. Multi-Cloud AI Infrastructure
The outage draws a sharper line than usual between how the major AI labs approach infrastructure. OpenAI’s relationship with Microsoft goes back to a multi-billion-dollar investment and deep Azure integration, which gives ChatGPT enormous compute access but also ties its uptime tightly to Azure’s regional health. Anthropic has historically run a multi-cloud strategy spanning AWS, Google Cloud and Azure, yet Thursday’s numbers suggest a meaningful share of Claude’s production traffic still routes through Azure-linked paths, enough to produce a visible spike in reports during the East US failure. xAI’s Grok showed a similar pattern.
Google is the clear structural outlier. Because Gemini runs on infrastructure Google owns and operates itself, from custom silicon up through the data center network, it has no equivalent single point of failure tied to a third-party cloud vendor. That’s not a guarantee against Gemini ever going down (Google Cloud has had its own multi-service outages in 2026), but it does mean Gemini’s fate isn’t coupled to Microsoft’s operational decisions the way ChatGPT’s, Claude’s and Grok’s currently are.
What This Means for Enterprise Customers
Enterprises that built critical workflows on top of ChatGPT or Claude got a real-world stress test on Thursday, whether they wanted one or not. Customer support bots, coding assistants, internal search tools and automated document pipelines built on either platform went dark for up to two hours during a weekday morning, a window that overlaps with peak business hours across the US East Coast and much of Europe’s afternoon. For companies running these tools in anything resembling a mission-critical capacity, the incident is a reminder that a vendor’s AI model quality and a vendor’s infrastructure resilience are two separate things that need to be evaluated separately.
Market and Business Impact
None of the four companies involved has published a financial estimate of the outage’s cost, and none should be expected to for an incident this short. But the indirect impact is easier to reason about. ChatGPT alone handles hundreds of millions of interactions per day across consumer chat, the API, and Codex; even a partial, 15-component degradation lasting roughly two hours translates into a meaningful volume of failed requests, retried API calls, and interrupted developer workflows. For businesses that pay per API call and build retry logic into their own products, a spike in failed requests during peak hours is a direct, if modest, cost.
There’s also a trust dimension that doesn’t show up on a balance sheet immediately. This is not the first time a major cloud or AI outage has made headlines in 2026, and each additional incident adds to a cumulative narrative that AI infrastructure, however impressive the underlying models have become, is still built on the same fallible cloud regions that have produced repeat failures at AWS, Google Cloud and now Azure. That narrative matters most for enterprise sales conversations, where procurement teams increasingly ask AI vendors pointed questions about uptime history and disaster-recovery plans before signing multi-year contracts.
How Downdetector and Status Pages Track These Incidents
Most of the numbers behind Thursday’s coverage come from Downdetector, a crowdsourced outage-tracking service that counts user-submitted reports in real time and layers them against a rolling baseline for each tracked platform. It isn’t a perfect measure of technical severity (a service can be fully down with relatively few reports if its user base is small, or show a report spike from unrelated causes) but it remains the fastest public signal available when several companies go quiet on their own status pages during an active incident. OpenAI’s own status page, which logs incidents at the individual-component level, offered the more precise technical picture once the company started posting updates, naming the specific ChatGPT and Codex components affected rather than describing the outage in blanket terms.
Predictions: What Comes Next After This Outage
- Multi-cloud redundancy becomes a selling point. Expect OpenAI, Anthropic and xAI to face renewed pressure from enterprise customers to demonstrate failover capability across more than one cloud provider, not just Azure.
- Status page transparency improves. Component-level incident reporting, the kind OpenAI used on Thursday, is likely to become the norm across all four companies rather than the exception, since vaguer status updates draw more public criticism during fast-moving incidents.
- Google leans into the “we don’t depend on a third party” narrative. Gemini’s relative stability during a rival-wide outage is a marketing point Google is unlikely to leave on the table in future enterprise pitches.
- Regulators and enterprise risk teams start asking sharper infrastructure questions. As AI tools move deeper into regulated workflows, expect procurement and compliance teams to start requesting concrete uptime SLAs and cloud-dependency disclosures as a condition of contracts.
- More simultaneous multi-platform outages are likely before infrastructure catches up. With OpenAI, Anthropic and xAI all carrying meaningful Azure exposure, another regional Azure incident could reproduce Thursday’s pattern until these companies diversify their hosting further.
The Bigger Picture: Concentration Risk in AI Infrastructure
Step back from the report counts and status-page language, and Thursday’s outage is really a story about how consolidated AI infrastructure has become underneath a market that looks, on the surface, wildly competitive. OpenAI, Anthropic and xAI spend enormous effort differentiating their models on benchmarks, pricing and safety posture, yet a single regional failure inside one cloud provider was enough to degrade all three at nearly the same moment. That’s a very different kind of risk than a bad model update or a bug in a specific feature. It’s systemic, and it doesn’t respect competitive rivalries.
The incident also lands at a moment when AI assistants have moved well past novelty status into daily infrastructure for millions of people and thousands of businesses. Two hours of degraded ChatGPT, Claude and Grok access on a Thursday morning is a very different event than the same outage would have been two years earlier, when far fewer companies had wired these tools into production systems. As that dependency deepens, incidents like this one are likely to draw more scrutiny, not less, from the businesses and regulators watching how resilient the AI industry’s foundations actually are.
Frequently Asked Questions
Were ChatGPT, Claude and Gemini all actually down at the same time on September 3, 2026?
ChatGPT, Claude and Grok experienced significant, confirmed disruptions within the same roughly 90-minute to two-hour window. Gemini showed a much smaller spike in reports, around 500 at its peak, and largely continued functioning because it runs on Google’s own cloud rather than the Azure infrastructure implicated in the other outages.
What caused the September 3 AI outage?
Reporting points to a failure inside Microsoft Azure’s East US region as the common root cause behind the ChatGPT, Claude and Grok disruptions, since all three rely on Azure-hosted infrastructure. Cloudflare and AWS were also reported as showing issues during the same window.
How many people were affected by the ChatGPT outage?
Downdetector logged more than 37,000 reports for ChatGPT as of around 11 a.m. ET, with some snapshots showing a combined ChatGPT-and-Codex total above 66,000 as the incident built through the morning. Report counts, not confirmed user totals, are the standard measure used for these incidents.
Why did Gemini stay up while ChatGPT, Claude and Grok went down?
Gemini runs on Google’s own cloud infrastructure rather than a third-party provider, so it had no exposure to the Azure East US failure that affected the other three platforms.
How long did the outage last?
OpenAI’s status page indicated elevated errors beginning around 10:58 a.m. UTC (roughly 6:58 a.m. ET), with report volumes peaking near 11 a.m. ET and most services reported as returning to normal by 12:42 p.m. ET.
Was Microsoft Copilot affected too?
Yes. Copilot, Microsoft’s own AI assistant, was reported to be suffering stability and downtime issues during the same window, despite running on Microsoft’s own Azure infrastructure.
Is this the first major AI outage of 2026?
No. 2026 has seen several major cloud-related outages, including a 28-hour AWS us-east-1 disruption, a Google Cloud outage that hit 33 services, and a stretch of 13 Cloudflare outages in eight days. Thursday’s incident is notable for hitting three competing AI labs at once rather than a single provider’s own customers.
Should businesses avoid relying on a single AI provider after this outage?
The incident strengthens the case for building fallback logic across more than one AI provider or cloud region for any workflow considered business-critical, since Thursday demonstrated that even well-resourced, competing AI labs can share the same underlying point of failure.




