Briefs · V · Ideatives Inc. · Updated
Your provider's outage is yours.
The longest 2026 provider incident reviewed here ran 22 hours 6 minutes: lightning caused power disturbances that impaired Azure's West US 2 region on May 29–30, per Microsoft's review. Across 19 provider incidents from 2026, causes ranged from config changes and automation defects to power and cooling failures, cut fiber and a provider suspending a customer's account.
Microsoft Azure PIR GHRP-84G, May 2026 · postmortem.io · 19 incidents selected by us; not a census
How long did 2026 provider incidents last?
Across the 16 2026 incidents with a stated duration, impact ran from 1 hour 41 minutes (Bitbucket Cloud, March) to 22 hours 6 minutes (Azure West US 2, May), per each company's own report.
Incident log, 2026
Sep 2026
Azure, multiple regions: gateway connectivity
A gateway manager change overloaded during OS servicing; degraded or interrupted connectivity for 5 hours 45 minutes. Preliminary review.
Sep 2026
Azure Sweden Central: AI services
Intermittent failures and 5XX errors for Azure OpenAI, Foundry and Cognitive Services for about 6 hours. Preliminary review.
Sep 2026
Google Cloud us-central1: fiber cut during maintenance
Fiber-optic cables were disconnected by mistake during routine hardware work. Network isolation in two zones for 4 hours 11 minutes.
Aug 2026
DigitalOcean: auth database migration
A planned migration of the identity database set off retry amplification. Control panel and API errors peaked at 25% over 3.5 hours; running Droplets were unaffected.
Show 15 earlier 2026 incidents, incl. Azure West US 2 (May, 22 hours)
Aug 2026
Google Cloud us-west1: fiber maintenance
Congestion after fiber-optic maintenance degraded Compute Engine, GKE, Cloud Run, Cloud Storage and IAM for 2 hours 22 minutes.
Jul 2026
Azure West US: repair scope bug
A defect in blast-radius analysis widened a single-device repair to every optical device leaving a datacenter. Connectivity failures for about 5 hours.
Jul 2026
Google Cloud VMware Engine: route drops
A network control-plane config update dropped inter-site traffic for stretched clusters for 10 hours 40 minutes.
May 2026
Azure West US 2: lightning, power and cooling
Lightning caused utility power disturbances across datacenters in two Availability Zones; services were impaired for 22 hours 6 minutes.
May 2026
Azure OpenAI: retry storm
Retry traffic from an upstream internal workload overloaded the routing layer for 7 hours 26 minutes.
May 2026
Railway: Google Cloud suspended its account
Google Cloud suspended Railway's production account. API, control plane and databases went offline for about 8 hours.
May 2026
AWS us-east-1: thermal event
Rising temperatures in one data center cut power to hardware in a single Availability Zone (use1-az4).
Apr 2026
Azure East US: control-plane lock contention
Lock contention in the PubSub networking control plane caused provisioning failures for 11 hours 52 minutes.
Mar 2026
Azure OpenAI: GPT-5.2 config mismatch
A GPT-5.2 configuration was incompatible with the deployed engine version; degradation lasted 20 hours 12 minutes.
Mar 2026
Bitbucket: hosting provider rate limit
Bitbucket Cloud hit a regional provisioning API rate limit at its hosting provider. Web, API and Pipelines were disrupted for 1 hour 41 minutes.
Feb–Mar 2026
GitHub: repeated availability incidents
Significant incidents on February 2, February 9 and March 5. GitHub published a review of the causes.
Feb 2026
Google Cloud Vertex AI: safety-filter change
A configuration change to the safety-filtering service for Gemini models caused 429 and 503 errors for 1 hour 58 minutes.
Feb 2026
Azure West US: transformer failure
An onsite transformer failure cut utility power to a datacenter; service was impaired for 20 hours 26 minutes.
Feb 2026
Azure: storage policy remediation
A policy workflow disabled anonymous access on Microsoft-managed storage, breaking VM and managed-identity operations across regions for about 6.5 hours.
Jan 2026
Cloudflare 1.1.1.1: CNAME ordering change
A memory optimization changed the order of CNAME records in DNS answers and broke resolution for some clients.
From the companies' own reports. Several 2026 incidents were physical (heat, fiber) rather than configuration changes. Our selection, not a census. More on postmortem.io, #outage.
For comparison: what caused the major 2025 cloud outages?
Per the providers' own postmortems, the five major 2025 cloud outages reviewed here came from configuration changes or latent automation bugs. None was an attack. AWS us-east-1 on Oct 19–20 lasted longest: 14 hours 32 minutes, after a latent race condition in DynamoDB's DNS automation.
Causes and costs
No attack
"Not caused, directly or indirectly, by a cyber attack." Cloudflare first suspected a hyper-scale DDoS.
Cloudflare postmortem, Nov 18 2025
2/3
Of publicly reported outages are attributed to third-party IT and data center providers. Sample size not disclosed.
Uptime Institute, May 2026, via ITWeb
57%
Of respondents said their most recent major outage cost over $100,000. 1 in 5 reported over $1 million.
Uptime Institute survey, 2026, via ITWeb
Uptime also says power remains the leading cause of impactful outages. Config changes are not the main cause overall.
How a safe change goes global
Change
Looks valid
Azure: valid customer config changes produced incompatible metadata.
Azure PIR YKYN-BWZ
Change
Passes the gates
CrowdStrike, 2024: the bad update "passed validation."
CrowdStrike PIR, Jul 2024
Change
Spreads globally
Google: the bad policy "replicated globally within seconds."
Google Cloud, Jun 2025
Failure
Latent bug fires
Cloudflare: a feature file doubled in size and went over a 200-feature limit.
Cloudflare, Nov 2025
Failure
Recovery overloads
Google: a "herd effect" in us-central1 made recovery worse.
Google Cloud, Jun 2025
Our view: the five steps form a common pattern. Each step is cited; linking them is our interpretation. Staged rollouts, validators and health gates did not catch these failures.
What to do
- Treat your provider's change pipeline as a dependency. Map the control planes you rely on: DNS, IAM, CDN.
- Keep a degraded mode that works without the provider's control plane: cached auth, static failover, a second region or CDN.
- Rehearse the failover.
Questions
- What were the major cloud outages of 2026?
- Incidents reviewed here include Azure West US 2 (22 hours, May, power and cooling), Azure West US (20 hours, February, transformer failure), Azure OpenAI GPT-5.2 (20 hours, March), Azure East US (12 hours, April), Google Cloud VMware Engine (11 hours, July) and Railway (about 8 hours after Google Cloud suspended its account, May), per each company's report.
- What caused the AWS us-east-1 outage in October 2025?
- A latent race condition in DynamoDB's DNS automation, per AWS's post-event summary. Services were affected for 14 hours 32 minutes.
- Was the Cloudflare outage of November 18, 2025 a cyberattack?
- No. Cloudflare's postmortem attributes it to an internal change.
- How many outages come from third-party providers?
- About two-thirds of publicly reported outages are attributed to third-party IT and data center providers, per Uptime Institute's 2026 analysis.
Sources
- Microsoft, Azure, multiple regions: gateway connectivity · postmortem.io, Sep 2026. Vendor, first-party.
- Google Cloud, Google Cloud VMware Engine: route drops · postmortem.io, Jul 2026. Vendor, first-party.
- Microsoft, Azure West US 2: lightning, power and cooling · postmortem.io, May 2026. Vendor, first-party.
- Microsoft, Azure OpenAI: retry storm · postmortem.io, May 2026. Vendor, first-party.
- Microsoft, Azure East US: control-plane lock contention · postmortem.io, Apr 2026. Vendor, first-party.
- Microsoft, Azure OpenAI: GPT-5.2 config mismatch · postmortem.io, Mar 2026. Vendor, first-party.
- Google Cloud, Google Cloud Vertex AI: safety-filter change · postmortem.io, Feb 2026. Vendor, first-party.
- Microsoft, Azure West US: transformer failure · postmortem.io, Feb 2026. Vendor, first-party.
- Microsoft, Azure: storage policy remediation · postmortem.io, Feb 2026. Vendor, first-party.
- Cloudflare, Cloudflare 1.1.1.1: CNAME ordering change · postmortem.io, Jan 2026. Vendor, first-party.
- Microsoft Azure, PIR 7QL5-Z50 (preliminary), September 29, 2026. Vendor, first-party. · postmortem.io
- Microsoft Azure, PIR ZJV6-SGG, July 23, 2026. Vendor, first-party. · postmortem.io
- Google Cloud, us-central1 incident, September 1, 2026. Vendor, first-party. · postmortem.io
- DigitalOcean, Cloud Control Panel and API postmortem, August 24, 2026. Vendor, first-party. · postmortem.io
- Google Cloud, us-west1 incident, August 20, 2026. Vendor, first-party. · postmortem.io
- Railway, GCP account suspension incident report, May 20, 2026. Vendor, first-party. · postmortem.io
- AWS Health Dashboard, use1-az4 thermal event, May 2026. Vendor, first-party. · postmortem.io
- Atlassian, Bitbucket incident report, March 6, 2026. Vendor, first-party. · postmortem.io
- GitHub, Addressing GitHub's recent availability issues, March 11, 2026. Vendor, first-party. · postmortem.io
- AWS, DynamoDB service disruption summary, Oct 19–20 2025. Vendor, first-party; no publication date on the page. · postmortem.io
- Microsoft Azure, Front Door PIR YKYN-BWZ, Oct 29–30 2025. Vendor, first-party. · postmortem.io
- Cloudflare, Nov 18 2025 postmortem, M. Prince. Vendor, first-party. · postmortem.io
- Cloudflare, Dec 5 2025 postmortem. Vendor, first-party. · postmortem.io
- Google Cloud, incident report, Jun 12 2025. Vendor, first-party. · postmortem.io
- CrowdStrike, preliminary post-incident review, Jul 2024. Vendor, first-party; context, not a 2025 incident.
- Uptime Institute, 2026 outage analysis release, May 13 2026, via ITWeb (also Business Wire China). Industry body; full report paywalled; sample size not disclosed; publicly reported outages only.
- Uptime Institute, 2025 outage analysis release, May 6 2025.