Briefs · V · Ideatives Inc. · Updated

Your provider's outage is yours.

The longest 2026 provider incident reviewed here ran 22 hours 6 minutes: lightning caused power disturbances that impaired Azure's West US 2 region on May 29–30, per Microsoft's review. Across 19 provider incidents from 2026, causes ranged from config changes and automation defects to power and cooling failures, cut fiber and a provider suspending a customer's account.

Microsoft Azure PIR GHRP-84G, May 2026 · postmortem.io · 19 incidents selected by us; not a census

How long did 2026 provider incidents last?

Across the 16 2026 incidents with a stated duration, impact ran from 1 hour 41 minutes (Bitbucket Cloud, March) to 22 hours 6 minutes (Azure West US 2, May), per each company's own report.

  • Azure West US 2 (power and cooling), May 29–30: 22h06m
  • Azure West US (power), Feb 7–8: 20h26m
  • Azure OpenAI GPT-5.2, Mar 9–10: 20h12m
  • Azure East US control plane, Apr 24: 11h52m
  • Google Cloud VMware Engine, Jul 14: 10h40m
  • Railway (Google Cloud account suspension), May 19–20: 7h54m
  • Azure OpenAI routing, May 29: 7h26m
  • Azure VMs and managed identity, Feb 2–3: 6h27m
  • Azure Sweden Central, Sep 29: 5h55m
  • Azure gateways, multiple regions, Sep 30: 5h45m
  • Azure West US, Jul 23: 4h57m
  • Google Cloud us-central1, Sep 1: 4h11m
  • DigitalOcean control plane, Aug 24: 3h30m
  • Google Cloud us-west1, Aug 20: 2h22m
  • Google Cloud Vertex AI, Feb 27: 1h58m
  • Bitbucket Cloud, Mar 6: 1h41m
Duration from each company's stated start and end times. DigitalOcean: 3.5 hours of active degradation within a 6-hour window. Azure Sweden Central is a preliminary review. Not shown: the AWS us-east-1 thermal event (May), GitHub's February–March incidents and Cloudflare 1.1.1.1 (January), which have no single stated duration. Azure September 30 is a preliminary review. Our selection of 2026 incidents, not a census.

Incident log, 2026

  1. Sep 2026

    Azure, multiple regions: gateway connectivity

    A gateway manager change overloaded during OS servicing; degraded or interrupted connectivity for 5 hours 45 minutes. Preliminary review.

    Microsoft · postmortem.io

  2. Sep 2026

    Azure Sweden Central: AI services

    Intermittent failures and 5XX errors for Azure OpenAI, Foundry and Cognitive Services for about 6 hours. Preliminary review.

    Microsoft · postmortem.io

  3. Sep 2026

    Google Cloud us-central1: fiber cut during maintenance

    Fiber-optic cables were disconnected by mistake during routine hardware work. Network isolation in two zones for 4 hours 11 minutes.

    Google Cloud · postmortem.io

  4. Aug 2026

    DigitalOcean: auth database migration

    A planned migration of the identity database set off retry amplification. Control panel and API errors peaked at 25% over 3.5 hours; running Droplets were unaffected.

    DigitalOcean · postmortem.io

Show 15 earlier 2026 incidents, incl. Azure West US 2 (May, 22 hours)
  1. Aug 2026

    Google Cloud us-west1: fiber maintenance

    Congestion after fiber-optic maintenance degraded Compute Engine, GKE, Cloud Run, Cloud Storage and IAM for 2 hours 22 minutes.

    Google Cloud · postmortem.io

  2. Jul 2026

    Azure West US: repair scope bug

    A defect in blast-radius analysis widened a single-device repair to every optical device leaving a datacenter. Connectivity failures for about 5 hours.

    Microsoft · postmortem.io

  3. Jul 2026

    Google Cloud VMware Engine: route drops

    A network control-plane config update dropped inter-site traffic for stretched clusters for 10 hours 40 minutes.

    Google Cloud · postmortem.io

  4. May 2026

    Azure West US 2: lightning, power and cooling

    Lightning caused utility power disturbances across datacenters in two Availability Zones; services were impaired for 22 hours 6 minutes.

    Microsoft · postmortem.io

  5. May 2026

    Azure OpenAI: retry storm

    Retry traffic from an upstream internal workload overloaded the routing layer for 7 hours 26 minutes.

    Microsoft · postmortem.io

  6. May 2026

    Railway: Google Cloud suspended its account

    Google Cloud suspended Railway's production account. API, control plane and databases went offline for about 8 hours.

    Railway · postmortem.io

  7. May 2026

    AWS us-east-1: thermal event

    Rising temperatures in one data center cut power to hardware in a single Availability Zone (use1-az4).

    AWS · postmortem.io

  8. Apr 2026

    Azure East US: control-plane lock contention

    Lock contention in the PubSub networking control plane caused provisioning failures for 11 hours 52 minutes.

    Microsoft · postmortem.io

  9. Mar 2026

    Azure OpenAI: GPT-5.2 config mismatch

    A GPT-5.2 configuration was incompatible with the deployed engine version; degradation lasted 20 hours 12 minutes.

    Microsoft · postmortem.io

  10. Mar 2026

    Bitbucket: hosting provider rate limit

    Bitbucket Cloud hit a regional provisioning API rate limit at its hosting provider. Web, API and Pipelines were disrupted for 1 hour 41 minutes.

    Atlassian · postmortem.io

  11. Feb–Mar 2026

    GitHub: repeated availability incidents

    Significant incidents on February 2, February 9 and March 5. GitHub published a review of the causes.

    GitHub · postmortem.io

  12. Feb 2026

    Google Cloud Vertex AI: safety-filter change

    A configuration change to the safety-filtering service for Gemini models caused 429 and 503 errors for 1 hour 58 minutes.

    Google Cloud · postmortem.io

  13. Feb 2026

    Azure West US: transformer failure

    An onsite transformer failure cut utility power to a datacenter; service was impaired for 20 hours 26 minutes.

    Microsoft · postmortem.io

  14. Feb 2026

    Azure: storage policy remediation

    A policy workflow disabled anonymous access on Microsoft-managed storage, breaking VM and managed-identity operations across regions for about 6.5 hours.

    Microsoft · postmortem.io

  15. Jan 2026

    Cloudflare 1.1.1.1: CNAME ordering change

    A memory optimization changed the order of CNAME records in DNS answers and broke resolution for some clients.

    Cloudflare · postmortem.io

From the companies' own reports. Several 2026 incidents were physical (heat, fiber) rather than configuration changes. Our selection, not a census. More on postmortem.io, #outage.

For comparison: what caused the major 2025 cloud outages?

Per the providers' own postmortems, the five major 2025 cloud outages reviewed here came from configuration changes or latent automation bugs. None was an attack. AWS us-east-1 on Oct 19–20 lasted longest: 14 hours 32 minutes, after a latent race condition in DynamoDB's DNS automation.

  • AWS us-east-1, Oct 19–20, 2025: 14h32m
  • Azure Front Door, Oct 29–30, 2025: 8h24m
  • Cloudflare, Nov 18, 2025: 5h46m
  • Google Cloud, Jun 12, 2025: 3h
  • Cloudflare, Dec 5, 2025: 25m
Durations computed from each provider's own start and end times; definitions differ by provider, and customer-side recovery ran longer. AWS is the last affected service (ECS/EKS/Fargate); DynamoDB itself was out 2h52m. Cloudflare Nov 18 core traffic was largely normal after 3h10m. Causes: AWS, a latent race condition in DynamoDB DNS automation; Azure, valid customer config changes exposing a latent bug; Cloudflare and Google, internal changes. All vendor-reported. Sources: AWS, Microsoft Azure, Cloudflare, Google Cloud post-incident reports.

Causes and costs

  • No attack

    "Not caused, directly or indirectly, by a cyber attack." Cloudflare first suspected a hyper-scale DDoS.

    Cloudflare postmortem, Nov 18 2025

  • 2/3

    Of publicly reported outages are attributed to third-party IT and data center providers. Sample size not disclosed.

    Uptime Institute, May 2026, via ITWeb

  • 57%

    Of respondents said their most recent major outage cost over $100,000. 1 in 5 reported over $1 million.

    Uptime Institute survey, 2026, via ITWeb

Uptime also says power remains the leading cause of impactful outages. Config changes are not the main cause overall.

How a safe change goes global

  1. Change

    Looks valid

    Azure: valid customer config changes produced incompatible metadata.

    Azure PIR YKYN-BWZ

  2. Change

    Passes the gates

    CrowdStrike, 2024: the bad update "passed validation."

    CrowdStrike PIR, Jul 2024

  3. Change

    Spreads globally

    Google: the bad policy "replicated globally within seconds."

    Google Cloud, Jun 2025

  4. Failure

    Latent bug fires

    Cloudflare: a feature file doubled in size and went over a 200-feature limit.

    Cloudflare, Nov 2025

  5. Failure

    Recovery overloads

    Google: a "herd effect" in us-central1 made recovery worse.

    Google Cloud, Jun 2025

Our view: the five steps form a common pattern. Each step is cited; linking them is our interpretation. Staged rollouts, validators and health gates did not catch these failures.

What to do

  1. Treat your provider's change pipeline as a dependency. Map the control planes you rely on: DNS, IAM, CDN.
  2. Keep a degraded mode that works without the provider's control plane: cached auth, static failover, a second region or CDN.
  3. Rehearse the failover.

Questions

What were the major cloud outages of 2026?
Incidents reviewed here include Azure West US 2 (22 hours, May, power and cooling), Azure West US (20 hours, February, transformer failure), Azure OpenAI GPT-5.2 (20 hours, March), Azure East US (12 hours, April), Google Cloud VMware Engine (11 hours, July) and Railway (about 8 hours after Google Cloud suspended its account, May), per each company's report.
What caused the AWS us-east-1 outage in October 2025?
A latent race condition in DynamoDB's DNS automation, per AWS's post-event summary. Services were affected for 14 hours 32 minutes.
Was the Cloudflare outage of November 18, 2025 a cyberattack?
No. Cloudflare's postmortem attributes it to an internal change.
How many outages come from third-party providers?
About two-thirds of publicly reported outages are attributed to third-party IT and data center providers, per Uptime Institute's 2026 analysis.

Sources