Briefs · II · Ideatives Inc. · Updated

Stability drops as AI use rises.

According to Google's DORA 2024 report, every 25% increase in AI adoption is associated with an estimated 7.2% decrease in software delivery stability, and DORA 2025 (nearly 5,000 respondents) found the negative relationship persisted. This is a modeled correlation from survey data; it does not measure cause.

DORA 2024, Google Cloud · survey, modeled estimate

Does AI adoption hurt delivery stability?

In Google's DORA 2024 model, each 25% rise in AI adoption goes with an estimated 7.2% drop in delivery stability and a 1.5% drop in throughput, while documentation quality rises 7.5%. These are survey-based correlations.

  • Documentation quality: +7.5%
  • Code quality: +3.4%
  • Code review speed: +3.1%
  • Delivery throughput: −1.5%
  • Delivery stability: −7.2%
Modeled estimates from one survey; they show correlation only. On throughput the sources split: DORA 2025 found it rose, Cortex saw cycle time rise 9%. DORA is funded by Google, which sells AI tools. Source: DORA 2024, Google Cloud blog, October 2024.

Incidents and review

  • 2025

    “AI adoption does continue to have a negative relationship with software delivery stability.”

    DORA 2025, Google Cloud · nearly 5,000 respondents, Sept 2025

  • 3.4×

    The chance of a production incident per merged change. Incidents per PR up 242.7%.

    Faros AI, April 2026 · vendor telemetry, 22,000 developers

  • +31.3%

    Pull requests merged with no review, human or agentic, from low to high AI adoption.

    Faros AI, April 2026 · vendor telemetry

Output, review and incidents

  1. More output

    Epics per developer

    +66% from each organization's lowest-AI to highest-AI period.

    Faros AI, 2026

  2. Bigger changes

    PR size

    +51% over the same comparison.

    Faros AI, 2026

  3. Less review

    No-review merges

    +31.3%. Cortex: “pressure to maintain velocity often leads to rushed reviews.”

    Faros AI, 2026 · Cortex

  4. More failures

    Incidents per PR

    More than tripled, up 242.7%.

    Faros AI, 2026

  5. Less stability

    Delivery stability

    −7.2% per 25% rise in AI adoption, modeled.

    DORA 2024

Each step is correlated with the next; no study here shows that one causes another. Our view: shipping faster without review builds up debt, and the debt is paid in incidents.

What to do

  1. Gate AI-written PRs on size and require a review. Track no-review merges as a metric.
  2. Report change failure rate and incidents per PR alongside throughput.
  3. Spend the time AI saves on tests and smaller batches before adding scope.

Questions

Does AI adoption reduce delivery stability?
DORA's 2024 model estimates a 7.2% drop in delivery stability for each 25% rise in AI adoption. DORA 2025 (nearly 5,000 respondents) found the negative relationship persisted.
Does AI increase incidents?
Faros AI's 2026 telemetry shows incidents per PR up 242.7% (3.4×) in teams' highest- versus lowest-AI-use periods. It is vendor data and shows correlation only.
Does AI improve delivery throughput?
DORA 2024 estimated a 1.5% drop in throughput per 25% rise in AI adoption. DORA 2025 found throughput rose.

Sources

  • DORA 2024, Google Cloud blog, October 2024. Survey; funded by Google, which sells AI tools. Sample size not confirmed.
  • DORA 2025, Google Cloud blog, September 2025. Survey, nearly 5,000 respondents; funded by Google.
  • Faros AI, Engineering Report 2026 “Acceleration Whiplash”, April 2026, and research page. Vendor telemetry, 22,000 developers, 4,000+ teams, two years. Full report gated. Faros says its data contradicts DORA 2025 on strong foundations. A newer Faros report (Q3 2026) was not reviewed.
  • Cortex, 2026 Benchmark Report, November 2025. Vendor; 50 surveyed leaders plus metrics from an unstated number of organizations. Year over year, no control group: PRs per author +20%, incidents per PR +23.5%, change failure rate +30% (rollbacks only).
  • METR, July 2025: 16 developers, 246 issues, 19% slower with AI. February 2026 update, 57 developers: −18% and −4%, intervals crossing zero; the authors call it “only very weak evidence”.