The Operate blog

Field notes on AI SRE, root-cause analysis, and running self-hosted AI in production.

Stuck state: the failure mode behind both July 2 postmortems
SRE & Platform Engineering · July 19, 2026

Stuck state: the failure mode behind both July 2 postmortems

Railway and CircleCI both published July 2, 2026 postmortems with the same hidden failure mode: state captured during instability that never corrected itself.

Reduce MTTR by attacking the 80 percent: diagnosis, not repair
SRE & Engineering Efficiency · July 18, 2026

Reduce MTTR by attacking the 80 percent: diagnosis, not repair

About 80 percent of MTTR is spent finding the cause, not fixing it. A diagnosis-first playbook for reducing MTTR across code, database, infra, logs, and CI.

Beating On-Call Alert Fatigue: Triage by Evidence, Not by Threshold-Tuning
SRE & Incident Management · July 16, 2026

Beating On-Call Alert Fatigue: Triage by Evidence, Not by Threshold-Tuning

On-call alert fatigue is a triage problem, not a threshold problem. A practical guide to cutting page volume by investigating alerts to root cause across the wh

Replay-Verified Is Not Production-Safe: Where the Agentic Database's Autonomy Ladder Breaks
Engineering Strategy · July 16, 2026

Replay-Verified Is Not Production-Safe: Where the Agentic Database's Autonomy Ladder Breaks

A response to 'Agentic databases aren't agentic': replaying traffic on a fork proves answer-equivalence, not production safety. Why the reviewable diff, not the

Automating Root Cause Analysis Without Giving AI Write Access to Production
Engineering Leadership · July 15, 2026

Automating Root Cause Analysis Without Giving AI Write Access to Production

A practical, vendor-neutral guide to automated root cause analysis across the whole stack, and why the safe default is patch-file-only with no write access to p

Knight Capital's 97 Unread Emails: A $460M Lesson for the AI-Agent Era
Postmortems · July 15, 2026

Knight Capital's 97 Unread Emails: A $460M Lesson for the AI-Agent Era

Knight Capital lost $460M in 45 minutes. The SEC order reads like a checklist for teams deploying AI agents against production: unreviewed changes, unread signa

Idle in Transaction in Postgres: What It Means, How to Find It, and Why Autovacuum Can't Save You
Postgres · July 15, 2026

Idle in Transaction in Postgres: What It Means, How to Find It, and Why Autovacuum Can't Save You

What idle in transaction means in Postgres, why it silently blocks autovacuum while dashboards stay green, the 60-second diagnosis SQL, and the one-line timeout

The flaky test runbook: detect, quarantine, diagnose, verify
Engineering Operations · July 14, 2026

The flaky test runbook: detect, quarantine, diagnose, verify

A practical runbook that treats flaky tests as incidents: detection thresholds, a five-cause taxonomy, quarantine exit criteria, and the math of verifying a fix

The 21% Problem: Why Second Opinions are Critical for Root Cause Analysis
Engineering Leadership · July 14, 2026

The 21% Problem: Why Second Opinions are Critical for Root Cause Analysis

When Mayo Clinic re-examined referred diagnoses, 21% changed completely. Root cause analysis has the same failure mode and none of the safeguards.

Every Component Is Green and the System Is Down
Engineering & SRE · July 13, 2026

Every Component Is Green and the System Is Down

Pod healthy, service healthy, endpoint healthy - and the request times out. Four real incidents where every dashboard was green, and a cross-layer protocol for

When Recovery Doesn’t Recover: Railway, CircleCI, and Stuck State
Engineering & SRE · July 12, 2026

When Recovery Doesn’t Recover: Railway, CircleCI, and Stuck State

Railway and CircleCI both went down on July 2, 2026. The interesting failures came after recovery: connections and workflows that silently kept bad state.

Recovery Is Not Convergence: Anatomy of an Outage That Outlived Its Root Cause
Incident Analysis · July 11, 2026

Recovery Is Not Convergence: Anatomy of an Outage That Outlived Its Root Cause

Railway's July 2026 outage report shows three systems that captured bad state during 20 minutes of instability and held it silently. Here's the failure class.

Every On-Call Model Has an Expiry Headcount: Lessons from Monzo’s Public Record
Engineering Management · July 9, 2026

Every On-Call Model Has an Expiry Headcount: Lessons from Monzo’s Public Record

Monzo published three on-call models at three company sizes - with pay figures and honest failure notes. A teardown of the scaling cliffs every org hits.

Automating Root Cause Analysis Without Giving AI Write Access to Production
Engineering · July 9, 2026

Automating Root Cause Analysis Without Giving AI Write Access to Production

Manual RCA eats 30-40% of engineering time. Here's how to automate the investigation across your whole stack without giving AI write access to production.

Automated Root Cause Analysis in 2026: A Vendor-Neutral, Whole-Stack Guide for Engineering Leaders
Engineering · July 9, 2026

Automated Root Cause Analysis in 2026: A Vendor-Neutral, Whole-Stack Guide for Engineering Leaders

A vendor-neutral guide to automated root cause analysis across the whole stack, and why the safe default is patch-file-only with no write access to production.

The Buyer's Guide to Automated Root Cause Analysis: What the AI Answers Leave Out
Engineering Leadership · July 8, 2026

The Buyer's Guide to Automated Root Cause Analysis: What the AI Answers Leave Out

A vendor-neutral buyer's guide to automated root cause analysis: the features to look for, whole-stack coverage, and why the safe default is patch-file-only, no

The Monitoring Pilot Never Flies: Aviation's Rule for Verifying AI Diagnoses
Engineering Leadership · July 6, 2026

The Monitoring Pilot Never Flies: Aviation's Rule for Verifying AI Diagnoses

Airbus builds redundant flight computers with different chips, teams, and languages. Aviation's independence principle is the standard AI root-cause analysis sh

They Added Capacity and It Did Not Help: What WorkOS's Connection-Pool Starvation Says About Hold-and-Wait Failures
Postmortem Analysis · July 4, 2026

They Added Capacity and It Did Not Help: What WorkOS's Connection-Pool Starvation Says About Hold-and-Wait Failures

WorkOS added capacity and nothing improved. Hold-and-wait connection-pool starvation, why it self-sustains, and the audit to run on your own critical path.

When to Hire an SRE (According to People Not Selling You One)
Engineering Leadership · July 4, 2026

When to Hire an SRE (According to People Not Selling You One)

Everyone ranking for 'when to hire an SRE' is selling the answer. A disinterested framework: the real signals, what must exist before the hire, and the alternat

Toil Went Up, Not Down: The AI-Ops Paradox in the Survey Data
Engineering Leadership · July 3, 2026

Toil Went Up, Not Down: The AI-Ops Paradox in the Survey Data

Survey toil fell for years, then reversed as AI adoption went mainstream. The mechanism: verification burden, new ops surfaces, and AI aimed at the wrong loop.

Zero-Downtime Database Migrations: The Playbook and the Failure Modes Nobody Writes Down
Engineering Operations · July 2, 2026

Zero-Downtime Database Migrations: The Playbook and the Failure Modes Nobody Writes Down

The expand-contract playbook plus what guides skip: foreign keys, stored procedures, backup mismatches, and how to verify a live migration is safe step by step.

AI-Assisted Infra Changes Need a Paper Trail, Not a Chat Log
Engineering Strategy · July 1, 2026

AI-Assisted Infra Changes Need a Paper Trail, Not a Chat Log

AI-assisted terraform and pipeline changes ship fast, but the reasoning dies with the chat session. A practical convention for keeping the why in your repo.

Retries Are Eating Your Signal: Flaky Tests Are System Bugs Wearing a Costume
Reliability Engineering · June 27, 2026

Retries Are Eating Your Signal: Flaky Tests Are System Bugs Wearing a Costume

Google found 84% of CI pass-to-fail transitions are flakes. Retrying hides the race conditions and shared state that later cause production incidents.

Nobody Will Cut the Requests: A Blame-Safe Playbook for Rightsizing Kubernetes at 25 Percent Utilization
Engineering Management · June 25, 2026

Nobody Will Cut the Requests: A Blame-Safe Playbook for Rightsizing Kubernetes at 25 Percent Utilization

Your cluster runs at 25 percent because nobody wants to own the next latency blip. A rollout playbook for cutting padded Kubernetes requests safely.

The Delhi Data Center Fire Is a Blast-Radius Audit You Haven't Run Yet
Incident Response · June 24, 2026

The Delhi Data Center Fire Is a Blast-Radius Audit You Haven't Run Yet

A fire in one Delhi data center destroyed decades of customer data and degraded Google Cloud for weeks. Eight audit questions to ask about your own blast radius

The Pipeline Returned 200 OK and Did Nothing: The Agent Failure Class Your Monitoring Reports as Healthy
Engineering & SRE · June 22, 2026

The Pipeline Returned 200 OK and Did Nothing: The Agent Failure Class Your Monitoring Reports as Healthy

Agent pipelines fail while every dashboard stays green because 200 OK measures transport, not work. Five shapes of silent failure and how to detect each.

Schema Drift Is an Incident Waiting for a Deploy
Database Operations · June 22, 2026

Schema Drift Is an Incident Waiting for a Deploy

Schema drift turns deploys, failovers, and restores into incidents. How drift happens, what Snowflake's 13-hour outage shows, and how to detect and triage it.

Your Average MTTR Is a Fiction: What 1.8 Million Outages Say About the Tail
Engineering Management · June 22, 2026

Your Average MTTR Is a Fiction: What 1.8 Million Outages Say About the Tail

Fresh data from 1.8 million outages: median resolution 1.9 minutes, mean 21.9. The gap explains why your MTTR dashboard misleads, and what to report instead.

Knight Capital's 97 Unread Emails: A $460M Lesson for the AI-Agent Era
Engineering Leadership · June 21, 2026

Knight Capital's 97 Unread Emails: A $460M Lesson for the AI-Agent Era

Knight Capital lost $460M in 45 minutes. The SEC order reads like a checklist for teams deploying AI agents against production: unreviewed changes, unread signa

Your Observability Bill Is a Hedge Against Bad Investigations
Engineering Management · June 21, 2026

Your Observability Bill Is a Hedge Against Bad Investigations

Every cost guide says collect less telemetry. The real lever is why you store it: an investigation model that can only query pre-collected data. Fix that instea

How to Reduce MTTR: Fix the 80% of the Clock Nobody Instruments
Engineering Leadership · June 19, 2026

How to Reduce MTTR: Fix the 80% of the Clock Nobody Instruments

Most MTTR guides optimize detection. The clock is lost in diagnosis. A phase-by-phase breakdown of where incident time goes and how to shrink the 80%.

Is It Us or Is It AWS? Triage for the First Fifteen Minutes of an Upstream Outage
Incident Response · June 19, 2026

Is It Us or Is It AWS? Triage for the First Fifteen Minutes of an Upstream Outage

AWS's July 16 CloudFront outage took 59 minutes to reach the status page. A triage protocol for deciding, in the first fifteen minutes, whether an incident is y

The Andon Cord Almost Never Stops the Line: What Software Got Wrong About Toyota's Most Borrowed Idea
Engineering Operations · June 18, 2026

The Andon Cord Almost Never Stops the Line: What Software Got Wrong About Toyota's Most Borrowed Idea

Toyota's andon cord rarely stopped the line—it summoned help in seconds. What the real mechanics teach ops teams about escalation, freezes, and cheap signals.