Back to blog
    Incident Response
    By Jai Jalan16 min read

    How to Run Root Cause Analysis in Software Without Stopping at the First Why

    We walk through root cause analysis in software with six evidence steps, method choice, and a fault tree of Cloudflare's 2025 outage.

    Quick Answer

    Root cause analysis in software works best as an evidence procedure, not one favorite method. Freeze logs and deploy history, build a timeline from machine timestamps, list every change in the window, and test hypotheses that could fail. Then match the method to the incident. Use 5 whys for a single chain, fault trees when several conditions had to coincide, and change analysis after a deploy. No root cause analysis software replaces that evidence.

    Cloudflare's November 2025 outage flickered on and off every five minutes. Root cause analysis is the work of tracing a failure back through every condition that made it possible, so the fix prevents the next incident too.

    The flicker, plus a status page that failed by coincidence, pointed the team toward an attack, says Cloudflare's own postmortem. The trigger was a database permissions change. We think the first plausible story winning is the most common way software investigations go wrong.

    In this post we'll cover the six steps we follow, which method fits which incident, a worked example on that outage, and how to tell when you're done.

    How to Run Root Cause Analysis on a Software Incident, Step by Step

    We run root cause analysis on a software incident in six steps. We freeze the evidence, build a timeline, list the changes, test hypotheses, map every necessary condition, and turn each condition into a fix. The method you choose (5 whys, fishbone or fault tree) only matters at step five.

    Most guides start with the method. We start with evidence, because a method applied to a guess still produces a guess. This is root cause analysis in production, where the evidence is perishable and the system keeps changing while you look at it.

    1. Freeze the Evidence Before Retention Deletes It

    Evidence decays. Logs rotate, metrics get downsampled into coarser buckets, and the pod that crashed gets rescheduled with a clean slate. Before anyone writes a theory, we export the raw material for the incident window.

    Our freeze list for a typical web service:

    • Application and proxy logs for the window, plus 30 minutes before onset
    • Deploy history, feature flag changes and config diffs for the same span
    • Database activity, including slow query logs, lock waits and connection counts
    • Infrastructure events, including node restarts, autoscaling and certificate renewals
    • The incident chat, with its timestamps

    Retention rules decide what you can still prove a week later. We wrote about that trade-off in what the FAA's 25-hour rule teaches about log retention.

    2. Build the Timeline From Machine Timestamps

    We build the timeline from system timestamps in UTC, not from memory. AWS gives the same advice in its Correction of Error guidance. Start at the first trigger, not at the first page, and explain any gap longer than 10 to 15 minutes.

    The first trigger is usually earlier than the first alert. In Cloudflare's case, the triggering change landed at 11:05 UTC and customer traffic started failing at 11:28 UTC. A timeline that starts at the alert would never show the change.

    3. List Every Change in the Window

    Changes are the highest-yield lead in software. The Google SRE book says that roughly 70% of outages are due to changes in a live system. We list all of them, not only code deploys:

    Change type Where we look
    Code deploys CI/CD history, release tags
    Configuration and feature flags Config repo diffs, flag audit log
    Permissions and grants Database grants, IAM changes
    Schema and data Migrations, backfills, sudden table growth
    Dependencies Library upgrades, base image rebuilds, provider incidents
    Traffic New customers, batch jobs, retries from a client

    Drift counts as a change too. We cover that case in our configuration drift triage guide. Not every outage follows a change, though. GitHub's August incidents are a good counterexample, which we unpacked in what GitHub's August postmortem breaks about triage order.

    4. Write Hypotheses That Could Be Proven Wrong

    A useful hypothesis predicts something you can check. "It's an attack" is a story. "If it's a volumetric attack, edge request rates rise before the error rate does" is a hypothesis, because one graph can kill it.

    We write each hypothesis with its prediction and its disproving check, then run the cheapest check first. When a check fails, we cross the hypothesis off in the incident document instead of quietly dropping it.

    Pro tip: We ask the person who proposed a hypothesis to name the observation that would make them abandon it. If they can't name one, the hypothesis isn't ready to test.

    5. Map Every Necessary Condition, Not One Cause

    Most software incidents need several conditions at once. Even Wikipedia's root cause analysis entry notes that one or several factors may make up the root cause, and that a causal factor can shape the outcome without being the root cause.

    For each condition, we ask if the incident still would have happened, or would have been smaller, without it. Conditions that pass that test go on the map. This is where the method choice in the next section comes in.

    6. Turn Each Condition Into a Check or a Guardrail

    Every condition on the map gets an owner and a dated action. Some actions prevent the trigger. Others contain the blast radius or shorten detection.

    We prefer guardrails over reminders. "Validate generated config files like user input" survives staff turnover. "Be more careful with grants" does not.

    Which Root Cause Analysis Method Fits Which Kind of Software Incident

    Pick the root cause analysis method by the shape of the incident, not by habit. One linear failure path suits 5 whys, a wide-open search suits a fishbone, and an outage that needed several conditions at once suits a fault tree. Anything that broke right after a change starts with change analysis.

    These four cover most of the root cause analysis techniques we see used in software teams. The table is our decision aid:

    Incident shape Method we reach for Why it fits Where it misleads
    One component, one failure path 5 whys Fast and easy to run alone Stops at the first answer that sounds right
    No strong lead, many candidate causes Fishbone (Ishikawa) diagram Forces breadth across categories Lists candidates but proves none of them
    Several conditions had to be true together Fault tree analysis AND gates show which conditions were jointly necessary Slow, and overkill for a one-line bug
    Broke right after a deploy, flag flip or config push Change analysis Most outages follow a change Blind to latent conditions and slow data growth

    • 5 Whys for One Linear Failure Chain

    The five whys technique came out of Toyota and was described by Taiichi Ohno. You ask "why" of each answer until you reach a cause you can act on.

    We use a 5 why root cause analysis when the failure lives in one place, such as a missing null check or a migration that ran out of order.

    For root cause analysis for software defects contained in one function, it is usually enough. The discipline that makes it work is branching, so when a "why" has two true answers, follow both.

    • Fishbone Root Cause Analysis With Software Categories

    The Ishikawa diagram puts the problem at the fish's head and groups candidate causes along the bones. The classic categories come from manufacturing (manpower, machine, material, method, measurement), which map poorly to software.

    For a fishbone root cause analysis on a service, we use six bones. They are code, configuration, data, dependencies, infrastructure and traffic. If you're deciding when to use 5 whys vs fishbone, our rule is simple. Use the fishbone first when you have no lead, then run 5 whys down each bone that the evidence supports.

    • Fault Tree Analysis for Failures That Needed Several Conditions

    Fault tree analysis was developed at Bell Laboratories in 1962 for the Minuteman I launch control system. It starts from the undesired top event and works down through AND and OR gates.

    The idea we borrow most is the minimal cut set. It is the smallest combination of events that still causes the top event.

    In root cause analysis in software engineering, each event in a minimal cut set is a place a guardrail could have stopped the outage. That turns a postmortem from "find the culprit" into "find the cheapest gate to close."

    • Change Analysis for Anything That Broke After a Change

    Change analysis compares the last known good state with the first known bad state, then explains every difference. We diff deploys, config, flags, grants and schema between the two points.

    It is the fastest method when it applies. It fails quietly when the cause is latent, such as a table that grew past a threshold, because nothing in the change log looks related.

    Pro tip: We keep the change list in the incident document even after a cause is confirmed. Unrelated changes in the window are the first thing a reviewer asks about, and an empty list invites doubt.

    Running Root Cause Analysis on the Cloudflare November 2025 Outage

    Cloudflare's 18 November 2025 outage shows why the method matters. 5 whys points to one fix, while a fault tree shows three separate places the outage could have been stopped. Every fact below comes from Cloudflare's postmortem. The root cause analysis structure is ours.

    Time (UTC) Event
    11:05 Database permissions change begins rolling out to the ClickHouse cluster
    11:20 Network starts failing to deliver core traffic
    11:28 Errors reach customer HTTP traffic
    13:05 Workers KV and Access bypass reduces impact
    14:30 Known good feature file deployed; main impact resolved
    17:06 All services restored

    • What 5 Whys Produces

    A straight 5 whys chain reads cleanly:

    1. Why did requests return 5xx? The FL2 proxy panicked.
    2. Why did it panic? The Bot Management feature file had more than 200 features, its preallocated limit.
    3. Why was the file too big? The generating query returned duplicate rows.
    4. Why the duplicates? The query didn't filter by database name, and the 11:05 change made tables in the r0 database visible.
    5. Why was that possible? The query assumed it would only ever see the default database.

    The fix it suggests is "add a database filter to the query." That fix is correct and far too small.

    • What a Fault Tree Produces

    A fault tree puts "global 5xx errors" at the top, under an AND gate with three branches. All three had to be true:

    • A bad file was generated. This branch is itself an AND condition because the grant change exposed r0, and the query lacked a database filter.
    • The bad file reached the whole network within minutes. Cloudflare says the file is regenerated every five minutes and pushed network-wide, by design, so bot defenses stay current.
    • The proxy treated an oversized file as fatal. The FL2 Rust code hit an unhandled error, Result::unwrap() on an Err value, and the worker thread panicked.

    Each branch is a separate place to stop the next outage. Cloudflare's follow-ups line up with them. They include hardening ingestion of internally generated config files "in the same way we would for user-generated input," adding global kill switches, and reviewing failure modes across core proxy modules.

    • What Change Analysis Would Have Put on the List

    The intermittent pattern had an explanation that only the change list held. The permissions change rolled out gradually, so the five-minute job produced a good or bad file depending on which node ran it.

    That flicker, plus the status page outage, looked like an attack. A change list for the window would have carried the 11:05 grant from the start. We'd have tested the attack hypothesis against it. An attack predicts rising edge traffic, while a bad config file predicts errors that track the file's generation cycle.

    Why 5 Whys Stops Too Early in Distributed Systems

    5 whys stops too early in distributed systems because it follows one branch and ends where the investigator's knowledge ends. Both limits are documented by people who used it seriously, inside and outside software. We still use 5 whys in root cause analysis, but only after the evidence has told us which branches exist.

    • It Follows One Branch When the Incident Has Several

    Each answer has exactly one "why" asked of it, so parallel conditions drop out. Teruyuki Minoura, formerly of Toyota, criticized the technique as too basic to reach root causes at the needed depth, as Wikipedia's five whys entry records.

    Richard Cook's How Complex Systems Fail goes further. Point 7 argues that post-accident attribution to a single root cause is fundamentally wrong, because overt failure needs multiple faults acting together.

    • It Ends Where the Investigator's Knowledge Ends

    You can't ask "why" about a mechanism you don't know exists. If nobody on the call knows the query reads system metadata tables, the chain ends at "the file was too big."

    John Allspaw made the software version of this argument in The Infinite Hows (2014). He asked teams to replace "why?" with "how?" questions. Those questions draw out second stories, the conditions and context around a decision rather than a culprit.

    • Two Investigators Reach Two Different Answers

    Results are not repeatable. Alan J. Card's 2017 paper in BMJ Quality & Safety, The problem with '5 whys', argues that the fifth "why" has an arbitrary depth that is unlikely to match the real root cause.

    In software, we see this as postmortems that name whichever layer the author owns. The database engineer finds a query bug. The platform engineer finds a rollout bug. Both are right, and neither is the whole answer.

    How to Tell When a Root Cause Analysis Is Finished

    A root cause analysis is finished when every necessary condition has evidence, every odd symptom has an explanation, and every condition has an owner with a dated action. "We found a cause" is not the bar. We run this checklist before closing the document:

    Check Passes when
    Evidence Each condition links to a log line, graph, diff or query result
    Necessity Removing any one condition would have prevented or shrunk the incident
    Odd symptoms Intermittency, recoveries and coincidences each have an explanation or are marked unexplained
    Detection The gap between first trigger and first alert is explained
    Actions Every condition has an owner and a date
    Effectiveness A review date is set to confirm the actions worked

    The effectiveness check matters more than it looks. Wikipedia's root cause analysis entry describes scheduling an effectiveness review after corrective actions ship, and reopening the analysis if the problem returns.

    We write the result as a blameless postmortem, so the people closest to the incident keep sharing what they saw. We also keep the dead ends. A postmortem that records only the final answer throws away the search that found it, which we argued in the postmortem records the answer and throws away the search.

    What Root Cause Analysis Software Can and Cannot Do in the Investigation

    Root cause analysis software can collect and correlate the evidence for steps one to four, but the judgment in steps five and six stays with engineers. We think that split is the honest way to evaluate any tool in this category.

    • Where It Saves the Most Time

    Good root cause analysis software pulls logs, deploys, database state and code into one timeline, matches changes to the onset of errors, and proposes hypotheses with the evidence attached. That cuts the slowest part of an investigation, which is usually finding the right data rather than reading it.

    • What It Cannot Do for You

    We'd hold any tool to these limits:

    • See what isn't instrumented. Wikipedia's entry notes that IT analysis is often limited to what has monitoring interfaces.
    • Decide the corrective action. Choosing between a query filter and a validation gate is a design decision with trade-offs.
    • Escape anchoring. A confident first answer from a tool anchors a team the same way a confident engineer does. We covered that risk in why second opinions matter for root cause analysis.

    • How We'd Compare Tools

    If you're comparing tools, our vendor-neutral guide to automated root cause analysis lists the questions we'd ask. We'd also check how a tool verifies its own answers, a topic we explored in what aviation cross-checks teach about verifying AI root cause analysis.

    Start Your Next Root Cause Analysis From the Timeline and the Change List

    Root cause analysis in software goes wrong when a method replaces evidence. We freeze the data, build a UTC timeline, list every change, and test hypotheses that could fail. Only then do we pick 5 whys, a fishbone, a fault tree or change analysis to fit the incident's shape.

    Root cause analysis software helps most with the evidence half. Operate does that part inside your own infrastructure. It reads logs, the database and code, finds the root cause with evidence, and drafts a pull request for an engineer to review.

    Frequently Asked Questions

    In testing, we treat root cause analysis as asking why a defect escaped, not only why it exists. We tag each escaped bug with where it entered (requirements, design, code or environment) and which test layer should have caught it. The pattern across releases shows where tests are missing.

    We look for three skills. Read system evidence fluently, write hypotheses that can be falsified, and run a blameless interview. The interview matters as much as the debugging, because part of the root cause analysis timeline lives in people's heads and only comes out when they feel safe describing what they saw.

    Neither number is a rule. Semco under Ricardo Semler practiced "three whys," according to Wikipedia. In our root cause analysis work we stop when the next answer leaves our control or stops pointing at something we can change. We count evidence, not questions.

    Amazon runs 5 whys inside its Correction of Error process, described in an AWS operations post. It says you may need more than five, insists on asking why any human error was possible, and notes COEs average 6 to 8 action items. We like that root cause analysis stance.

    In Six Sigma, root cause analysis sits in the Analyze step of DMAIC, which means Define, Measure, Analyze, Improve and Control. Teams list candidate causes, often on a fishbone, shortlist three or four, and collect data to validate each. We borrow that validation step when a software cause is still only plausible.

    Share: X LinkedIn
    #Root Cause Analysis
    #Incident Response
    #postmortem
    #SRE
    #Reliability Engineering

    Keep reading