Back to blog
    Reliability Engineering
    By Jai Jalan12 min read

    What Is a Good MTTR for a Software Team (and How to Set Your Own Target)

    What a good MTTR looks like for software teams, using DORA's recovery bands and our target from the error budget and incident rate.

    Quick Answer

    A good mean time to resolution for a software team comes from its own error budget and incident rate, not a universal number. DORA's 2024 survey puts elite teams under one hour to recover from a failed deployment. The practical method divides the quarterly error budget by expected incidents, sets a target per severity level and reports percentiles instead of one average.

    Mean time to resolution (MTTR) is the average time a team needs to close out an incident, from the moment it starts to the moment the fix is confirmed. When we searched Google for "what is a good MTTR" on 28 September 2026, the top result said an ideal score is under five hours.

    That page, from LLumin, is written for maintenance teams fixing plant equipment. For a software team, the right number comes from what your users can tolerate, not from a factory floor.

    We understand why leaders ask. A board slide needs a target, and "it depends" is not a target. In this post we'll cover the most cited public software benchmark, which MTTR to measure, how to derive your own target from your error budget, and how to report it so the number does not mislead anyone.

    What a Good MTTR Looks Like for Software Teams, According to DORA

    The most widely cited public MTTR benchmark for software teams comes from DORA, and it says elite teams recover from a failed deployment in under one hour. Most other teams in the 2024 survey took up to a day, and the lowest cluster took between one week and one month.

    The 2024 Accelerate State of DevOps Report grouped respondents into four clusters using cluster analysis. Here is the recovery-time column, with the share of respondents in each cluster:

    Performance level Failed deployment recovery time Share of respondents
    Elite Less than one hour 19%
    High Less than one day 22%
    Medium Less than one day 35%
    Low Between one week and one month 25%

    The same report says elite performers recover from failed deployments 2,293 times faster than low performers. When people talk about MTTR DORA benchmarks, this is the table they mean. The gap between clusters is so wide that a team's cluster says more than its exact minutes.

    Inside a company, MTTR in software engineering usually means mean time to resolution across every kind of incident. For that wider number, we have not found a public dataset of the same quality. We read three limits into the DORA table before anyone puts it on a slide.

    • It Measures Deployment Failures, Not Every Incident

    DORA changed the metric in 2023. What it used to call mean time to recover became failed deployment recovery time, because the old definition mixed change-caused failures with outside events such as a data center outage.

    So MTTR in DevOps conversations now usually means recovery from a bad deploy. A cloud provider outage, an expired certificate or a slow memory leak is not in this benchmark, and those are often the long incidents.

    • The Clusters Come From Survey Answers

    Respondents picked a range for the primary application or service they work on. DORA says it does not set the levels in advance and lets them emerge from the answers each year. We treat the bands as a rough MTTR benchmark, not as audited measurements.

    • DORA Puts Improvement Ahead of the Level

    The report's own advice is that the best teams "achieve elite improvement, not necessarily elite performance." We agree with it. A team that moves from the low band to the medium band in two quarters is doing harder work than an elite team standing still.

    Pick Which Mean Time to Resolution You Mean Before Anyone Sets a Target

    A target means nothing until the team agrees where the clock starts and stops. Atlassian's incident metrics guide points out that the R in MTTR can stand for repair, recovery, respond or resolve, and each one measures a different stretch of the incident.

    Metric Clock starts Clock stops What it tells a leader
    Mean time to repair Repair work begins System is fully working, testing included Hands-on fix speed
    Mean time to respond First alert System is working again Speed once the team knows
    Mean time to recovery System fails System is fully working again Customer-facing downtime
    Mean time to resolution System fails Fix confirmed and recurrence prevented Full cost of the incident

    The MTTR calculation is the same for all four. Divide total minutes by the number of incidents. Only the start and stop points change, and they change the answer a lot.

    Mean time to resolution, as Atlassian defines it, includes the work that keeps the failure from happening again. That follow-up can run for weeks after users stop noticing anything.

    We set the headline target on the recovery clock, because it matches what customers feel and what the error budget counts. We still report full resolution as a second number, so prevention work does not quietly disappear.

    Pro tip: We write the start and stop events into the incident template itself ("impact started", "impact ended", "follow-up closed"). If the timestamps depend on memory after the fact, every MTTR number built on them drifts.

    How to Set Your Own MTTR Target From Your Error Budget

    Your MTTR target should be the average incident length your error budget can absorb at your real incident rate. That turns "be faster" into arithmetic a board can check.

    The Google SRE book describes a quarterly error budget. The gap between the service-level objective and 100% is how much unreliability a service may spend in a quarter. We use that error budget as the ceiling for incident minutes. Here is how to set MTTR target numbers with it, one step at a time.

    1. Convert the Service-Level Objective Into Minutes per Quarter

    A 90-day quarter has 129,600 minutes. A 99.9% availability service-level objective leaves 0.1% of that as error budget, which is 129.6 minutes. At 99.5% the budget is 648 minutes, and at 99.95% it is 64.8 minutes.

    If your SLO counts good requests rather than good minutes, a partial outage spends only part of the budget. We still start with time, because incidents are logged in time and leaders think in time.

    2. Count Customer-Impacting Incidents From the Last Four Quarters

    Use your own incident history, not a guess. Count only incidents that burned budget, meaning users saw errors or latency beyond the SLO.

    Take the quarterly average and round it up. We assume next quarter will be no calmer than the last one, because new services tend to add incidents before they remove any.

    3. Divide the Budget by the Incident Count

    The result is the longest average incident you can afford. That is your mean time to resolution target on the recovery clock.

    SLO Error budget per quarter 3 incidents 6 incidents 12 incidents
    99.5% 648 min 216 min 108 min 54 min
    99.9% 129.6 min 43.2 min 21.6 min 10.8 min
    99.95% 64.8 min 21.6 min 10.8 min 5.4 min

    Say a team runs its checkout API at 99.9% and logged six customer-impacting incidents per quarter last year. Its average incident has to end within about 21 minutes, or the SLO breaks before any other risk is counted.

    4. Keep Part of the Budget for Everything That Is Not an Incident

    We never plan to spend the whole budget on incidents. Slow rollouts, planned maintenance and short blips that never get a ticket all draw from the same pool.

    We plan incidents against half the budget, which halves every number in the table above. The checkout team from the example would then aim for about 11 minutes per incident, and that number starts a very different conversation.

    5. Check the Result Against What Your Team Can Actually Do

    A target of 5.4 minutes per incident is not an MTTR goal a human team can meet. Detection, paging and an engineer opening a laptop can take longer than that on their own.

    When the arithmetic gives a number your team cannot reach, the fix is fewer incidents or a looser SLO, not a faster on-call rotation. We think this is the most useful output of the whole exercise, because it shows leadership which lever to pull.

    We then compare the result with the DORA bands. A target under one hour asks for elite-level recovery. A target of several hours sits comfortably in the high and medium bands, where most teams in the 2024 survey landed.

    Pro tip: We run this division for every service that has an SLO, then sort services by the gap between target and actual. The service with the widest gap is where the next quarter's reliability work should go.

    Why One Average MTTR Misleads Leadership Every Quarter

    A single average hides the incidents that matter most, because a few very long outages skew incident durations. Google's report on Incident Metrics in SRE by Štěpán Davidovič tested this with a Monte Carlo method simulation. It found MTTR-style statistics "poorly suited for decision making or trend analysis" for production incidents.

    In practice, one six-hour outage can move a quarter's mean time to resolution more than twenty quick fixes. We looked at how wide that gap gets in real data in our study of median and mean MTTR across 1.8 million outages.

    For a board slide, we report three numbers per severity each quarter instead of one mean:

    Number What it shows Why we report it
    Median (p50) recovery time The typical incident Stable, hard to move with one outlier
    90th percentile (p90) recovery time The long tail Shows the incidents customers remember
    Incident count How often the clock started A flat MTTR with twice the incidents is a worse quarter

    A percentile needs enough incidents to mean anything. With five incidents in a quarter, p90 is close to the single worst incident, and we say so on the slide rather than pretend to precision.

    Set a Separate Mean Time to Resolution Target for Each Severity Level

    A SEV1 outage and a SEV3 bug should never share one target, because they spend different amounts of budget and wake different people. We tie each incident severity to the clock that matters for it and to what it costs.

    Severity Typical impact Clock we target How we set the number
    SEV1 Full outage or data at risk Recovery Error budget divided by expected SEV1 count
    SEV2 Major feature degraded Recovery Budget left after SEV1, divided by SEV2 count
    SEV3 Minor bug with a workaround Resolution Business-hours target agreed with product

    Where each severity line sits is a decision of its own. We've laid out one way to make it when facts are still unclear in our incident severity levels decision procedure.

    Here is what does not work. Copying a severity table from a vendor blog. When we asked Perplexity "what is a good MTTR for a 50 person software engineering team" on 28 September 2026, it gave a SEV1 target of 15 to 60 minutes. The answer cited blog posts and did not name a dataset behind that range.

    Where the Minutes Go Once an MTTR Target Exists

    Once a target exists, the next question is which stage of the incident eats the time. We split every incident into detect, acknowledge, diagnose, mitigate and verify, and we timestamp each stage in the incident record.

    Detection and paging are usually measured already, because alerting tools stamp them. Diagnosis is often the stage with no timestamp at all, which is why we look there first. Our piece on why diagnosis is the real MTTR bottleneck walks through how to instrument it.

    With stage timestamps in place, cutting mean time to resolution becomes a question about the slowest stage instead of the whole incident. That is a question a team can act on in a sprint.

    A target also invites gaming. Teams under pressure close incidents early and reopen them as new ones, which makes the mean look better and the service worse. We pair every MTTR target with a recurrence count, the approach we explain in our case against ticket closure rate.

    Set Your MTTR Target This Quarter With Three Inputs

    A good mean time to resolution is the one your error budget can absorb at your real incident rate. It is a number you derive, not one you copy.

    We'd take your SLO, your last four quarters of customer-impacting incidents and the DORA bands, then set one recovery target per severity and report p50, p90 and count. If the math asks for minutes nobody can hit, cut incidents first.

    Operate reads your logs, database and code to find out why something broke, then opens a draft PR for a human to review.

    Frequently Asked Questions

    Many teams use it as one, and we think it works best as a lagging indicator. DORA's own metrics guide lists "setting metrics as a goal" as a pitfall and cites Goodhart's law. We track mean time to resolution as a trend and never rank teams on it.

    An SLA is a contractual promise to a customer, while MTTR is an average across many incidents. In our view they answer different questions. A team can hit its mean time to resolution target and still breach an SLA, because the SLA applies to each incident on its own.

    We put the start and end timestamps in two columns, subtract them, and multiply by 1,440 to get minutes. AVERAGE gives the mean, MEDIAN gives p50, and PERCENTILE.INC with 0.9 gives p90. We filter by severity before calculating, never after.

    MTBF is the average time between failures, so it measures how often things break, while MTTR measures how long each failure lasts. We watch both, because availability depends on the pair. Atlassian's incident metrics guide gives the formula as MTBF divided by the sum of MTBF and MTTR.

    DORA's 2024 report gives a reference point for software deployments. The elite cluster had a 5% change fail rate, high 20%, medium 10% and low 40%. We treat a rate as acceptable when the failures it causes still fit inside the error budget.

    Share: X LinkedIn
    #MTTR
    #Incident Response
    #Reliability Engineering
    #SRE

    Keep reading