Back to blog
    Reliability Engineering
    By Jai Jalan14 min read

    What Is the Impact of Flaky Tests on Your Team (and How to Measure It)

    We break the impact of flaky tests into four cost lines, explain why published numbers disagree, and show how to measure your own team's cost in two weeks.

    Quick Answer

    The impact of flaky tests shows up in four places, including investigation time on every red build, developer time spent waiting on reruns, CI compute for the reruns, and regressions that ship because a real failure looked flaky. The one published industrial cost study found reruns cheap and human investigation expensive. Two weeks of measurement from existing CI data, triage logs and postmortems gives a team its own number instead of a vendor estimate.

    A flaky test is a test that passes on one run and fails on the next with no code change. Once that happens often, a red build stops telling you anything. Every engineering leader has heard that flaky tests are expensive, yet few can say how expensive for their own team.

    When we asked Perplexity on 2 October 2026 what flaky tests cost, every source it cited was a vendor blog. The numbers ranged from tens of thousands to millions of dollars a year. We think that range is a symptom, not an answer.

    In this post we'll cover the four lines of flaky test cost, why published numbers disagree, a two-week way to measure your own, a worked example, and which cost to cut first.

    What Flaky Tests Cost a Team, Line by Line

    A flaky test costs a team in four separate lines, including people investigating red builds, people waiting on reruns, machines running reruns, and real regressions dismissed as noise. We price each line separately because they move in different directions when you change policy.

    Ask a team "why are flaky tests bad?" and many answers collapse these into one complaint about wasted time. Many cost estimates do the same with one number. That hides the trade-off that matters. Turning on automatic reruns cuts the first line, adds a little to the third, and can grow the fourth.

    Cost line What it measures Published primary figure Measured or assumed
    Investigation Minutes a person spends deciding if a red build is real About 28 minutes to manually triage each test failure, from a developer survey at Slack Survey estimate
    Investigation and repair Share of developer time spent on flaky tests At least 2.5% of productive developer time on a project with about 30 developers (Leinen et al., TU Munich and CQSE) Measured from booked time
    Waiting on reruns Developer hours lost to rerunning builds Over 150,000 hours a year across Jira backend reruns (Atlassian) Reported by the company, method not detailed
    CI compute Runner minutes spent on reruns About $3 a month for automatic reruns in the same 30-developer project (Leinen et al.) Measured
    Escaped regressions Real failures dismissed as flaky No dollar figure published. Parry et al. found developers who see flaky tests more often may be more likely to ignore genuine failures Survey finding, not priced

    • Investigation Time on Every Red Build

    This line usually dominates. Someone has to open the log, compare it with the last green run and decide. The Leinen et al. case study split the 2.5% into 1.1% of total time investigating, 1.3% repairing flaky tests and 0.1% building monitoring tools.

    We find investigation is also the line teams often undercount. Nobody logs "15 minutes deciding the build was fine", so it never reaches a sprint report.

    • Developer Wait Time on Reruns

    A rerun is cheap for the machine and expensive for the person staring at it. GitHub reported that 1 in 11 commits in its monolith once had a red build caused by a flaky test. Anyone deploying a handful of commits was likely to hit one.

    Wait time is material where a red build blocks a merge or a deploy. It matters less where reruns happen in the background.

    • CI Compute for Reruns

    This is the line many calculators lead with, but the primary data points elsewhere. The same case study found a failed pipeline cost $5.67 in manual investigation, while an automatic rerun cost a small fraction of that.

    At GitHub's list price of $0.006 per minute for a Linux 2-core hosted runner (checked 2 October 2026), a 12-minute rerun costs about 7 cents. We would not build a business case on this line alone.

    • Regressions That Ship Because the Failure Looked Flaky

    This is the line nobody prices, and we think it is the one leaders should worry about. Google's testing team wrote in 2016 that about 84% of pass-to-fail transitions it observed involved a flaky test, and that developers sometimes dismissed a real failure as flaky.

    Once a team learns that red usually means noise, the one real software regression in the batch gets the same shrug. That is alarm fatigue applied to continuous integration.

    Pro tip: We price the escaped-regression line from the team's own postmortems, not from a formula. One incident where a failing check was rerun until green is worth more in a budget meeting than any per-minute estimate.

    Why Published Flaky Test Cost Numbers Disagree So Widely

    Published numbers on the cost of flaky tests disagree because they measure different lines, at companies of very different sizes, and many vendor figures are models built on assumed inputs rather than measurements. We read the top-ranking cost pages on 2 October 2026, and the spread follows directly from those choices.

    The clearest example is reruns. The TU Munich and CQSE study concluded that rerunning tests was negligible and inexpensive, and the team moved effort from investigation to automatic reruns of up to five attempts. Atlassian, writing in December 2025, reported that reruns caused by flaky tests in its Jira backend repository waste over 150,000 developer hours a year.

    Both can be true. One counted runner cost. The other counted people waiting. Here is how the main sources compare.

    Source Company size or scope What was counted How it was produced
    Leinen et al. One commercial project, about 30 developers, about 1M lines of code Booked investigation and repair time, rerun compute Measured from Jira, GitLab CI and Slack data
    Google Testing Blog, 2016 Google-wide test corpus 1.5% of test runs flaky; almost 16% of tests show some flakiness Measured, no dollar cost
    GitHub, 2020 GitHub monolith Share of commits with a flaky red build was about 9%, cut to under 0.5% Measured
    Slack, 2020 Mobile codebases, 120+ developers Triage minutes per failure, and test job failures cut from 57% to under 5% of failing builds Survey plus build data
    Atlassian, 2025 Jira repositories Up to 21% of Jira Frontend master build failures; developer hours on reruns Reported totals
    Vendor cost calculators Hypothetical teams Compute plus assumed engineer hours Modeled. For example, one 2026 model assumes a 12% flake rate and $0.008 per CI minute

    Here's the thing about scaling a single-company percentage to your headcount. The 2.5% in Leinen et al. came from one team's process, which already used reruns and a defined triage habit. Your number could be lower or several times higher.

    That is why we would not take a published flaky test cost to a board. We would take our own.

    How to Measure the Impact of Flaky Tests on Your Own Team in Two Weeks

    You can measure the impact of flaky tests with data most teams already have, including CI run history, a short triage log and your postmortems. We run it for two weeks, because one week is too easy to dismiss as unusual and a month is long enough that people stop logging.

    1. Count Flaky Failures From CI Data You Already Have

    Count builds that failed and then passed on the same commit. This is how to identify flaky tests in CI without buying a tool first, and it gives you a flaky test rate per week.

    On GitHub Actions, every workflow run carries a run_attempt field and a head_sha, per the workflow runs REST API. A run that succeeded on attempt two or later, on an unchanged SHA, is a flaky failure candidate.

    # Workflow runs that needed more than one attempt (first page of results)
    gh api "repos/OWNER/REPO/actions/runs?per_page=100" \
      --jq '.workflow_runs[] | select(.run_attempt > 1) | [.id, .head_sha, .run_attempt, .conclusion] | @tsv'
    

    Test runners give a finer, per-test signal. Maven Surefire with rerunFailingTestsCount set writes flakyFailure elements into the XML report for a test that failed and then passed on a rerun (Surefire docs). Most test runners and build tools now expose a similar label.

    Expect noise. A rerun after a runner outage is not a flaky test, so we tag infrastructure failures separately.

    2. Log Triage Minutes for Two Weeks

    Ask whoever looks at red builds to log one line per failure with the build link, minutes spent and verdict. Slack's 28-minute figure came from a survey. Yours should come from a log.

    Keep it light. A shared form with three fields gets filled in. A ticket per failure does not.

    3. Price the Reruns at Your Runner Rate

    Multiply rerun minutes by your runner rate. For hosted GitHub runners, the per-minute price table lists $0.006 for Linux 2-core and $0.062 for macOS (checked 2 October 2026). Self-hosted runners need your own cost per minute.

    We include this line because leaders ask for it, not because we expect it to be large.

    4. Measure How Long Authors Wait

    For pull requests where the red build blocked a merge, take the time from the first failure to the green rerun. That is waiting time, even if the author switched tasks, because the switch has its own cost.

    Pro tip: We split waiting time by queue position. A rerun that lands in a busy runner queue at 4 p.m. can cost more waiting than the original job took to run.

    5. Check Postmortems for Failures Dismissed as Flaky

    Search the last six to twelve months of postmortems and incident timelines for "flaky", "rerun" and "retried". Each hit where a failing test predicted the incident goes into the escaped-regression line. Often the count is zero. In our view, a single hit usually outweighs the other three lines combined.

    A Worked Example of Flaky Test Cost for a 40-Engineer Team

    A worked example shows why investigation dominates for many teams and compute barely registers. The inputs below are a hypothetical team, not a real company, with two published reference points borrowed for scale.

    Say a 40-engineer team runs 400 pipelines a week. We assume 5% of them hit a flaky failure, close to the 4.8% of test suite executions with at least one flaky failure in the TU Munich and CQSE study. Each pipeline takes 12 minutes on a Linux 2-core hosted runner.

    Line Calculation Per week Per 48-week year
    Flaky red builds 400 pipelines × 5% 20 builds 960 builds
    Investigation 20 × 28 minutes (Slack's survey figure) 560 minutes, about 9.3 hours About 448 hours
    Waiting on reruns 20 × 12 minutes 240 minutes, 4 hours About 192 hours
    Rerun compute 20 × 12 minutes × $0.006 $1.44 About $69
    Escaped regressions Count from postmortems Not modeled Not modeled

    At a hypothetical loaded rate of $100 an hour, investigation alone is about $44,800 a year and compute is about $69. The ratio, not the total, is the useful finding. Your inputs will change the totals, but we'd be surprised if the order of the lines changed.

    What this example cannot show is the escaped-regression line. We leave it blank rather than guess, and we'd advise you to do the same in front of leadership.

    Which Flaky Test Cost to Cut First, and When Reruns Are the Cheaper Answer

    Cut the largest line in your own measurement first, and treat automatic reruns as a cost decision rather than a quality failure. When investigation dominates, a logged automatic rerun is usually a low-cost first move, as long as every pass-on-retry is recorded.

    If your largest line is Cut it with Watch for
    Investigation Automatic reruns with a logged flaky verdict, as CQSE and GitHub did Retries hiding real intermittent bugs
    Waiting on reruns Rerun only the failed job, and move known-flaky suites off the merge path Suites that never come back onto the gate
    Compute Smaller rerun scope, cheaper runners Rarely worth a project on its own
    Escaped regressions Quarantine with a named owner and a fix-by date, plus regression testing that stays on the gate Quarantine lists that only grow

    Reruns buy back investigation time. They do not fix the cause, and many flaky tests trace to a race condition in the product rather than in the test. We wrote about that in why flaky tests are system bugs, not noise, and about configuring retries so a regression cannot slip through.

    The opportunity cost argument also has limits. Not every minute saved becomes feature work, and the TU Munich and CQSE authors scoped their findings to their own context. We'd present the saving as hours returned to the team, not as revenue.

    Pro tip: We make every automatic rerun write the test name and the word "flaky" to the build log. The rerun saves investigation time today and still feeds next sprint's fix list.

    Read more on handling a single flaky test as an incident in our flaky test runbook.

    When a pipeline goes green and still ships nothing useful, the cost moves somewhere else entirely. Our taxonomy of green no-op pipeline steps covers that failure class.

    Measure the Four Flaky Test Costs Before You Fund a Fix

    The impact of flaky tests in CI is real, but the number that justifies a fix is your own, measured across investigation, waiting, compute and escaped regressions. The primary data points to people's time as the dominant line and compute as a much smaller one.

    We'd spend two weeks counting from CI data, a triage log and postmortems, then cut the largest line first. For flaky tests that keep coming back, Operate reads CI logs, code and infrastructure context to find the root cause and drafts a patch for an engineer to review.

    Frequently Asked Questions

    We make the failure reproducible first. We run the test in a loop, in random order and under load, then compare a failing run with a passing one. Shared state, timing assumptions and external calls are the usual causes. We fix the cause, never by adding a sleep.

    We replace fixed sleeps with explicit waits on the condition the test needs. We treat the race between browser state and test code as a common source of flaky tests. We also avoid mixing implicit and explicit waits, because that can produce unpredictable timeouts.

    We set a retry count in the Playwright config and read the report. Playwright marks a test "flaky" when it fails on the first run and passes on a retry, so each flaky test is listed by name. We then fix the cause, usually a timing assumption or shared test data.

    It means Nx saw the same task fail and then succeed with an identical hash of inputs. Because the hash pins the inputs, a failure after a code change is not counted. We treat the message as a lead because Nx Cloud links the failed and successful attempts so we can compare them.

    We define it as the loop of detecting each flaky test, tracking it with an owner, quarantining it from the merge gate when needed, and returning it once fixed. The part teams often skip is the return step, so quarantine lists keep growing. We review ours monthly.

    Share: X LinkedIn
    #Reliability Engineering
    #ci-cd
    #engineering management

    Keep reading