How to Retry Flaky Tests Without Letting a Real Regression Through
We show what retries do to flaky tests and real regressions, how six test runners report a pass on retry, and which failures should fail the build.
Quick Answer
Retrying flaky tests is safe only when a pass on retry is recorded and acted on. Blind retries also wave through intermittent regressions because a bug failing 30% of runs clears a two-retry gate about 97% of the time. The safer setup caps retries at one or two, scopes them to known flaky tests or infrastructure errors, fails the build on new flakes, and uses retries to classify causes.
A flaky test passes and fails on the same code, so its result depends on something other than your change. The common answer is a retry setting that nobody looks at again. The retry turns the build green and the pipeline moves on.
The trouble is that a retry cannot tell a noisy test from a new race condition that fails some of the time.
In this post we'll show the math of what a retry does, how six test runners report a pass on retry, which failures deserve a retry, and a policy you can roll out in one sprint.
What a Retry Does to a Flaky Test and to a Real Regression
A retry makes a known flaky test almost harmless and makes an intermittent regression almost invisible, at the same time. We think you should retry flaky tests only when the pass on retry is recorded and gated, because the retry setting itself cannot tell those two cases apart.
The math is short. If a test fails a given run with probability x, and attempts are independent, a gate with N retries fails only when all N + 1 attempts fail. That probability is x^(N+1).
| Per-run failure rate | What it usually is | Gate catches it, 0 retries | 1 retry | 2 retries |
|---|---|---|---|---|
| 1% | Known flaky test | 1% | 0.01% | 0.0001% |
| 5% | Known flaky test | 5% | 0.25% | 0.0125% |
| 30% | Intermittent regression | 30% | 9% | 2.7% |
| 50% | Intermittent regression | 50% | 25% | 12.5% |
| 100% | Hard regression | 100% | 100% | 100% |
The first row matches the numbers in Mill's post on managing flaky tests in CI. One retry turns a 1% flake into 0.01%, and two retries into 0.0001%. The third row is the one we worry about.
• A Known Flaky Test Stops Blocking Merges
For a test you already know is noisy, two retries take a 5% failure rate down to roughly one build in 8,000. That is the whole appeal of retries, and it is real.
Google's testing team reported in 2016 that about 1.5% of all test runs produced a flaky result, and almost 16% of their tests showed some flakiness. At that scale, a gate without retries blocks almost every change for reasons unrelated to it.
• An Intermittent Regression Slips Through
Say your change adds a race condition in a cache write that loses the update in 3 of every 10 runs. Without retries, CI catches it 30% of the time. With two retries, CI catches it 2.7% of the time, so the change merges about 97% of the time.
That is a real software regression that now looks exactly like a flaky test. Google's post found about 84% of observed pass-to-fail transitions in post-submit testing involved a flaky test, and it notes that developers sometimes dismiss a failure as flaky and later learn it was caused by the code.
We also treat the independence assumption as too generous. GitHub's engineering team points out that a test failing because of a leap year still fails when rerun two minutes later. Correlated failures make the table above look better for retries than reality is.
• A Hard Failure Costs Three Runs and Tells You Nothing New
When a change breaks a test outright, every retry fails too. You pay for three runs and learn nothing a single run would not have told you. Mill's post makes the same point about blanket retries. An actually failed test runs three times before giving up.
The cost grows when a change breaks hundreds of tests at once. That is why two runners in the next section ship a brake that stops retrying after too many failures.
Pro tip: We compute the table above for our own suite before choosing a retry count. If the flaky tests you care about fail under 2% of runs, one retry already removes almost all the noise, and a second retry mostly buys you blindness to intermittent regressions.
How Six Test Runners Report a Flaky Test That Passes on Retry
Most test runners can record a pass on retry as a flaky result, but only three of the six we checked document a switch that fails the build on it. They are Playwright Test, Maven Surefire and the Gradle test-retry plugin.
With pytest-rerunfailures, Jest and Bazel, we found no such switch, so gating on a flaky test is your job.
We checked each runner's official documentation on 26 September 2026.
| Runner | Retry setting | Default | How a pass on retry is reported | Documented way to fail the build on it |
|---|---|---|---|---|
| Playwright Test | retries, --retries=N |
No retries | Categorized as flaky, counted as "1 flaky" in the summary |
failOnFlakyTests or --fail-on-flaky-tests (added in v1.52) |
| Maven Surefire | rerunFailingTestsCount |
Ignored at 0 or below | Counted as flaky, Flakes: 1 in the summary, flakyFailure in the XML report |
failOnFlakeCount (since 3.0.0-M6) |
| Gradle test-retry plugin | retry { maxRetries } |
maxRetries is 0 |
Each execution appears in the XML and HTML reports; the task passes | failOnPassedAfterRetry = true |
| pytest-rerunfailures | --reruns N or @pytest.mark.flaky(reruns=N) |
Opt-in | RERUN lines in the summary; earlier tracebacks dropped unless --rerun-show-tracebacks |
None found in the README |
| Jest | jest.retryTimes(n, options) |
Opt-in per file or block | Errors from failed attempts are logged only with logErrorsBeforeRetry: true |
None found on the docs page |
| Bazel | --flaky_test_attempts |
default gives 1 attempt, or 3 for rules marked flaky |
Marked FLAKY in the test summary |
None found in the flag docs |
1. Playwright Test Marks the Test as Flaky and Can Fail the Run
Playwright retries are off by default. With retries on, a test that fails and then passes lands in the flaky category, and the summary line counts it. The failOnFlakyTests option makes the run exit with an error when any test is flaky.
import { defineConfig } from '@playwright/test';
export default defineConfig({
retries: process.env.CI ? 2 : 0,
failOnFlakyTests: !!process.env.CI,
});
With both settings on, retries become a detector. The run still fails, but the report tells you the test passed on a second attempt. We run specs that are already quarantined in a separate CI job without the flag.
2. Maven Surefire Counts Flakes and Keeps the First Failure in XML
Surefire reruns failing tests when rerunFailingTestsCount is above 0, for JUnit 4.12+, JUnit 5 and TestNG. A test that passes on a rerun is counted as flaky, the build succeeds, and the summary prints a Flakes count. The XML report keeps each failed attempt inside a flakyFailure element.
<plugin>
<groupId>org.apache.maven.plugins</groupId>
<artifactId>maven-surefire-plugin</artifactId>
<configuration>
<rerunFailingTestsCount>2</rerunFailingTestsCount>
<failOnFlakeCount>1</failOnFlakeCount>
</configuration>
</plugin>
The Surefire docs say failOnFlakeCount fails the build when more than the specified number of tests are flaky. We keep that number small and raise it only with a written reason.
3. The Gradle Test Retry Plugin Lets Flakes Pass Unless You Say Otherwise
The JUnit-friendly Gradle plugin (org.gradle.test-retry) retries failed tests after the whole task runs. By default, tests that pass on retry do not fail the task. Its README says, "Retrying tests alone is not a viable flaky test mitigation strategy."
test {
retry {
maxRetries.set(2)
maxFailures.set(20)
failOnPassedAfterRetry.set(true)
}
}
maxFailures stops retrying once a round has that many failures, which the README ties to cases such as a full disk or a missing database. For a Gradle test retry setup on CI, this is the brake we always set.
4. Pytest-Rerunfailures Hides Earlier Tracebacks by Default
For pytest rerun failed tests support, the usual plugin is pytest-rerunfailures. Its README says only the final attempt of a flaky test produces a traceback unless you pass --rerun-show-tracebacks. It also lets you rerun only failures that match an exception pattern.
pytest --reruns 2 --reruns-delay 1 \
--only-rerun ConnectionError --only-rerun TimeoutError \
--max-suite-reruns 10 --rerun-show-tracebacks
--max-suite-reruns caps reruns across the whole suite. We did not find a documented option that fails the run on a pass after rerun, so we parse the RERUN lines in CI and fail the job ourselves.
5. Jest Retries Quietly Unless You Ask It to Log
jest.retryTimes() reruns failed tests up to the count you set. The errors that caused earlier failures reach the console only when logErrorsBeforeRetry is on. By default, Jest retries after the other tests in the file finish, unless you set retryImmediately.
jest.retryTimes(2, { logErrorsBeforeRetry: true });
We found no flaky status on the Jest object docs page, so we treat retried Jest suites as unreported until our own reporter proves otherwise.
6. Bazel Flags FLAKY in the Summary and Retries Rules Marked Flaky
Bazel's --flaky_test_attempts retries each test up to the given number. Tests that needed more than one attempt are marked FLAKY in the test summary. With the value default, regular tests get one attempt and tests marked flaky on their rule get three.
bazel test //... --flaky_test_attempts=default
That default is a sensible shape because retries apply only to tests someone has explicitly labeled, and everything else fails on the first attempt.
Which Flaky Test Failures Deserve an Automatic Retry
Retry a failure automatically only when it matches a known flaky test or a known infrastructure error, and fail the build on everything else, including any new flaky test. This is how to handle flaky tests without teaching the gate to ignore your own bugs.
| Symptom in CI | Retry? | Gate rule | Why we choose it |
|---|---|---|---|
| Test is on the flaky test quarantine list, with an owner and a ticket | Yes, up to 2 | Allow a pass on retry, log it to the flake tracker | Known noise, already assigned to someone |
| Failure matches a known infrastructure error, such as a refused connection to a test container | Yes, once | Filter by exception, for example pytest --only-rerun |
The code under test never ran its assertion |
| A test with no flaky history passes on retry in a pull request | Yes, to classify | Fail the build | New nondeterminism usually arrives with a change |
| Dozens of tests fail in the same run | No | Stop retrying with Gradle maxFailures or pytest --max-suite-reruns |
One shared cause, retries only add runtime |
| Failure depends on the date, time zone or clock | No | Fix it or quarantine it | A rerun minutes later sees the same clock |
| Test fails on every attempt | No | Fail | Hard regression, retries cost time |
The row that matters most is the third one. A test that has never flaked and suddenly passes on retry inside your pull request is the cheapest regression you will ever find. The author knows the change, the diff is small, and the failure is fresh.
A 2020 Hacker News thread makes the same argument from the other side. A commenter asks who checks that the software under test, not the test, is the flaky part. We agree, and failing on new flakes is how you make someone check.
For the process after a test lands on the list, our flaky test runbook covers quarantine with exit criteria and how many green runs prove a fix.
Pro tip: We keep the quarantine list in the repository next to the tests, not in a CI setting. A retry exception then shows up in code review, with an author and a date, instead of living in a dashboard nobody opens.
How to Use Retries for Flaky Test Detection Instead of Hiding Failures
A retry is the cheapest experiment you can run on a failure. Change one condition, rerun, and record what changed. Used that way, retries do flaky test detection and point at the cause, instead of only turning the build green.
GitHub's engineering team described this in December 2020. Their CI had detected flaky failures by comparing builds on the same git tree hash and by retrying in the same build. Those two approaches identified only 25% of flaky failures.
They then reran each failed test three times, each attempt targeting one common cause. That identified 90% of flaky failures, and commits with a flaky build dropped from about 1 in 11 to 1 in 200, which they called an 18x improvement. The three attempts map cleanly to these causes.
• Same Process and Host Points to Randomness or a Race
If the test passes when rerun under the same conditions, GitHub reads it as randomness in the code or a race condition. We look at shared caches, async waits and anything that depends on thread scheduling.
• Same Process With the Clock Shifted Forward Points to Time
If the test passes only when simulated time moves forward, the test assumes something about time. GitHub's example is assuming February has 28 days.
• A Different Host and Database Points to Order or Shared State
If the test passes on a fresh host but fails on both local retries, the likely cause is test order dependence or other shared state. In continuous integration runners that reuse workers, this is the class we suspect first when a failure gets blamed on "the network".
You can build a smaller version of this in any runner. Playwright exposes testInfo.retry, so a fixture can change one condition on a retry, such as turning on tracing. The Playwright docs show clearing server-side caches on retry; if you do that, log it, because "passes only after cleanup" points straight at shared state.
GitHub also found that flakiness is uneven because most flaky tests failed fewer than ten times, and only 0.4% failed 100 times or more. We track a flaky test rate per test (flaky results divided by runs) and send the top of that list to owners, not the long tail.
A Five-Step Retry Policy for Flaky Tests You Can Ship This Sprint
A retry policy that protects your gate needs five changes. Find existing retries, make retried passes visible, cap and scope retries, fail on new flakes, and review the list weekly. None of them needs a new tool, and each one is a small pull request.
1. Find Every Retry Setting Already in Your Repositories
Retry settings accumulate in config files over the years. We start with a search across the repository.
grep -rnE "retries|rerunFailingTestsCount|maxRetries|--reruns|retryTimes|flaky_test_attempts" .
Add any CI-level "rerun failed jobs" habits to the list too. A rerun of a whole job replaces the record of the first failure, which is harder to audit than a test-level retry.
2. Make a Pass on Retry Visible in the Test Report
Turn on the reporting each runner offers, including the Playwright flaky category, Surefire's flakyFailure XML, one report entry per execution with Gradle, --rerun-show-tracebacks for pytest and logErrorsBeforeRetry for Jest. If nobody can see the first failure, the retry already hid it.
3. Cap Retries at Two and Scope Them to Known Cases
Use the table in the first section to pick the count. Scope retries to the quarantine list and to infrastructure exceptions, and set the suite-wide brake (maxFailures or --max-suite-reruns) so a broken build does not triple its runtime.
4. Fail the Build on Any New Flaky Test
Set failOnFlakyTests, failOnFlakeCount or failOnPassedAfterRetry on the main merge gate. For runners without a switch, fail the job when the report shows a retried pass for a test that is not on the list.
5. Review the Flaky Test List Weekly With Named Owners
Sort by flaky test rate, assign the top entries, and remove tests from the list only after a fix is verified.
We treat recurring flaky tests the same way we treat recurring production errors. Each one gets an owner, evidence and a root cause.
For the view that flakes often expose real system nondeterminism, see our post on flaky tests as system bugs, and for pipelines that pass while doing nothing, our taxonomy of green no-op steps.
Pro tip: We add the retry count and the flaky test rate to the same weekly review as the incident list. When a test that flakes in CI shows up near a production incident in the same code path, it moves to the top.
Make Every Retry Leave a Record Before You Trust a Green Build
We think retries are fine for flaky tests you already know about, and dangerous for everything else. Cap them at two, scope them to known cases, and fail the build on any new flaky test.
Keep every retried pass visible so your regression testing stays honest. When a flake turns out to be a real bug, Operate reads logs, database and code to find out why, then drafts the fix as a pull request for a human to review.
Frequently Asked Questions
For Cucumber on the JVM, we let the build tool do it. The Maven Surefire docs say Cucumber JVM 2.0.0 and higher supports rerunning scenarios through the JUnit Platform, so rerunFailingTestsCount covers scenarios the same way it covers JUnit tests. We still fail the build on any flaky test scenario outside the quarantine list.
We run it many times in a row on the same code and count failures. Playwright has --repeat-each <N>, and Bazel's flag docs point to --runs_per_test for the same job. We run the loop on a CI runner, not a laptop, because a flaky test often depends on CPU contention your laptop does not have.
We start from the first failing attempt and find the condition that differed. Then we fix the code or the test so that condition stops mattering, such as waiting on a real signal instead of a sleep. We call a flaky test fixed only after hundreds of clean repeated runs on CI.
A failed test case is a test whose run did not meet its expected result. Maven Surefire's summary line counts failures and errors separately, next to skips and flakes. We read the first failure's stack trace first, because later attempts on the same flaky test often fail differently or not at all.
We give every failing test an owner within a day, attach the logs and the commit it failed on, and reproduce it with the same seed and test order. If the owner cannot reproduce it, the test goes on the flaky test list with a ticket instead of being deleted or silently skipped.


