October DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsSlow PC?RecommendedPC slow today? Run a repair scan before it gets worseResolve common Windows issues and optimize system performance.Scan NowOctober DealsAmazon USDeal season is back - check today's better picksAmazon US: current deals, useful picks and tech finds.See Picks×
Skip to content
World desk7 min

How to Detect and Customize Flaky Test Detection

Detect flaky tests by preserving first-attempt and retry outcomes, then tune retries and CI gates without hiding persistent failures.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Detect flaky tests by preserving each test’s first result and retry results, then looking for tests that change outcome across runs. A test that fails once and passes on retry is evidence of flakiness—not proof of its cause. Customize retry scope and CI policy so the signal remains visible, and investigate the conditions that made the result vary.

What flaky-test detection tells you

A flaky test produces inconsistent outcomes across runs in a way that appears non-deterministic. That inconsistency makes failures harder to trust and can lead to repeated runs and investigations. The key signal is variation: a test that fails and then passes under a retry is different from one that fails on every attempt.

Detection identifies instability; it does not diagnose the root cause. Keep the first attempt and every retry as separate results. If reporting shows only the final green status, a retry can hide the original failure and remove the evidence you need to fix the test.

A workflow for finding intermittent tests

  1. Preserve attempt-level results. Record whether the initial run passed or failed, each retry outcome, and the final classification. Keep logs and failure artifacts attached to the relevant attempt.
  2. Repeat selectively. Use the runner’s retry behavior to identify fail-then-pass cases. For an investigation, deliberately repeat tests where supported and compare results. Repetition can reveal inconsistency, but a test that passes repeatedly is not thereby guaranteed reliable.
  3. Compare the conditions. Check test order, shared state, concurrency, environment, and whether the test behaves differently when run alone. Look for a prior test that leaves state behind, or a race between the test and the application.
  4. Retain useful diagnostics. For UI tests, save screenshots or video on failure so you can reconstruct what the page showed. Keep enough attempt context to compare the initial failure with a later pass.
  5. Classify accurately. A test that fails initially and passes on retry is inconsistent. A test that fails on all attempts remains a failure; do not relabel it as flaky just because retries occurred.
  6. Fix, then verify. Address the cause and rerun under the conditions that exposed it. If equivalent coverage already exists, consider deleting or rewriting the unstable test, or moving coverage to a more reliable lower-level test.

Customize detection without hiding failures

Make four decisions explicitly: what counts as the detection signal, which tests are in scope, what CI should do with a detected flake, and how retries affect test isolation and run time. Do not assume every framework gives retries the same meaning.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Choose a small, explicit retry budget

Retries are useful as a diagnostic and can distinguish a retry-pass from a persistent failure. They also add runtime, and an overly permissive policy can make unstable tests look harmless. Choose a limited budget based on the suite’s runtime and the impact of missed defects; there is no universal retry count. Keep retry outcomes in reports even when the job is allowed to continue.

Set the scope to match the evidence

Begin with the affected tests or group rather than silently applying special treatment to the whole suite. Expand or narrow the scope as you learn whether instability is local or shared. In Playwright Test, retries can be configured globally and for test groups. In pytest, rerun, random-order, failure-replay, and failure-classification capabilities are provided through plugins; plugin behavior and configuration depend on the plugin you use.

Separate detection from the CI gate

A report can flag a flaky result without failing the job, or the pipeline can fail when a test is classified as flaky. Decide which behavior you want rather than letting a retry’s final pass make the decision implicitly. A stricter gate makes instability visible to contributors; a non-blocking report can be a temporary containment policy while a team investigates. In either case, assign follow-up and avoid treating quarantine as a permanent fix.

Account for retry isolation and runtime

Retry behavior can affect whether a test runs immediately after failure or later in a more isolated context. Playwright’s current configuration reference describes immediate retry and isolated retry at the end of the suite; the isolated strategy can reduce interference but increases total run time. Configuration options are version-sensitive, so check the reference against the Playwright version installed in your project before using them.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Playwright Test: configure retries and flaky-test reporting

Playwright Test retries are off by default. A test that fails on its first attempt and passes on retry is reported as flaky; a test that still fails after retries remains failed. The following example enables a limited retry budget in CI and makes flaky classifications fail the run. failOnFlakyTests is documented as available since Playwright v1.52; verify your installed version before adding it.

// playwright.config.ts
import { defineConfig } from '@playwright/test';

export default defineConfig({
  retries: process.env.CI ? 2 : 0,
  failOnFlakyTests: !!process.env.CI,
});

The number above is an example policy, not a universal recommendation. If you want to observe flakes without blocking CI, leave failOnFlakyTests disabled while still collecting the retry classification. To investigate repeatability, Playwright also documents repeatEach, which repeats each test and is useful for debugging flaky tests. Avoid enabling repeated runs indiscriminately in a large suite: it increases execution time and should answer a specific diagnostic question.

Configuration and version notes are in the Playwright retries guide and the TestConfig reference. The current reference lists retryStrategy as available since v1.62; confirm version support before relying on it.

pytest: rerun and quarantine carefully

pytest’s core guidance points to plugins for rerunning failures, randomizing test order, replaying observed failures, or classifying failures. These mechanisms are not interchangeable: randomizing order helps expose state dependencies, while rerunning a failure tests whether the outcome varies. Choose a plugin and configure it according to the behavior you need, preserving attempt-level results so a pass after rerun remains visible.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

pytest also cautions that non-strict xfail can act like manual quarantine: a known failure may stop breaking the build, but leaving it in place permanently is dangerous. If you quarantine a test, keep the status visible, identify an owner, and track a path back to a reliable test rather than allowing the exception to disappear from view. For UI failures, capture screenshots or video when tests fail.

See pytest’s Flaky tests documentation for its discussion of causes and mitigation.

Azure Pipelines: report flakes separately from build policy

Azure Pipelines documents automatic flaky-test detection using reruns as well as custom detection. Its management options distinguish identifying and reporting flaky tests from deciding whether they should fail a build. Flaky data availability can depend on the branch, and teams can use the flaky tag to troubleshoot, then manually create bugs or mark and unmark tests based on analysis.

Use that separation deliberately: inspect the attempt history, decide whether the test is genuinely inconsistent, and choose a reporting or build policy that does not erase the result. See Microsoft Learn’s Manage flaky tests – Azure Pipelines.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Find and fix the cause

Race conditions and timing

Synchronize tests with meaningful application state rather than relying on arbitrary delays. A sleep may make one run pass, but application timing can change and the test can become flaky again; the delay also makes the suite slower. Log access to shared resources and wait for a specific condition that represents readiness.

Order dependence and shared state

Run the test independently and vary test order to see whether another test is setting up or corrupting state. Make tests independent, reset shared resources where necessary, and avoid relying on execution order. pytest’s documentation describes random-order testing as one way to expose state problems.

Environment and test boundaries

Uncontrolled system state and inadequate environment isolation can also cause inconsistent outcomes. Compare environments and resource contention between attempts. Where appropriate, split unit and integration suites so a slow or stateful integration path does not obscure a deterministic unit test. If the test is inherently brittle but equivalent coverage exists, rewrite it or move the assertion to a more reliable layer.

Google’s guidance discusses synchronization, test independence, and why arbitrary delays are a poor long-term remedy in Test Flakiness – One of the main challenges of automated testing (Part II).

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Common detection problems and fixes

Symptom Likely issue What to do
The CI job is green, but a test failed before passing on retry. The report or gate emphasizes only the final attempt. Retain attempt-level results and report the flaky classification; decide separately whether that classification should fail CI.
A test fails on every retry. This is a persistent failure, not a retry-pass flake. Keep it as a failure and investigate the failing assertion or environment.
Adding retries makes the suite much slower. The retry scope or budget is too broad, or deliberate repeats are running across too many tests. Limit retries to the needed scope and use repeat-each behavior for targeted investigation.
A test passes alone but fails in the suite. Order dependence, shared state, or concurrency may be involved. Vary order, inspect shared resources, and make setup and cleanup independent.
A fixed delay seems to help, but failures return. The delay does not synchronize on the application condition that matters. Wait for meaningful state or a specific condition instead of extending an arbitrary sleep.
A quarantined test disappears from attention. Containment has become permanent and unowned. Keep the flake visible, assign follow-up, and remove the quarantine when the test is repaired or replaced.

Or skip the browser setup

If intermittent UI failures are difficult to inspect because captures contain overlays or require custom browser infrastructure, ScreenshotNeo offers a one-request screenshot API. Cookie banners, newsletter popups, and chat widgets are removed before the shot; bot checks, blank pages, timeouts, failed loads, and cache hits are not billed. Its MCP server lets AI agents take screenshots, and the free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

See the ScreenshotNeo API documentation for request options. ScreenshotNeo also supports PNG, JPEG, WebP, or PDF captures. Sign up free for 1,000 screenshots a month with no card.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

More from the Wire

  1. World desk4 min
    How to Spot an AI Voice Scam Before Sending MoneyDon’t rely on how a caller sounds. Pause, call back through a known number, and verify the emergency with another trusted person before sending money.
  2. Mountain View desk4 min
    Google’s SynthID Detector: How to Check AI-Generated Images, Video and AudioGoogle’s SynthID Detector looks for an embedded watermark in supported images, video and audio. Here is what its results do—and do not—show.
  3. Shenzhen desk3 min
    HONOR Expands Beyond Smartphones With Humanoid Robot RevealHONOR said it unveiled its first humanoid robot at MWC 2026 and named shopping assistance, workplace inspections, and supportive companionship as intended uses. Later Robotics D1 claims and a reported…
Recommended PC Tool
Recommended PC Tool
PC Slower Than It Used to Be?Free scan - under a minute
Crashes, No Sound, or Screen Glitches?Free driver scan

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.