DriversRecommendedOutdated drivers can make a good PC feel brokenScan driver issues before chasing fixes manually.Scan NowOctober DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsClean PCRecommendedOne scan can reveal what keeps slowing WindowsLook for cleanup and repair opportunities.Run Scan×
Skip to content
World desk5 min

How to Find and Fix Flaky Tests

A passing rerun is a clue, not a fix. Trace flaky test failures to uncontrolled state, timing, or environment dependencies, then validate the repair under the conditions that exposed them.
Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A flaky test passes and fails without a meaningful change to the code under test or its inputs. Treat that inconsistency as a symptom to investigate, not a reason to dismiss the test: identify the uncontrolled dependency, repair or contain it, and verify the test under the conditions that exposed the failure.

What makes a test flaky?

A test is nondeterministic when it produces different results without a noticeable change in the code, test, or relevant inputs. The difference usually points to a dependency the test does not control: shared state, timing, the clock, a remote service, browser behavior, or managed resources. Martin Fowler’s guide to eradicating nondeterminism and Mike Bland’s discussion of testing culture describe this underlying pattern.

As an Amazon Associate I earn from qualifying purchases.

A rerun that passes is useful evidence that the failure may be intermittent. It does not establish that the product code is correct, nor does it repair the test. First confirm that you are comparing runs of the same revision under sufficiently similar conditions.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How to investigate a flaky failure

1. Record the failure and its context

Capture the test name, assertion or error, commit, environment, test order, and whether the same revision passes on rerun. Note relevant differences between runs, such as parallel execution or external service availability. Without this context, a failure after a code or environment change can be mistaken for flakiness.

2. Compare an isolated run with a suite run

Run the failing test by itself, then in the suite where it originally failed. If it fails only in the suite, investigate order dependence: an earlier test may leave behind database records, altered global or static state, a modified singleton, or an incomplete cleanup. Check for collisions when tests run in parallel.

Where practical, rebuild a known starting state for each run. A transaction rollback can help contain database changes if the test does not need to commit. Cleanup and shared fixtures can reduce setup cost, but they also create opportunities for one test’s leftover state to affect another.

3. Make the failure observable

Repeat runs under controlled conditions and record useful logs and state. Change one suspected variable at a time—such as test order, parallelism, starting data, or an external dependency—so the result helps distinguish causes instead of adding more uncertainty. A seed can help when the test framework or application uses randomized behavior, but it is not a universal fix: record the seed and the conditions it controls.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

4. Inspect timing and asynchronous boundaries

Look for fixed sleeps used to wait for background work, UI changes, or network responses. A short sleep can expire before a slow response arrives; a long one wastes time and still cannot guarantee the response has arrived. Prefer a callback when the system supports one, or poll for the expected condition with a bounded timeout. The timeout should fail with a useful message if the condition never occurs. Fowler recommends callbacks or polling rather than bare sleeps in his treatment of nondeterministic tests.

5. Check environmental dependencies

Look for dependencies that can vary between runs, including:

  • Direct reads of wall-clock time, date boundaries, or time zones.
  • Remote services, network conditions, and data that changes outside the test.
  • Browser timing, animations, popup dialogs, or other UI behavior.
  • Pre-existing test data, shared database records, and global state.
  • Resources such as database connections that are not reliably released.

Control or narrow the dependency where possible. For example, provide stable test data or a controllable clock instead of relying on the live environment. If a dependency cannot be controlled, make the test’s dependence explicit and check that the failure is not caused by setup, teardown, or resource leaks.

6. Fix the cause and validate the original failure conditions

Choose a repair that stabilizes the test without dropping the regression it is meant to catch. Rerun the test alone and in the relevant suite, including the order, concurrency, timing, or dependency conditions that previously triggered the failure. Preserve an assertion for the original defect when possible. A fix that merely makes the test pass once has not established that its signal is dependable.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Choose a repair that keeps useful coverage

Compare possible fixes by how confidently they address the diagnosed cause, how stable they are under the known failure conditions, what regression coverage they retain, and their effects on runtime, maintenance, and fidelity to production behavior.

Isolation and test data

Rebuilding fixture state is usually easier to reason about when its setup cost is acceptable. If setup is expensive, cleanup or shared immutable fixtures may be appropriate, but cleanup must itself be reliable. Roll back database work when the test does not need to commit; where it does need to verify committed behavior, rollback is not a substitute for exercising that behavior.

Asynchronous work

Use a callback if the system offers a reliable completion signal. Otherwise, use bounded polling for the actual condition the test needs—not an arbitrary delay. The timeout exposes a genuinely missing response rather than letting the test wait indefinitely.

Third-party services and browser boundaries

Stubbing an unstable service or GUI boundary can make a test repeatable, but it also removes some end-to-end confidence. Keep another verification method for behavior excluded by the stub. For browser suites, focus end-to-end coverage on important user journeys and move detailed rules into faster, lower-level tests. The practical test pyramid discusses balancing test levels; Fowler’s microservice testing guidance covers trade-offs around test boundaries.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

When to quarantine a flaky test

Quarantine can keep one unreliable test from obscuring the rest of a suite’s signal while a repair is underway. But a quarantined test is no longer an ordinary regression check in that suite. Keep the test visible in a separate queue or later pipeline stage, and record the failure reason, a named owner, and a removal deadline. Fowler gives a one-week limit as an example, not a universal standard; choose a deadline that fits the team’s workflow and do not let quarantine become permanent.

Browser screenshots without browser setup

For browser failures, screenshots can make the visible state easier to inspect alongside logs and other captured state. ScreenshotNeo is a website screenshot API and MCP server for developers. If you capture pages yourself, retain the test’s relevant environment and timing details so the screenshot helps diagnose the failure rather than masking it.

Or skip the browser setup

A single GET request can capture a URL as an image or PDF. Example cURL request:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

See the ScreenshotNeo API documentation for request options. Cookie banners are accepted and removed before capture, along with known consent banners, newsletter popups, and chat widgets; each cleanup step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers report the page verdict and billing status. Its MCP server provides screenshot and page-information tools for AI agents. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 screenshots. Sign up for 1,000 free screenshots a month with no card.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Frequently Asked Questions

How can I tell a flaky test from a real regression?

Compare runs of the same revision under similar conditions and reproduce the failure while varying one suspected dependency at a time. A pass on rerun alone does not rule out a real defect.

Should I quarantine a flaky test?

Only as a temporary containment measure while work proceeds. Keep it visible, assign an owner, and set a removal deadline so it does not silently stop providing regression coverage.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

More from the Wire

  1. World desk4 min
    How to Spot an AI Voice Scam Before Sending MoneyDon’t rely on how a caller sounds. Pause, call back through a known number, and verify the emergency with another trusted person before sending money.
  2. Mountain View desk4 min
    Google’s SynthID Detector: How to Check AI-Generated Images, Video and AudioGoogle’s SynthID Detector looks for an embedded watermark in supported images, video and audio. Here is what its results do—and do not—show.
  3. Redmond desk20 min
    How to create a link to File or Folder in Windows 11Windows 11 gives you several ways to point to a file or folder without moving or duplicating it. You can create a desktop shortcut,…
Recommended PC Tool
Recommended PC Tool
PC Slower Than It Used to Be?Free scan - under a minute
Outdated Drivers Are Slowing You DownFree scan - exact matches

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.