Recommended Free Tools
A flaky test passes and fails without a meaningful change to the code under test or its inputs. Treat that inconsistency as a symptom to investigate, not a reason to dismiss the test: identify the uncontrolled dependency, repair or contain it, and verify the test under the conditions that exposed the failure.
What makes a test flaky?
A test is nondeterministic when it produces different results without a noticeable change in the code, test, or relevant inputs. The difference usually points to a dependency the test does not control: shared state, timing, the clock, a remote service, browser behavior, or managed resources. Martin Fowler’s guide to eradicating nondeterminism and Mike Bland’s discussion of testing culture describe this underlying pattern.
As an Amazon Associate I earn from qualifying purchases.
A rerun that passes is useful evidence that the failure may be intermittent. It does not establish that the product code is correct, nor does it repair the test. First confirm that you are comparing runs of the same revision under sufficiently similar conditions.
How to investigate a flaky failure
1. Record the failure and its context
Capture the test name, assertion or error, commit, environment, test order, and whether the same revision passes on rerun. Note relevant differences between runs, such as parallel execution or external service availability. Without this context, a failure after a code or environment change can be mistaken for flakiness.
2. Compare an isolated run with a suite run
Run the failing test by itself, then in the suite where it originally failed. If it fails only in the suite, investigate order dependence: an earlier test may leave behind database records, altered global or static state, a modified singleton, or an incomplete cleanup. Check for collisions when tests run in parallel.
Where practical, rebuild a known starting state for each run. A transaction rollback can help contain database changes if the test does not need to commit. Cleanup and shared fixtures can reduce setup cost, but they also create opportunities for one test’s leftover state to affect another.
3. Make the failure observable
Repeat runs under controlled conditions and record useful logs and state. Change one suspected variable at a time—such as test order, parallelism, starting data, or an external dependency—so the result helps distinguish causes instead of adding more uncertainty. A seed can help when the test framework or application uses randomized behavior, but it is not a universal fix: record the seed and the conditions it controls.
4. Inspect timing and asynchronous boundaries
Look for fixed sleeps used to wait for background work, UI changes, or network responses. A short sleep can expire before a slow response arrives; a long one wastes time and still cannot guarantee the response has arrived. Prefer a callback when the system supports one, or poll for the expected condition with a bounded timeout. The timeout should fail with a useful message if the condition never occurs. Fowler recommends callbacks or polling rather than bare sleeps in his treatment of nondeterministic tests.
5. Check environmental dependencies
Look for dependencies that can vary between runs, including:
- Direct reads of wall-clock time, date boundaries, or time zones.
- Remote services, network conditions, and data that changes outside the test.
- Browser timing, animations, popup dialogs, or other UI behavior.
- Pre-existing test data, shared database records, and global state.
- Resources such as database connections that are not reliably released.
Control or narrow the dependency where possible. For example, provide stable test data or a controllable clock instead of relying on the live environment. If a dependency cannot be controlled, make the test’s dependence explicit and check that the failure is not caused by setup, teardown, or resource leaks.
6. Fix the cause and validate the original failure conditions
Choose a repair that stabilizes the test without dropping the regression it is meant to catch. Rerun the test alone and in the relevant suite, including the order, concurrency, timing, or dependency conditions that previously triggered the failure. Preserve an assertion for the original defect when possible. A fix that merely makes the test pass once has not established that its signal is dependable.
Choose a repair that keeps useful coverage
Compare possible fixes by how confidently they address the diagnosed cause, how stable they are under the known failure conditions, what regression coverage they retain, and their effects on runtime, maintenance, and fidelity to production behavior.
Isolation and test data
Rebuilding fixture state is usually easier to reason about when its setup cost is acceptable. If setup is expensive, cleanup or shared immutable fixtures may be appropriate, but cleanup must itself be reliable. Roll back database work when the test does not need to commit; where it does need to verify committed behavior, rollback is not a substitute for exercising that behavior.
Rank #4
Asynchronous work
Use a callback if the system offers a reliable completion signal. Otherwise, use bounded polling for the actual condition the test needs—not an arbitrary delay. The timeout exposes a genuinely missing response rather than letting the test wait indefinitely.
Third-party services and browser boundaries
Stubbing an unstable service or GUI boundary can make a test repeatable, but it also removes some end-to-end confidence. Keep another verification method for behavior excluded by the stub. For browser suites, focus end-to-end coverage on important user journeys and move detailed rules into faster, lower-level tests. The practical test pyramid discusses balancing test levels; Fowler’s microservice testing guidance covers trade-offs around test boundaries.
When to quarantine a flaky test
Quarantine can keep one unreliable test from obscuring the rest of a suite’s signal while a repair is underway. But a quarantined test is no longer an ordinary regression check in that suite. Keep the test visible in a separate queue or later pipeline stage, and record the failure reason, a named owner, and a removal deadline. Fowler gives a one-week limit as an example, not a universal standard; choose a deadline that fits the team’s workflow and do not let quarantine become permanent.
Best Value
Browser screenshots without browser setup
For browser failures, screenshots can make the visible state easier to inspect alongside logs and other captured state. ScreenshotNeo is a website screenshot API and MCP server for developers. If you capture pages yourself, retain the test’s relevant environment and timing details so the screenshot helps diagnose the failure rather than masking it.
Or skip the browser setup
A single GET request can capture a URL as an image or PDF. Example cURL request:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
See the ScreenshotNeo API documentation for request options. Cookie banners are accepted and removed before capture, along with known consent banners, newsletter popups, and chat widgets; each cleanup step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers report the page verdict and billing status. Its MCP server provides screenshot and page-information tools for AI agents. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 screenshots. Sign up for 1,000 free screenshots a month with no card.
Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minuteWindows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallFrequently Asked Questions
How can I tell a flaky test from a real regression?
Compare runs of the same revision under similar conditions and reproduce the failure while varying one suspected dependency at a time. A pass on rerun alone does not rule out a real defect.
Should I quarantine a flaky test?
Only as a temporary containment measure while work proceeds. Keep it visible, assign an owner, and set a removal deadline so it does not silently stop providing regression coverage.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




