A flaky visual test captures a different page state on different runs even though no intended UI change explains the difference. Diagnose the capture before updating the baseline: inspect the changed pixels, trace, network requests, console, DOM, and capture metadata. Then stabilize data, fonts, assets, and motion; mask only genuinely volatile regions. Retries can reveal intermittency, but a passing retry is not a fix.
What makes a visual test flaky?
A screenshot mismatch can signal a real regression, but it can also result from the same interface being captured in different conditions. A changing animation frame, late-loading font, unfinished request, missing image, or variable data can alter pixels without a corresponding code change. Playwright describes a test that fails on its first run but passes on retry as “flaky.” That label identifies inconsistent results; it does not diagnose the cause. Playwright’s retry documentation
Chromatic’s guidance identifies animations, late-loading resources, changing data, layout behavior, late font loading, and unfinished network requests as common sources of instability. Its trace can help you inspect network requests, console logs, DOM snapshots, and snapshot metadata. Chromatic: Unstable tests
How to diagnose a flaky visual test
- Reproduce it without changing the baseline. Use the same browser, viewport, fixtures, and environment where possible. If CI is the only place it fails, investigate differences between CI and local runs rather than assuming the baseline is wrong.
- Inspect the expected and actual screenshots. Locate the changed region. A shifted text block may suggest a font or layout change; a missing image points toward a resource issue; inconsistent values may indicate nondeterministic data; a partly changed component may be animation.
- Open the trace or equivalent capture diagnostics. Look at requests, console errors, DOM state, and capture metadata. Check which assets were still loading or had failed when the screenshot was taken. Chromatic recommends starting with its trace for unstable tests. Chromatic: Unstable tests
- Classify the difference before changing code or the baseline. Decide whether it reflects a real UI regression, uncontrolled application input, resource timing or failure, motion, or intentionally variable content.
- Fix the source of variation, then rerun. Keep the test conditions and baseline unchanged until you can explain the mismatch. Update the baseline only when the changed appearance is intended.
Stabilize the common sources of variation
Make test data deterministic
Use fixed values or a seeded generator instead of unseeded random values. Control dates, counters, identifiers, and other values that can change between runs if they affect the pixels. Keep the application state and fixtures consistent so a screenshot compares the same state each time.
Quick wins for a faster PC:
Clear out junk files and repair common Windows errorsFree Scan →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Repair Windows errors before they cause bigger problemsFix Now →#1 Best Overall
Make fonts and other assets available reliably
A screenshot taken before a web font loads can differ in text width, line breaks, and overall layout. Ensure the intended fonts load reliably and preload them when appropriate. Prefer stable local or static image and font resources to unpredictable remote hosts where you can. Keep image optimization and compression behavior consistent. Chromatic’s resource-loading guidance covers retries for assets that do not load in time as well as missing images, fonts, stylesheets, and domain issues. Chromatic: Resource loading
Handle animation according to what the test asserts
For a static visual comparison, disable or pause animation deliberately so the capture does not depend on the frame timing. If animation is what you are testing, retain it and assert its expected behavior rather than hiding it. Chromatic documents pausing CSS transitions, CSS and SVG animations, and videos; its default CSS behavior pauses at the end of the animation cycle, and its configuration can change where animation is paused. Those behaviors describe Chromatic’s capture environment and should not be assumed for another tool. Chromatic: Animations
Rank #2
Wait for the actual condition, not an arbitrary delay
If layout changes after a request or a component renders late, identify what the test needs to wait for and wait for that condition. A fixed sleep can hide timing problems on one machine while remaining too short on another. Use the trace and DOM state to determine whether the relevant element, font, or resource is ready rather than adding delay without evidence.
Mask only intentionally volatile regions
Hide, mask, or normalize a small region only when its variation is outside the behavior the test is meant to verify—for example, a deliberately changing timestamp. Broad masks can conceal real regressions. Playwright’s screenshot assertion options include ways to hide or modify dynamic regions; check the API reference for syntax supported by the version installed in your project. Playwright: Visual comparisons and PageAssertions API
PC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minuteUse retries as a signal, not a repair
Playwright retries are off by default unless configured. When enabled, they rerun a failing test; a test that fails initially and passes on retry is reported as flaky. That result helps surface intermittency, but the successful retry does not tell you why the original screenshot differed. Fix the underlying variation and verify it across controlled reruns instead of treating a retry-only green result as proof of stability. Playwright: Retries
Choose a visual testing approach by its diagnostics and controls
Playwright, Chromatic, and Percy document different approaches to screenshot comparison and stabilization. The vendor materials cited here do not establish a complete, current comparison of features, compatibility, or pricing, so verify details with each provider before choosing.
Rank #4
| Approach | What the cited documentation establishes | What to check for your team |
|---|---|---|
| Playwright screenshot assertions | Playwright documents native screenshot comparisons and options to hide or modify dynamic regions. Visual comparisons | Confirm the API syntax and behavior for your installed version; inspect the diagnostics available in your runner and CI setup. |
| Chromatic | Chromatic documents trace-based diagnosis, resource-loading guidance, and capture behavior for animation. Unstable tests, Resource loading, Animations | Check how its capture environment, framework integration, and hosted review workflow fit your tests. |
| Percy | Percy’s article describes integration with Jest, Cypress, Playwright, and Selenium, plus stabilization that freezes animations, disables blinking cursors, and normalizes dynamic rendering. Percy: Integration and stabilization | Verify current integrations, stabilization controls, diagnostics, and pricing with Percy; the cited article is not a full current product comparison. |
Or skip the browser setup
If you need a screenshot for a one-off check or a capture workflow outside your existing visual test runner, ScreenshotNeo is a website screenshot API and MCP server. A GET request returns a PNG, JPEG, WebP, or PDF. This does not replace diagnosing flaky assertions in your test suite; it is another way to capture a page.
For example, save a WebP screenshot of a page with cURL:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
See the ScreenshotNeo API documentation for parameters and setup. Before a capture, ScreenshotNeo can accept the cookie or consent banner and remove more than 60 known consent platforms, newsletter popups, and chat widgets; each of those steps can be turned off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and the response reports the page verdict and billing status in headers. Its MCP server provides take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients.
The Free plan includes 1,000 screenshots per month with no card required; paid plans start at $5 for 3,000 shots. Sign up for ScreenshotNeo’s free plan.
Troubleshooting by symptom
| Symptom | Likely source to investigate | Next step |
|---|---|---|
| Text wraps differently or the page shifts vertically | A web font loaded late or did not load; layout changed after capture began. | Check font and stylesheet requests and DOM state in the trace; ensure fonts are ready and wait for the relevant layout condition. |
| An image, icon, or background is missing | A resource failed, arrived late, or came from an inconsistent host. | Inspect network failures and domains; make the asset source stable and verify loading behavior. Chromatic documents resource-loading retries and missing-asset cases. Resource loading |
| A component differs only in a small animated area | The capture occurred at a different animation frame. | Pause or disable motion for a static comparison, or test the animation deliberately. Verify the behavior of your specific capture tool. |
| Numbers, labels, or content change between runs | Random, time-dependent, or otherwise variable application data. | Use fixed fixtures or seeded values and control the rendered state. |
| The test passes only after retry | An intermittent condition remains, such as timing, data, or resource availability. | Use the failing attempt’s trace and screenshot to isolate it; do not treat the retry as the repair. |
| A visual mismatch appears only in CI | The capture conditions may differ from local conditions, or a request may behave inconsistently in CI. | Reproduce with the same browser, viewport, fixtures, and CI environment where possible, then inspect trace, network, console, and DOM evidence before changing the baseline. |
Frequently asked questions
Should I update the baseline when a visual test fails?
Only after you have established that the changed appearance is intentional. A mismatch alone does not distinguish a UI regression from an unstable capture.
Can masking fix a flaky test?
Masking can remove deliberately variable pixels from a comparison, but it does not fix unstable rendering. Keep masks as narrow as possible so they do not hide changes the test should catch.
Free tools Windows power users keep installed
One-click scans. No signup required.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




