Quick wins for a faster PC:
Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Clear out junk files and repair common Windows errorsFree Scan →Automated tests become flaky or expensive to maintain when too much coverage depends on the UI, tests share mutable data, failures are masked with retries, or assertions track details that change more often than the behavior they are meant to protect. The remedy is a risk-based mix of test levels, isolated data, condition-based waits, focused assertions, and diagnostics that make failures reproducible—not a different framework alone.
1. Putting too much of the suite through the UI
End-to-end (E2E) browser tests exercise real user journeys across multiple layers. That makes them valuable for checking a small number of critical flows, but also exposes them to browser behavior, timing, test data, network dependencies, and changes in the interface. A large UI-heavy suite tends to run more slowly, fail in more places, and make it harder to locate the cause.
As an Amazon Associate I earn from qualifying purchases.
Use the least expensive test level that can reliably answer the question:
- Unit tests: check focused logic in isolation.
- Service or API tests: cover behavior across service boundaries without exercising the entire interface.
- Integration tests: check that components work together where that interaction is the risk.
- UI/E2E tests: verify a small set of essential journeys through the system as a user experiences them.
Martin Fowler’s Practical Test Pyramid describes why a portfolio with more focused, fast tests and fewer broad-stack GUI tests is often easier to maintain. Treat the pyramid as a heuristic, not a required count or ratio: test boundaries and architecture differ between products.
2. Treating a test-pyramid percentage as a target
A test mix should reflect the product’s risks, architecture, and the feedback engineers need. Google’s 2015 testing guidance offered a 70/20/10 split as a starting guess, not a universal standard; Fowler also notes that teams use test-level labels differently. Copying a percentage without considering what the tests actually verify can produce a suite that looks balanced on paper but leaves important behaviors untested.
For each test, ask what failure it would catch, whether a smaller-scope test could catch it more reliably, and whether the result provides useful feedback quickly. Add E2E coverage when the complete user journey or cross-system behavior is the thing at risk, rather than to duplicate every lower-level assertion.
3. Ignoring flaky failures—or hiding them with retries
John Micco’s 2016 post, “Flaky Tests at Google and How We Mitigate Them,” defines a flaky result as a test that exhibits both a passing and failing result with the same code. Micco reported that about 1.5% of Google test results were flaky in the context of that post. That is a historical, organization-specific figure—not a current rate for Google or an estimate for the industry.
Retries, reruns, and quarantine can help separate transient failures from blocking failures, but they have costs. Retries delay diagnosis and feedback; quarantine can hide a genuine product defect or race condition. Use them as explicit, temporary containment while investigating the root cause:
- Track repeated flaky failures and assign follow-up ownership.
- Record when a test is quarantined and make the affected coverage visible.
- Do not count a retry-passing result as evidence that the underlying test is healthy.
- Look for causes such as uncontrolled timing, shared state, external dependencies, and environment differences.
4. Using fixed sleeps or asserting before the application is ready
A hard-coded delay assumes the application will always finish a task within a chosen number of seconds. If the task finishes sooner, the test wastes time; if it takes longer, the test can fail intermittently. Synchronize on the state the scenario needs instead: for example, wait for a specific element to appear or for a request-driven result to become available. Google’s end-to-end testing guidance recommends sound waiting practices and cautions against placing every behavior check in the UI suite.
Make the wait specific to the condition being tested, and give it a bounded timeout so a real failure returns useful feedback rather than hanging indefinitely. If a test still races, capture the state around the wait and investigate what is preventing the expected condition.
5. Asserting volatile presentation instead of important behavior
Tests coupled to exact copy, layout, or internal page structure can fail after harmless design changes. Prefer assertions about the scenario’s meaningful outcome: for example, that a user can complete a purchase and receives a confirmation, rather than that a button has a particular position or a frequently revised sentence appears verbatim.
Presentation is the right target when visual fidelity itself is the requirement. In that case, use a focused visual comparison and constrain the viewport and page region being compared. Keep the visual check separate from behavioral assertions where possible so a layout change does not obscure whether the underlying flow still works.
6. Sharing mutable state or persistent test data
Tests that depend on leftover records, shared accounts, or state created by another run can pass or fail according to execution order. They may also contaminate later tests or affect external systems. Create isolated, preferably ephemeral, test data and clean it up when the run ends. Avoid relying on a particular test being run first.
Rank #4
Fakes and stubs can make tests faster and more predictable, but they are another source of drift: a fake that no longer behaves like the real dependency may let a test pass while integration fails. Keep doubles aligned with the contract that matters, and retain suitable tests against real integrations for behaviors the double cannot establish.
7. Making failures difficult to diagnose
A failed test should leave enough evidence to answer what happened, where, and under which conditions. Keep logs readable and retain relevant state—such as a screenshot for a browser failure or a database snapshot when data state is material—where the test environment and privacy requirements permit. Include the failing step and useful context in the test output rather than only reporting that an assertion failed.
Do these 3 things before closing this tab:
1Fix the driver behind crashes, sound loss and screen glitches2Clear out junk files and repair common Windows errors3Scan for outdated or missing drivers - takes under a minuteDocument known failure modes when that helps responders, but documentation is not a substitute for fixing recurring instability. A failure that routinely requires a person to rerun a suite and guess at the cause is a signal to improve isolation, synchronization, or diagnostics.
Best Value
8. Treating automation as the whole testing strategy
Automated checks are strong at repeatable verification and regression protection. They do not reliably answer every question about usability, design, or surprising edge cases. Keep exploratory testing in the quality strategy, then turn valuable discoveries into automated regression tests when the behavior can be checked repeatably. Fowler’s Practical Test Pyramid discusses exploratory testing alongside the trade-offs among test layers.
9. Choosing a test approach by fit, not test count
When deciding whether a behavior belongs in a unit, service/API, integration, or browser test, compare the approaches against the job the test must do:
| Consideration | Question to ask |
|---|---|
| Scope and fidelity | Which real behavior or dependencies does the test exercise? |
| Feedback speed | How long does it take to run locally and in CI? |
| Reliability | How exposed is it to timing, shared state, external services, browser behavior, or environment differences? |
| Maintenance | How likely are ordinary product changes to require rewriting it? |
| Debuggability | Does failure point toward a component, and is there enough evidence to reproduce it? |
| Purpose | Is it checking focused logic, an integration contract, or an essential customer journey? |
The Selenium project’s Test Practices page says, “No one approach works for all situations.” Its guidance was last modified on 2022-10-19; adapt testing practices to the system and environment rather than treating any single tool or pattern as a universal fix.
Recommended Free Tools
Or skip the browser setup
If you need screenshots to inspect a browser failure or capture a page in an automated workflow, ScreenshotNeo offers a one-request option. For example, using cURL:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
See the ScreenshotNeo documentation for request options. Before capture, it accepts cookie or consent banners like a visitor and removes 60+ known consent platforms, newsletter popups, and chat widgets; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers report the page verdict and billing status. Its MCP server provides screenshot tools for AI agents, including Claude and Cursor. The free plan includes 1,000 screenshots per month without a card; paid plans start at $5 for 3,000 shots.
Sign up for ScreenshotNeo’s free plan to get 1,000 screenshots a month with no card.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.
The Tool Desk
Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →




