The Tool Desk
Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →A flaky microservice test passes on some runs and fails on others even though the relevant code has not changed. A retry can confirm that outcomes vary; it cannot show whether the cause is a test defect, an unstable environment, or a real service problem. Capture the failed run first, compare it with a passing run, then follow the evidence across the service boundaries the test exercises.
What makes a microservice test flaky?
Flakiness is inconsistent outcomes across executions under effectively unchanged relevant code. It is not the same as a test that reliably exposes a regression. In a distributed system, the test may depend on network communication, independently deployed services or dependencies, orchestration, timing, or shared test data. Those are possible sources of variation—not diagnoses. Establish the cause from the failing system’s evidence.
A green retry is useful evidence that the outcome is intermittent, but it does not prove the service is healthy or identify what varied. There is no universal number of reruns that establishes flakiness or makes a test trustworthy; compare runs in a controlled way and preserve what happened on the first failure.
How should you investigate a flaky test?
1. Preserve the first failure
Before rerunning, record enough context to compare the failure with other executions:
- The test name, shard, commit, and CI build identifier.
- Versions of the services and dependencies involved, plus relevant configuration.
- Timestamps, logs, and any trace, request, or correlation identifiers.
- Resource pressure and whether other tests or services failed around the same time.
Repeat the test in a controlled way, keeping the code revision and relevant environment as consistent as practical. Compare a failure with a pass rather than treating a successful retry as a clean bill of health. AWS Well-Architected DevOps guidance recommends investigating root causes, refining test design, and using a stable, reproducible testing environment.
2. Identify the behavior and boundary under test
State what behavior the test is meant to prove, then check whether its scope is larger than necessary. A test that proves local logic does not need to wait on a networked service; a test intended to verify a cross-service interaction needs an appropriate higher-level boundary.
| Test level | What it covers | Useful role and trade-off |
|---|---|---|
| Unit | Local logic, usually without service-to-service calls. | Typically offers the most control and quickest feedback, but cannot establish that real service interactions work. |
| Component or integration | A service or component working with selected dependencies. | Checks behavior across a meaningful boundary; its setup and repeatability depend on how those dependencies and test data are controlled. |
| Contract | Whether a service’s API expectations agree with those of its consumers or providers. | Targets compatibility at an API boundary without requiring every check to exercise a full end-to-end journey. |
| End-to-end | A user journey across multiple services. | Provides broader interaction coverage, but includes more moving parts and is generally more costly to set up, observe, and maintain. |
These levels complement one another. Keep most checks at the smallest boundary that can prove the behavior, while retaining selected higher-level tests for interactions and failure modes that local tests cannot validate. Google Cloud architecture guidance similarly recommends a large share of unit tests alongside automated higher-level integration and system tests. Infrastructure as code can help create and tear down dedicated environments and resources for those higher-level checks.
3. Follow the run across services
Use timestamps and a test-run or transaction identifier to line up test output with service telemetry. Metrics show trends such as request rate, error rate, and latency; logs record discrete events; traces show a transaction’s path across components and where errors or time accumulated. Google Cloud’s observability guidance describes these signals as complementary. A trace is especially useful when a failure could have occurred in any one of several service interactions.
Check whether the failing run coincides with service restarts, dependency errors, delayed or reordered work, shared-data collisions, resource saturation, or deployment and configuration changes. Treat each as a hypothesis to test against the timeline and telemetry, not as a standard explanation for every flaky test. Monitoring service interactions for rising errors or latency can help identify where to look next.
4. Repair the cause and make the setup repeatable
Use the evidence to correct the unstable assumption or setup. Depending on what the failing and passing runs show, practical changes may include controlling test data and cleanup, waiting for an explicit asynchronous completion condition instead of relying on timing, isolating shared state, pinning dependency versions, or provisioning a repeatable environment. These are possible remedies, not interchangeable fixes; apply the one that addresses the observed cause.
Rank #4
5. Keep unresolved failures visible
If a test cannot be repaired immediately, use an explicit team policy to quarantine it until resolved, as AWS guidance recommends. Keep it recorded as flaky, assign a path back to normal gating, and make the status of a retry-passed build distinguishable from a clean deterministic pass. Ownership, expiration, escalation, and gating rules are team decisions; there is no universal value for them.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.When is a flaky result a resilience signal?
Some intermittent failures expose a real system behavior under dependency or infrastructure disruption rather than a defective test. If the evidence points to recovery behavior, create a deliberate resilience test with a controlled scope, monitoring, safety measures, and rollback preparation. Google Cloud guidance describes testing scenarios such as regional failover, release rollback, and data restoration, and assessing recovery against recovery time objective (RTO) and recovery point objective (RPO). That is planned resilience coverage—not simply rerunning a flaky functional test.
Best Value
Why test boundaries matter in microservices
Microservices add network partitions and independently changing services to the interactions a test may need to consider. Testing Strategies in a Microservice Architecture, Toby Clemson’s foundational 2014 guidance, separates unit, integration, component, contract, and end-to-end approaches. The practical implication is not to eliminate higher-level checks, but to use each where it provides evidence that a narrower test cannot: local tests for local behavior, contract checks for API expectations, and a small set of end-to-end tests for important cross-service journeys.
Flakiness also has a suite-level cost: when a build contains many tests, even a small chance of an individual false failure can undermine confidence in the aggregate result. Google SRE’s testing chapter illustrates this with a calculation in which 42,000 results would each need correctness above 99.9999% to keep a stated aggregate false-rejection rate below 1%, under the chapter’s assumptions. This is a worked example, not a measured reliability statistic or a target that applies unchanged to every CI suite.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




