No—not on every change. A team can run a trusted, carefully selected set of affected tests for quick feedback, then run broader tests after integration or before release. Expand the run when a change touches shared or core code, its impact is unclear, or the selection system cannot reliably identify the tests at risk. Selective testing is a way to schedule coverage, not a reason to leave important behavior untested.
How selective test runs work
A test-selection system tries to connect changed code with the tests that exercise it. With dependency information, it can identify tests affected indirectly through other components and run those rather than the entire suite. Google described this approach in its 2011 account of its own system: Testing at the Speed and Scale of Google. That example shows a possible design, not proof that every project can achieve the same accuracy.
Another approach is test-impact analysis, which uses information about code changes and test execution to select tests. Microsoft’s Azure Pipelines documentation describes the feature’s behavior and limitations, including cases where it cannot analyze changed files and falls back to running all tests. Teams should inspect the selection reports and confirm current product support and configuration before relying on it: Use Test Impact Analysis.
What selective execution can and cannot tell you
A selected run offers evidence about the tests the system believes are relevant. It does not establish that unselected tests are unnecessary, that the selection was correct, or that the change is ready to release. The value depends on how well the project tracks dependencies and represents generated files, configuration, assets, and cross-component contracts.
Use different testing scopes at different stages
A practical pipeline can give developers fast feedback without treating a narrow presubmit result as the whole qualification process. Google’s documented example runs affected tests before submission and all project tests in continuous build after commits. Google Cloud also describes presubmit checks that can include unit, fuzz, hermetic integration, and static and dynamic analysis; its documentation says core or widely used code may receive a global presubmit. These are examples of Google practice, not universal requirements: Efficacy Presubmit and Google Cloud’s approach to change.
| Pipeline stage | Useful scope | What it contributes |
|---|---|---|
| Local development and presubmit | Fast checks and tests selected as affected by the change | Frequent feedback while a change is being written and reviewed |
| After integration | Broader project testing, including tests not selected for presubmit | A wider check for missed dependencies and interactions |
| Release qualification | Testing appropriate to the release’s impact and risk | Evidence that the candidate has been evaluated beyond a narrow change-specific run |
The exact staging depends on how the software is built and deployed. Google’s guidance on test strategy emphasizes that the right qualification process depends on the software’s purpose and audience; it does not prescribe one test volume for every project: How Much Testing is Enough?.
Keep the layers of testing distinct
- Unit tests provide focused feedback on local behavior.
- Integration tests check interactions between components.
- End-to-end tests exercise critical user journeys.
- Coverage measures can help identify code or features that lack attention, but coverage alone does not demonstrate that behavior is correct.
A selective unit-test pass should not be mistaken for evidence from integration or end-to-end checks. Decide which layers belong at each pipeline stage based on the software’s risks and qualification needs.
When to broaden the test run
Use a broader scope when the cost of a missed regression is high or when the selection system has weak information about the change. The trigger may be a project rule, a reviewer’s judgment, or an automatic fallback.
Recommended Free Tools
- Core or widely used code: A change to a shared library or central component can affect many consumers. Google Cloud describes using a global presubmit for some core or widely used code.
- Shared dependencies or public interfaces: Broaden coverage when downstream components, API contracts, or cross-component behavior may be affected.
- Build, test, or common configuration changes: These can alter what gets compiled or tested, or how tests run, making ordinary impact mapping less dependable.
- Unknown or incomplete impact: If the selector cannot reason about changed files or dependencies, use the broader fallback rather than treating missing information as proof of no risk. Microsoft documents an all-tests fallback in certain analysis gaps.
- High-impact changes or releases: Choose the scope according to likely consequences and the confidence required before deployment, rather than applying a universal test-count threshold.
Apache Airflow’s selective CI documentation is an example of explicit project policy: it identifies changes that trigger full-test checks and narrower edits that can receive selective checks. Those rules are specific to Airflow, but the approach—writing down which changes expand the run—can make a team’s own policy more consistent: Selective CI Checks.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Make the selection system part of the testing strategy
Track misses and fallback behavior
Selection can produce false negatives: a relevant test may be omitted. Google’s account of presubmit efficacy explicitly notes the risk of an incorrect prediction. Review selection reports, investigate failures discovered only in broader runs, and make sure the system’s fallback behavior is visible to developers rather than silently narrowing coverage: Efficacy Presubmit.
Rank #4
Do not treat flaky tests as irrelevant tests
A flaky test can make either a selective or a full run harder to interpret. Unreliable results are a signal-quality problem to address; simply excluding a test can conceal a real regression. Google’s 2016 account describes separating presubmit gating from post-submit release evaluation and reports that about 1.5% of its test runs reported a flaky result in that historical account. That figure is specific to Google’s 2016 experience, not a current or industry-wide rate: Flaky Tests at Google and How We Mitigate Them.
Reduce runtime without confusing speed with scope
If the full suite is slow, infrastructure can reduce elapsed time as well as selective execution can reduce the number of tests run. Bazel documents features such as sharding and remote execution. Those affect how tests are scheduled or executed; they do not decide whether a particular behavior needs coverage: The Bazel Code Base.
Do these 3 things before closing this tab:
1Clear out junk files and repair common Windows errors2Scan for outdated or missing drivers - takes under a minute3Repair Windows errors before they cause bigger problemsBest Value
Choose a policy you can validate
There is no universal cutoff for when selective testing is safe. Base the policy on the project’s impact model, the quality of its dependency and test information, and the consequences of a missed regression. Measure whether broader runs find failures missed by selection, how long feedback takes, and whether the selection reports explain what was included. Keep a broader testing stage in the pipeline, and increase presubmit scope whenever the selector’s confidence or the change’s risk warrants it.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




