Driver FixRecommendedSound, Wi-Fi or graphics acting up? Check drivers firstFind missing or outdated drivers fast.Check DriversOctober DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsPC HealthRecommendedCrashes, freezes, slowdowns? Check your PC nowSpot repairable issues before they interrupt work.Check PC×
Skip to content
World desk6 min

False Positives vs. False Negatives in Software Testing

A red test is not always a code defect, and a green run does not prove software is defect-free. Learn to distinguish false alarms from missed defects and investigate each.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A false positive says a defect exists when the tested software behaves as it should; a false negative lets a real defect go undetected. In a test runner, red means an assertion failed under the conditions of that run—not necessarily that production code is wrong. Green means the executed assertions passed—not that the software is defect-free. To classify either result, compare the behavior with the intended specification.

What the terms mean in software testing

The ISTQB glossary defines a false-positive result as one that reports a defect although none exists in the test object, and a false-negative result as one that fails to identify a defect that is actually present. The terms depend on what counts as “positive”: here, a positive finding means the test reports a defect.

Test outcome Actual condition Interpretation
Reports a defect No defect under the intended behavior False positive: a false alarm
Does not report a defect A defect is present False negative: a missed defect
Reports a defect A defect is present Correct detection
Does not report a defect No defect is present Correct non-detection

A red test may reflect a code regression, but it may also result from a faulty test, incorrect fixture, unstable environment, or expectation that does not match the specification. Conversely, passing assertions establish only that the tested conditions passed; untested cases and defects outside the assertions remain possible. A 1995 FDA-hosted software terminology glossary provides historical testing vocabulary, not current regulatory guidance: FDA software terminology glossary.

How flaky tests create false alarms

A flaky test passes and fails intermittently without a clear deterministic cause. When unchanged code fails once and passes on another run, the failure may be a false alarm rather than evidence of a newly introduced defect. pytest warns that unreliable failure signals can weaken trust in test results and consume time through reruns and investigation: pytest: Flaky tests.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Common causes

  • Uncontrolled state: tests share global data, depend on leftover files or services, or fail to clean up between cases.
  • Order dependence: one test changes state that another test assumes has a default value.
  • Parallel execution: concurrent tests contend for shared resources or interfere with one another.
  • Timing assumptions: an assertion expects an operation to complete within a brittle fixed interval.
  • Floating-point comparisons: exact equality is used where a suitable tolerance is needed.

Diagnose before suppressing

  1. Preserve the initial failure, including logs, inputs, environment details, and test ordering.
  2. Re-run or replay the test to determine whether the result is intermittent. A later pass does not explain the first failure.
  3. Check isolation, cleanup, shared state, external dependencies, parallelism, timing assumptions, and numeric tolerances.
  4. Try randomized test order where appropriate to expose hidden dependencies, then reproduce the issue under controlled conditions.
  5. Fix the underlying condition. If temporary quarantine is needed to unblock work, assign an owner and follow-up; pytest describes permanent non-strict expected-failure quarantine as dangerous.

Reruns can reduce disruption but should remain an investigative aid, not a substitute for root-cause work. Terminology also varies by organization: Chromium’s CQ documentation uses “false negative” locally for a flaky failure that should have passed, rather than the ISTQB definition used here. Its retry policy discusses a trade-off between letting flaky tests through more often and reducing disruption to unrelated changes: Chromium CQ documentation.

Why tests miss real defects

A test can pass defective code when it does not exercise the faulty behavior, assert the relevant outcome, or distinguish correct from incorrect results. Weak assertions are especially easy to overlook: a test may check that a response exists without checking its contents, or verify a function call without validating its effect.

  • Map tests to specified behaviors and important boundary conditions, not just lines of code.
  • Check that each assertion would fail if the behavior it claims to protect were wrong.
  • Include negative cases, error handling, boundary values, and interactions that matter to users or downstream systems.
  • Review test data and fixtures: they can accidentally omit the conditions that reveal a defect.

Use mutation testing to probe test sensitivity

Mutation testing makes small, deliberate changes to code and checks whether the test suite detects them. In Microsoft’s .NET guidance for Stryker.NET, a mutant is “killed” when tests catch the change and “survived” when they do not. A survivor is a useful prompt to inspect whether a behavior is untested or an assertion is too weak; it is not automatically proof of a production defect. Some mutants are equivalent with respect to observable behavior, and mutation operators sample only some possible faults. See Microsoft Learn: mutation testing with Stryker.NET.

Google’s Testing Blog makes the related point that tests added to kill mutants must themselves be valuable: Google Testing Blog: Mutation Testing. Treat mutation score as a diagnostic for test-suite sensitivity, not as the probability that defects will be detected or as a goal to reach 100%. Prioritize high-risk and business-critical behavior rather than chasing a score in isolation.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How to investigate a suspicious failure or possible escape

  1. Establish the reference behavior. Identify the relevant specification or product requirement and decide what behavior is correct for the exact input and conditions.
  2. Check what changed. Compare code, environment, inputs, test order, dependencies, and configuration. Save the original failure evidence.
  3. Assess repeatability. Re-run or replay to look for intermittency, while retaining the first result as part of the evidence.
  4. Inspect the test and implementation together. If the failure is deterministic, determine whether the code violates the requirement, the test encodes the wrong expectation, or its setup is invalid.
  5. For a suspected missed defect, find the coverage gap. Identify the unasserted behavior or boundary condition, add a targeted test, and consider mutation testing to check whether meaningful code changes are detected.
  6. Make quarantine temporary and visible. Record an owner and follow-up rather than allowing a known unreliable test to disappear from attention.

Which error matters more?

Neither type is universally worse. A false negative can let a harmful defect reach users; a false positive can block a safe change, consume investigation time, and make teams less responsive to future alerts. The right balance depends on the decision the test informs.

  • Impact: What is the consequence if the defect ships, and what is the consequence if a correct change is blocked?
  • Decision point: Is the test quick local feedback, a merge gate, or a release or safety gate?
  • Likelihood and other detection: How likely is this defect class, and can monitoring, review, or downstream tests catch it?
  • Investigation cost: How much time does a noisy failure consume, and can it be reproduced quickly?
  • Recovery: Can a shipped defect be rolled back or detected downstream, or would its effects be hard to reverse?

These are practical decision questions, not a standardized scoring method. There is no general comparative cost figure that establishes one error type as always more expensive.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup

If a test or workflow needs a website screenshot as an input, ScreenshotNeo is a screenshot API and MCP server for developers. One GET request returns a PNG, JPEG, WebP, or PDF. For example, with cURL:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

See the ScreenshotNeo API documentation for parameters and setup. ScreenshotNeo accepts cookie or consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers report the page verdict and billing status. Its MCP server provides take_screenshot, get_page_info, and capture_pdf tools for AI agents. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. Sign up for ScreenshotNeo’s free plan.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Standards context

ISO/IEC/IEEE 29119-1:2022 is titled “Software and systems engineering — Software testing — Part 1: General concepts.” ISO describes Part 1 as informative; Parts 2, 3, and 4 are normative for organizations claiming conformance. Citing the overview does not mean a particular test suite is certified: ISO overview of ISO/IEC/IEEE 29119-1:2022.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

More from the Wire

  1. Shenzhen desk3 min
    HONOR Expands Beyond Smartphones With Humanoid Robot RevealHONOR said it unveiled its first humanoid robot at MWC 2026 and named shopping assistance, workplace inspections, and supportive companionship as intended uses. Later Robotics D1 claims and a reported…
  2. Cupertino desk5 min
    Apple Unveils AirPods Max 2: The Upgrade That Should Have Happened Years AgoAirPods Max 2 adds H2-powered audio features and Apple claims up to 1.5× more effective ANC, but its design, Smart Case, and 20-hour battery rating are unchanged. Wired lossless audio…
  3. Cupertino desk4 min
    Apple’s OLED Touch MacBooks Are Coming—but the Dynamic Island Is the Real GambleApple has not announced an OLED touchscreen MacBook, but reports point to high-end models arriving in late 2026 or early 2027. The reported Mac Dynamic Island could be useful, but…
Recommended PC Tool
Recommended PC Tool
PC Slower Than It Used to Be?Free scan - under a minute
Crashes, No Sound, or Screen Glitches?Free driver scan

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.