October DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsPC HealthRecommendedCrashes, freezes, slowdowns? Check your PC nowSpot repairable issues before they interrupt work.Check PCOctober DealsAmazon USDeal season is back - check today's better picksAmazon US: current deals, useful picks and tech finds.See Picks×
Skip to content
World desk5 min

How Generative AI Can Improve QA Testing

Generative AI can speed up test drafting and scenario exploration, but useful QA still depends on clear requirements, correct assertions, and tests run in the real project.
Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Generative AI can help QA teams draft tests, expand scenarios, and analyze failures—but it does not determine whether a test checks the right behavior. Give it clear requirements, relevant code, and existing test conventions; then run, inspect, and maintain every generated test as part of the real project.

Where generative AI can help in QA

Used as an assistant, generative AI can turn source code, specifications, and existing tests into draft unit tests or test cases. It can also suggest boundary conditions, examine failure output, and propose scenarios a team may have missed. These are ways to support testing work, not substitutes for executing tests or applying QA judgment. The IEEE Computer practitioner playbook discusses test generation, continuous testing and feedback, failure analysis, early prototyping, and simulation of varied users or conditions as possible uses: IEEE Computer practitioner playbook.

  • Drafting: Ask for tests based on a specific function, behavior, or requirement, using the project’s test framework and conventions.
  • Scenario expansion: Ask for boundary cases, invalid inputs, and combinations of conditions to consider.
  • Failure analysis: Provide a failing test and relevant error output to get hypotheses about causes or follow-up checks.
  • Exploration: Use suggestions to identify cases worth investigating, then decide which belong in the test suite.

The tool’s output is a proposal. The team remains responsible for defining expected behavior and deciding whether a test belongs in the suite.

Why specification and code context matter

A prompt that says only “write tests for this function” leaves the expected behavior underspecified. A model may infer plausible behavior from implementation details or common patterns, then encode the wrong expectation in an assertion. Supplying requirements, preconditions, postconditions, undefined behavior, relevant code, and existing tests gives it a stronger basis for proposing meaningful cases.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Google Research’s 2026 evaluation on production bugs compared a spec-driven agent—which first documented preconditions, postconditions, and undefined behavior—with a traditional test-generation agent. The spec-driven approach improved bug detection by 9.8 percentage points (p = 0.0352) and branch coverage by 2.5 percentage points (p = 0.0034) against that study’s baseline. These are findings for that evaluation, not a guarantee that any spec-first prompt will produce the same improvement. Google Research study.

In the same work, an LLM-as-a-Judge preferred the spec-driven generated suites in 77.8% of cases over the baseline suites and in 56.7% of cases over human-authored tests. Those results describe evaluator preference in the study; they do not show that AI universally writes better tests than people.

Generated tests need review and execution

Code that looks like a test may not compile, may fail for incidental reasons, or may contain no useful assertion. Even a test that passes can check the wrong thing. The IEEE Computer playbook cautions that a plausible but incorrect assertion can make a test pass or fail for the wrong reason. Review each assertion against a requirement or specification rather than accepting it because it looks reasonable.

A 2024 study by Khalid El Haji, Carolin Brandt, and Andy Zaidman examined 290 GitHub Copilot-generated tests for 53 sampled tests from open-source Python projects. In the study’s existing-suite setting, 45.28% of generated tests passed within the suite; 54.72% were failing, broken, or empty. Without an existing test suite, 92.45% were failing, broken, or empty. These figures apply to the study’s sample and setup, not to every Copilot version, language, or current use. TU Delft study record.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A practical AI-assisted test workflow

  1. State the intended behavior. Write down the requirement the software should satisfy, including what should happen for valid and invalid inputs.
  2. Provide relevant context. Include the target code, related types or dependencies, project test conventions, and existing tests where appropriate. Avoid sending secrets or sensitive data to a tool unless its data-handling terms allow it.
  3. Ask for a contract before asking for test code. Have the model list preconditions, postconditions, boundary cases, and behavior that is undefined or unspecified. Correct misunderstandings before generation.
  4. Generate a small, reviewable set. Ask for tests in the project’s actual framework and request a short explanation of what requirement each test checks.
  5. Inspect assertions and fixtures. Verify that expected values come from requirements, not merely from the implementation. Check that test setup is realistic and that cases are independent where needed.
  6. Run tests in the project environment. Confirm that they compile, execute, and pass for the intended reasons. Where practical, introduce a known defect or use mutation testing to check that the tests detect a meaningful failure.
  7. Look for gaps, not just more tests. Review branch coverage and important input boundaries, but do not treat test count or line coverage alone as evidence of test quality.
  8. Re-run and maintain. Keep generated tests under normal review and regression processes. Revisit them when requirements or implementation change.

Testing AI features and variable behavior

When the software under test contains an AI component, identical inputs may not always produce identical outputs. A single pass/fail result may miss variation across repeated runs or inputs. The IEEE Computer playbook recommends thinking beyond a simple pass/fail label for such systems, including repeated runs, broader input coverage, and metrics suited to the behavior being evaluated. Teams should define acceptable behavioral criteria for the feature and measure against them, rather than assuming one expected string is the only valid outcome. The playbook also discusses hallucinated assertions, nondeterminism, and bias as risks to consider: IEEE Computer practitioner playbook.

Choosing an AI test-generation approach

The cited studies do not establish a universally best vendor or model. When evaluating an approach, consider whether it can use relevant specifications and code context, whether it supports explicit contract reasoning, and whether its tests are readable and correctly assert intended behavior. Also examine how you will assess defect detection—not only coverage—and the human effort needed to correct and maintain the output.

For formal learning about testing with generative AI, the German Testing Board lists an English CT-GenAI syllabus, version 1.1 (2026). The listing establishes the syllabus’s availability; it does not by itself establish a particular course or provider. German Testing Board syllabi.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup

If part of your QA workflow needs website screenshots, you can request one directly from ScreenshotNeo, a screenshot API and MCP server for developers. A one-call cURL example is:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

See the ScreenshotNeo documentation for request options. It accepts cookie banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; bot checks, blank pages, and failed loads are not billed. Its MCP server provides screenshot and PDF tools for AI agents. The free plan includes 1,000 screenshots per month with no card, and paid plans start at $5 for 3,000 shots. Sign up free for 1,000 screenshots a month, with no card.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

More from the Wire

  1. Shenzhen desk3 min
    HONOR Expands Beyond Smartphones With Humanoid Robot RevealHONOR said it unveiled its first humanoid robot at MWC 2026 and named shopping assistance, workplace inspections, and supportive companionship as intended uses. Later Robotics D1 claims and a reported…
  2. Cupertino desk5 min
    Apple Unveils AirPods Max 2: The Upgrade That Should Have Happened Years AgoAirPods Max 2 adds H2-powered audio features and Apple claims up to 1.5× more effective ANC, but its design, Smart Case, and 20-hour battery rating are unchanged. Wired lossless audio…
  3. Cupertino desk4 min
    Apple’s OLED Touch MacBooks Are Coming—but the Dynamic Island Is the Real GambleApple has not announced an OLED touchscreen MacBook, but reports point to high-end models arriving in late 2026 or early 2027. The reported Mac Dynamic Island could be useful, but…
Recommended PC Tool
Recommended PC Tool
PC Slower Than It Used to Be?Free scan - under a minute
Crashes, No Sound, or Screen Glitches?Free driver scan

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.