AI is used in quality engineering both to assist testing work and to test products that contain AI. Generative AI can help engineers analyze requirements, draft tests and automation, and summarize results; quality engineering for AI-enabled products examines risks such as model performance and data representativeness. In both cases, AI output is a candidate to evaluate—not proof that a product or test is correct.
How AI assists software testing
Generative AI can contribute at several points in the testing lifecycle. Its useful role is usually to produce a draft, suggestion, or analysis that a person checks against requirements, evidence, and business context.
As an Amazon Associate I earn from qualifying purchases.
Requirements and acceptance criteria
A model can restate a requirement, flag ambiguous wording, suggest questions for stakeholders, or draft scenarios that exercise acceptance criteria. This can make unclear assumptions easier to spot, but the model cannot determine the intended business rule when the requirement leaves it unstated. Product owners, analysts, and engineers must confirm the interpretation before it becomes a test oracle.
The Tool Desk
Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Test cases and test data ideas
Given a requirement or feature description, an LLM can propose normal, boundary, negative, and unusual cases, or suggest data variations worth exploring. A reviewer should check that each case is valid, non-redundant, linked to a requirement or risk, and capable of distinguishing correct behavior from incorrect behavior. A longer generated test list is not, by itself, evidence of better coverage.
Automation scripts and regression suites
AI assistance can translate a behavior description into a candidate automation script, explain existing test code, suggest edits, or help identify tests that might be prioritized or maintained. Treat generated code like any other untrusted change: review the selectors, setup, assertions, cleanup, and expected results, then run it in the relevant environment. A script can execute reliably and still assert the wrong outcome.
Test results and defect reports
AI can draft a summary from logs or execution artifacts, identify recurring failure patterns, or assemble a defect report. Verify the draft against the raw logs, screenshots, environment details, and reproduction steps before using it as release evidence. Summaries can omit a relevant condition or describe correlation as cause.
Continuous improvement
Teams may use AI to look for repeated failure patterns and propose changes to test suites or processes. To assess whether that helps, compare the assisted workflow with an agreed baseline. Useful measures include reviewed test usefulness, requirement coverage, defects found, correction time, maintenance burden, and escaped defects. These are evaluation choices for a team, not published guarantees of AI productivity.
Rank #2
How to use AI in a test workflow responsibly
- Start with a defined question. Provide the model with the relevant requirement, constraints, terminology, and desired output. Do not ask it to infer undocumented business rules.
- Keep the source and traceability. Record which requirement or risk prompted a generated test idea, and identify generated material so reviewers can distinguish suggestions from approved tests.
- Review against an oracle. Check each expected result against an authoritative requirement, design decision, or domain expert. A model-generated expected result is not independent confirmation.
- Run and inspect the artifact. Execute scripts in the intended environment, inspect outputs and failures, and validate summaries against source evidence.
- Measure the workflow, not the volume. Compare quality-relevant outcomes with the team’s baseline, including time spent correcting generated work and the ongoing cost of maintaining it.
ISTQB’s updated CT-GenAI syllabus identifies prompt engineering, evaluating generated outputs, and applying GenAI across the testing lifecycle as practical areas of focus. It is an educational resource for teams seeking a structured approach; it does not make model output authoritative.
AI for testing and testing AI are different activities
AI for testing means using AI tools to assist with test design, automation, prioritization, analysis, or reporting. Testing AI means evaluating an AI component or AI-enabled system as the product under test. A team can do either without doing the other.
The distinction matters because there are two different sources of risk: an AI assistant may produce a flawed test artifact, while an AI-enabled product may have behavior or data-related risks that ordinary checks for deterministic software do not fully address. A team using an AI tool to write tests still needs to test the resulting product, and a team testing an AI product does not need to use generative AI to do so.
Rank #3
How to test an AI-enabled system
Testing should follow the system’s risks and intended use. ISO/IEC TS 42119-2:2025 describes applying the ISO/IEC/IEEE 29119 testing series to AI systems and components through a risk-based approach. It connects risk identification to choices such as test level, test type, test-design technique, static review, and coverage measure. The appropriate combination depends on the system and the consequences of failure.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Model and behavior risks
If model performance is a material risk, include model-level testing suited to the system’s intended task and context. Define what acceptable performance means before interpreting results; a score without a defined acceptance threshold or use context is difficult to act on.
Data representativeness risks
If input data may fail to represent the conditions in which the system will be used, consider data-representativeness testing. Examine whether the data and test conditions reflect the relevant users, inputs, environments, and operating conditions. The specific checks depend on the risk being addressed.
Functional, non-functional, and use-context risks
AI components remain part of larger systems. Test functional requirements and relevant non-functional concerns at the levels where failures can occur, including interactions with surrounding software and the product’s use context. Risk may also justify static review, continuous testing where behavior can change in production, and coverage measures or design techniques suited to the system.
ISO/IEC TS 42119-2:2025 states in section 5.4: “Risk-based testing (RBT) is a core concept in the ISO/IEC/IEEE 29119 series, which expects risks to be used as the prime driver for determining the test approaches included in the test strategy and therefore the consequent software testing.” Requirements remain important alongside risk when setting a strategy; risk is not a reason to disregard what the system is required to do.
Do these 3 things before closing this tab:
1Clear out junk files and repair common Windows errors2Fix the driver behind crashes, sound loss and screen glitches3Repair Windows errors before they cause bigger problemsBuild a risk-based test strategy
- Identify risks and requirements. Consider both the intended behavior and what could go wrong, including likelihood and consequence.
- Prioritize exposure. Decide which risks warrant the strongest or most frequent evidence rather than treating every test as equally important.
- Choose test treatments. Select appropriate test levels, types, design techniques, reviews, and coverage measures for each priority risk. Include model-level or data-representativeness work when those risks apply.
- Account for change. Consider continuous testing when production behavior or the system’s operating conditions may change.
- Define evidence and acceptance criteria. Specify what result will count as acceptable and what evidence reviewers need before a release or other decision.
- Reassess after changes or failures. New data, model behavior, system integrations, or observed incidents can change the risk picture and the test strategy.
NIST’s AI Resource Center provides materials for testing, evaluation, verification, and validation (TEVV), along with resources related to the AI Risk Management Framework. NIST describes the framework as voluntary guidance and says version 1.0 is under revision. It can inform a team’s approach; it does not replace product-specific risk decisions.
Best Value
- The Certified Quality Engineer Handbook, 4th Edition
Standards and guidance: distinguish published documents from drafts
| Document | What it addresses | Status described in the consulted material |
|---|---|---|
| ISO/IEC TS 42119-2:2025, Artificial intelligence — Testing of AI — Part 2: Overview of testing AI systems | Applying the ISO/IEC/IEEE 29119 testing series to AI systems and components, with risk guiding test choices. | Published. |
| ISO/IEC TS 25058:2024, Guidance for quality evaluation of artificial intelligence systems | Guidance for evaluating AI systems using an AI system quality model; applies to organizations developing or using AI. | Published. |
| ISO/IEC 25059:2023 | A previously published edition concerning the quality model for AI systems. | Previously published edition; not identified as the second edition. |
| Second-edition ISO/IEC FDIS 25059 | Draft material describing quality-model considerations including probabilistic outcomes, learning behavior, reliance on data, product quality, and quality in use. | Identified as a draft in the approval phase in the consulted ISO material, not a published replacement for ISO/IEC 25059:2023. |
| NIST AI Risk Management Framework | Voluntary risk-management guidance and related resources, including TEVV materials. | NIST’s AI Resource Center says version 1.0 is under revision. |
Publication and draft status are not interchangeable. Teams selecting a standard should confirm the current edition and status with its publisher before treating a draft as a published requirement.
What the evidence does—and does not—show
A 2025 secondary study mapping industry-context research on AI adoption in software testing reported that many use cases were proposed, while actual implementations and observed benefits in the literature it reviewed were limited. That finding qualifies the available studied evidence; it does not prove that organizations do not use AI in testing. It also does not support a broad adoption percentage or a universal productivity claim.
For a specific team, the useful question is whether an AI-assisted workflow improves outcomes that matter under that team’s conditions, after accounting for review, correction, and maintenance. The fact that a tool can generate a test, script, or summary is not evidence that it improves quality.
Quick wins for a faster PC:
Clear out junk files and repair common Windows errorsFree Scan →Scan for outdated or missing drivers - takes under a minuteDriver Scan →Use screenshots as test evidence when visual behavior matters
When a quality workflow needs visual evidence of a page, screenshots can help reviewers inspect a rendered state or compare what appeared during a test. Screenshot capture does not validate an AI model or establish correctness on its own; it is one artifact to assess alongside requirements, execution results, and other evidence.
For a browser-based capture workflow, a developer can use a browser automation setup to open the target page, wait for the required state, and save a screenshot. If the workflow instead needs a screenshot returned from an API, ScreenshotNeo is a website screenshot API and MCP server for developers. The example below requests a WebP capture; replace the sample URL with the page used in your test. See the ScreenshotNeo documentation for request options.
Quick Recap
Or skip the browser setup
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
ScreenshotNeo accepts cookie or consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each cleanup step can be turned off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and the response reports the page verdict and billing status in headers. Its MCP server provides take_screenshot, get_page_info, and capture_pdf tools for AI agents. The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 shots. Sign up for ScreenshotNeo’s free plan.
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




