AI can help testers design, run, prioritize, and maintain tests, but it does not make the results self-validating. People still need to define expected behavior, judge risk, review AI-generated work, and decide whether the evidence supports a release. When the software being tested contains AI, the test strategy must also account for data, models, and behavior that may be probabilistic rather than exactly repeatable.
Two different meanings of AI in software testing
The phrase covers two related but distinct activities:
- Using AI to test conventional software: AI tools assist with work such as generating test cases, analyzing code, prioritizing tests, or maintaining automation.
- Testing AI-based software: testers evaluate a product whose behavior depends on data, learned models, or generative AI. The AI system itself is part of the test object.
The distinction matters. A test-generation assistant may help test a conventional application, but that does not by itself test the assistant’s own reliability. Conversely, testing an AI-based feature requires attention to its data and model behavior, whether or not AI tools are used to help with the testing.
How AI may assist with testing conventional software
A 2025 secondary study by Katja Karhu, Jussi Kasurinen, and Kari Smolander maps potential and reported AI use cases in software testing. These are application areas, not a promise that a tool will work well in a particular team or project.
| Testing activity | Possible AI assistance | What still needs human checking |
|---|---|---|
| Requirements and test design | Analyze requirements, suggest scenarios, or draft test cases and scripts. | Confirm that cases reflect the actual requirement, include important boundary conditions, and do not invent expected behavior. |
| Code and failure analysis | Analyze code, summarize an error, suggest possible causes, or help investigate a root cause. | Reproduce the issue and validate explanations against the code, logs, and observed behavior. |
| UI testing and automation | Assist with UI tests or intelligent automation, including identifying interactions to exercise. | Check that selectors, actions, and assertions represent meaningful user behavior and remain robust as the interface changes. |
| Execution and prioritization | Help prioritize tests, execute tests, or predict where defects may occur. | Ensure high-risk requirements remain covered and that a ranking has not hidden a critical test. |
| Maintenance | Suggest updates to scripts or help identify tests that may need maintenance. | Determine whether an apparent test failure reflects a real regression, a changed requirement, or a brittle test. |
These examples describe possible assistance, not autonomous authority. A generated test can be syntactically plausible while asserting the wrong thing; a persuasive explanation of a failure can still be mistaken.
What human testers contribute
Practical guidance inferred from ISTQB’s coverage of evaluating generative-AI results and managing hallucinations, reasoning errors, bias, privacy, and security risks is to keep human responsibility explicit. The exact division of work will vary with the system and its risks; the sources do not establish one universal workflow or productivity gain.
- Frame expected behavior: resolve ambiguous requirements and identify what users and stakeholders actually need.
- Set risk priorities: decide which failures matter most, including safety, security, privacy, and business impact.
- Review generated work: inspect tests, scripts, summaries, and recommendations before relying on them.
- Judge evidence: determine whether a failure is meaningful, whether coverage is adequate, and whether remaining uncertainty is acceptable.
- Own release decisions: make clear who is accountable for conclusions and what evidence supports shipping.
This is not a claim that humans never make mistakes or that every AI suggestion requires the same review. It is a way to prevent generated output from being mistaken for verified evidence.
How to test software that contains AI
AI-based systems can exhibit probabilistic behavior, non-determinism, and dependence on data. Exact repeatability may therefore be difficult, and a test plan limited to conventional input-output checks can miss important failures. ISTQB’s CT-AI v2.0 outline organizes coverage around input data testing, model testing, and machine-learning development testing; it also includes generative AI and large language models.
Free tools Windows power users keep installed
One-click scans. No signup required.
Test input data
Check whether the data supplied to the system is suitable for the intended use. Consider its quality, coverage, and relevant edge cases, as well as whether sensitive information is handled appropriately. A model can behave poorly because of its inputs even when its software integration appears to work.
Test the model’s behavior
Define acceptance criteria that fit the system’s purpose, then select functional performance measures appropriate to the feature. Assess expected and undesirable behaviors across representative cases rather than treating one successful response as proof of reliability. For generative systems, include checks for incorrect or unsupported output, bias, and security or privacy concerns.
Test the development and integration lifecycle
Include the machine-learning development process in the test scope: how data and models are built, evaluated, changed, and integrated into the product. The relevant test levels and evidence depend on the system. Keep enough traceability to understand which data, model, configuration, and software version produced a result.
Because outputs may vary, define how to evaluate results across repeated runs or a set of test cases, where that is appropriate. Do not assume that a single fixed expected string is always the right oracle; use criteria that reflect the feature’s intended behavior and risk.
Do these 3 things before closing this tab:
1Fix the driver behind crashes, sound loss and screen glitches2Clear out junk files and repair common Windows errors3Scan for outdated or missing drivers - takes under a minuteA practical human–AI testing workflow
- Describe the behavior and risk. Start with requirements, user needs, and failure consequences. Identify what must be true, what must not happen, and which cases are most important.
- Choose the right kind of AI use. Decide whether AI is assisting the test process, is part of the product under test, or both. Keep those scopes visible in plans and results.
- Use AI for bounded tasks. Ask for draft cases, candidate scripts, code analysis, or test prioritization with enough context to make the output reviewable. Treat suggestions as candidates, not approved coverage.
- Verify before execution or reliance. Check generated tests against requirements and actual application behavior. For AI-based products, verify that data, model behavior, and lifecycle concerns are represented in the test design.
- Record evidence and decisions. Preserve what was tested, relevant versions and inputs, observed results, and the reasoning behind risk or release decisions. Protect private or sensitive material when using external AI services.
- Review failures and update coverage. Reproduce important findings, distinguish product defects from test or environment issues, and revise tests when requirements or models change.
What the evidence says about adoption and results
Karhu, Kasurinen, and Smolander’s study, dated April 7, 2025, mapped industry-context studies from 2020 onward. It describes potential use cases across testing but reports that industrial implementations and observed benefits in the mapped evidence were limited. This supports a cautious conclusion: interest and proposed applications should not be presented as proof of widespread adoption, better quality, or faster releases.
Rank #4
The study also repeats Perforce survey figures: for 2024, 48% of respondents were interested in AI but had not started initiatives, and 11% were already implementing AI techniques in software testing. It cites 2025 Perforce survey results in which over 75% of respondents identified AI-driven testing as pivotal to their 2025 strategy, while 16% reported adopting AI in testing. These are survey results attributed to Perforce and quoted by the secondary study, not measurements of all software organizations or causal evidence that AI improved testing outcomes.
The reviewed evidence does not provide a broadly generalizable estimate of how much a human–AI testing workflow improves speed or quality. Treat any expected benefit as a hypothesis to evaluate in your own context, not a guaranteed uplift.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Where to learn: CT-AI or CT-GenAI?
ISTQB’s certification tracks reflect the two different problems. Both list the Certified Tester Foundation Level (CTFL) as a prerequisite; check current availability and local exam arrangements with ISTQB.
Best Value
| Track | Purpose | Listed learning routes |
|---|---|---|
| CT-AI v2.0 | Testing AI-based systems, including data, model, and machine-learning development testing. | Syllabus, sample exam, and training-provider routes. |
| CT-GenAI | Using generative AI in the software testing process, including evaluating generated results and managing associated risks. | Accredited training and self-study. |
Capture UI evidence without confusing capture with testing
For a UI test, a screenshot can provide visual evidence for a human review or a separate analysis step. A capture service does not decide whether the interface is correct; the test still needs an expected behavior and a way to judge the result. If you need repeatable website captures as part of a test workflow, ScreenshotNeo is a screenshot API and MCP server for developers.
Or skip the browser setup
Make a screenshot request with cURL; replace the example target URL with the page you need to capture. See the ScreenshotNeo API documentation for request options.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
ScreenshotNeo accepts cookie and consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each of those steps can be turned off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and responses report the page verdict and billing status in headers. Its MCP server offers take_screenshot, get_page_info, and capture_pdf for AI agents and MCP clients. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000.
Sign up for 1,000 free screenshots a month with no card.
Windows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallCrashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minuteQuick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




