Generative AI is a practical assistant for drafting and expanding software tests—but it is not a reliable substitute for running, reviewing, and maintaining them. In one 2024 study of Copilot-generated Python tests, only about 45.28% passed when generated within an existing test suite; without one, 92.45% were failing, broken, or empty. Those figures describe a specific study setup, not AI test generation across all tools, languages, or test types.
What generative AI can—and cannot—do for testing
An AI coding assistant can turn a function, a written behavior, or an existing test into a first draft of additional tests. That can reduce blank-page work and help surface cases a developer may want to check. The useful output is a candidate test, not proof that the software behaves correctly.
A generated test can fail to run, assert the wrong thing, duplicate assumptions in the implementation, or pass without detecting meaningful defects. A high test count or line coverage alone does not establish that a suite protects important behavior. Tests need to run in the project’s normal environment, and a person needs to judge whether their assertions express the intended behavior.
What the evidence actually shows
Copilot-generated Python tests had mixed results
El Haji, Brandt, and Zaidman’s 2024 empirical study evaluated 290 Copilot-generated tests in a setup involving 53 sampled tests from open-source projects. When generation took place within an existing test suite, approximately 45.28% of generated tests passed; the other 54.72% were failing, broken, or empty. Without an existing suite, 92.45% were failing, broken, or empty. These results concern the study’s Python tasks, sample, tool, and evaluation setup. They are not a general pass rate for today’s AI models or for integration, UI, or other kinds of tests. Read the study.
The comparison makes context and workflow relevant: generation with an existing suite performed differently from generation without one in this study. It does not show that providing context guarantees correct tests. The remaining failures, and the need to judge what a test checks, make execution and review essential.
A separate Copilot trial measured code functionality, not test-generation quality
GitHub reported a randomized trial with 202 developers who each had at least five years of experience and were asked to write API endpoints. Participants with Copilot access were 53.2% more likely to pass all 10 unit tests in that coding task. This is evidence about whether Copilot-assisted code passed the task’s tests—not whether AI-generated tests are reliable or good at finding defects. GitHub published the finding in 2024 and updated its article in 2025. Read GitHub’s account of the trial.
Evaluation work is still being developed
NIST’s 2025 pilot plan describes a way to measure and evaluate AI-generated unit tests for elementary Python code. It is an evaluation plan, not a result showing that a model performs well. The evidence described here does not establish a vendor-neutral ranking or settle performance across languages, complex projects, integration tests, UI tests, or security testing. See NIST’s AI-generated code work.
How to judge an AI-generated test
Before keeping a generated test, check whether it is both executable and meaningful. A passing test is not automatically a useful test: it may simply repeat an assumption already embedded in the code.
The Tool Desk
Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →- Validity: Does the test run in the project’s usual environment and assert the intended behavior?
- Defect-finding value: Would it fail if a known or seeded defect changed that behavior, or does it merely execute lines?
- Assertions and expectations: Are the expected outcomes justified by requirements or documented behavior, rather than copied uncritically from implementation details?
- Coverage of relevant cases: Does it consider meaningful boundaries, invalid inputs, or other edge cases for this function?
- Maintenance cost: Is the test understandable and stable enough to maintain, or will it be tightly coupled to implementation details?
- Context and scope: What code, existing tests, requirements, or comments informed it, and does the evidence apply to this language and test type?
How to run a responsible team pilot
- Choose a bounded starting point. Select low-risk functions the team understands, rather than starting with a broad or safety-critical test suite.
- Supply explicit behavior and edge cases. Ask the assistant to draft tests against the expected behavior. Include relevant context, such as requirements or existing tests, while treating the resulting tests as proposals.
- Run tests normally. Use the project’s existing test environment and workflow. Separate tests that do not run from those that run but do not assert the right behavior.
- Review and repair. Inspect expected values and assertions for tautologies, copied assumptions, missed edge cases, and dependence on private implementation details. Have a developer review changes as part of the ordinary engineering process.
- Compare against a baseline. Measure outcomes on comparable tasks, and break results down by language, task, and test type rather than treating a small pilot as a universal verdict.
- Track value as well as volume. Record the fraction of generated tests that are valid, defect-finding value where it can be assessed, time spent drafting and reviewing, maintenance effort, coverage, post-deployment bugs, and developer confidence. More generated tests alone is not a success measure.
- Check governance before sharing code. Confirm that your organization’s policies permit sending the relevant code and prompts to the service you plan to use. The sources cited here do not establish the current privacy terms of individual AI products.
These safeguards are a practical way to evaluate a pilot, not proof that one specific workflow will outperform another. GitHub’s rollout guidance likewise recommends setting goals, measuring outcomes, piloting changes, and retaining engineering judgment and code review. Read GitHub’s organizational evaluation guidance.
Use AI for assistance, not sign-off
Generative AI is most defensible as a way to accelerate test drafting and exploration while keeping responsibility for correctness with the engineering team. The available results are too narrow to support a general reliability claim: one study measured Copilot-generated Python tests in a defined setup, GitHub’s randomized task measured code passing tests rather than test quality, and NIST’s cited work is a pilot plan. For a team, the meaningful question is whether generated tests are valid, detect relevant defects, and save time after review and maintenance—not whether the assistant can produce them quickly.
Rank #4
Or skip the browser setup
If your testing workflow needs screenshots of web pages for visual checks, you can capture them with a browser automation setup—or use ScreenshotNeo, a website screenshot API and MCP server. A GET request can return a PNG, JPEG, WebP, or PDF. Its capture process accepts cookie or consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each of those steps can be turned off. Bot checks and CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and the response identifies page verdict and billing status in headers. AI agents can use its MCP server tools, including take_screenshot, get_page_info, and capture_pdf.
For example, this cURL request captures a page as WebP; see the ScreenshotNeo API documentation for request options and setup:
Recommended Free Tools
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
The free plan includes 1,000 screenshots per month with no card required; paid plans start at $5 for 3,000 screenshots. Sign up for ScreenshotNeo’s free plan.
Quick Recap
Best Value
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




