October DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsSlow PC?RecommendedPC slow today? Run a repair scan before it gets worseResolve common Windows issues and optimize system performance.Scan NowOctober DealsAmazon USDeal season is back - check today's better picksAmazon US: current deals, useful picks and tech finds.See Picks×
Skip to content
World desk6 min

How to Generate Software Tests With AI: A Practical Developer Workflow

AI can draft tests, but useful coverage depends on repository context, explicit behaviors, and careful review. Use this workflow to generate, run, debug, and improve candidate tests.
Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

AI can draft unit, integration, and end-to-end tests, but it cannot establish by itself what your software is supposed to do or whether its tests are meaningful. The reliable approach is to give an assistant the implementation, relevant repository conventions, and explicit behaviors to test; then inspect, run, debug, and extend its draft.

What AI can—and cannot—do when generating tests

AI coding assistants can produce candidate tests for several levels of a test suite. GitHub’s documentation describes generating unit and integration tests with Copilot, while Visual Studio Code documents prompts for unit, integration, and end-to-end tests, along with running and debugging tests in the editor.

Think of the output as a draft, not proof. The assistant can suggest test cases and code, but you remain responsible for deciding the intended behavior, checking that assertions actually protect it, and confirming the tests run against the real project. GitHub explicitly cautions that generated tests may not cover all scenarios.

How do I generate tests with AI?

  1. Choose the behavior to protect. Identify observable inputs and outputs, important interactions, errors, and boundaries. If a requirement is ambiguous, resolve it before asking for tests; implementation details alone may not reveal product intent.
  2. Give the assistant relevant context. Open or reference the code under test and nearby tests. State the language, test framework, naming conventions, fixtures, and mocking approach. GitHub recommends making existing tests available to help suggestions fit the project’s framework and conventions; VS Code documents including file context in prompts.
  3. Request a focused draft. Ask for tests of named behaviors and scenarios rather than “complete coverage.” For example: Write tests for [function/module] using [framework] and the conventions in [existing test file]. Cover [normal cases], [boundary cases], and [failure behavior]. Test public behavior rather than private implementation details. Return the test code and list any assumptions.
  4. Inspect each test before running it. Confirm it calls the actual code under test and checks a meaningful result. Review imports, fixtures, setup, teardown, and mocks. Watch for assertions that merely repeat the implementation’s logic or depend on incidental implementation details.
  5. Run tests using the project’s normal workflow. Use the usual test command or the IDE’s test runner. Separate syntax or setup errors from genuine behavior failures. VS Code documents running and debugging discovered tests through Test Explorer and the editor.
  6. Iterate against specific failures and gaps. If a test fails, first decide what the correct behavior should be. Then give the assistant the failure and relevant context and ask for a targeted diagnosis. Do not weaken an assertion just to make the suite pass. Add missing cases based on requirements, not just suggestions the assistant made.

Can AI write unit tests for my code?

Yes. Provide the function or module, its expected public behavior, and an existing test file if the repository has one. Ask for a small set of tests covering representative valid inputs, invalid inputs, and relevant boundaries. Tell the assistant which framework and conventions to use; otherwise it may choose unfamiliar imports, assertion styles, or fixture patterns.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

For example, for a function that accepts a page number, spell out behaviors such as a normal positive page, the first valid page, zero or negative input, and the expected response for invalid input. The exact cases must come from the function’s contract, not from a guess about what the product should do. Review whether the tests would fail if the function returned a plausible but incorrect value.

How do I get AI to test edge cases?

Name the boundaries and failure conditions explicitly. A request for “edge cases” alone is underspecified: the assistant may choose cases that are irrelevant while missing the ones that matter to your API or user workflow.

  • Input boundaries: minimum and maximum accepted values, empty values, and values just outside the allowed range.
  • Invalid input: malformed data, missing required fields, unsupported types, or invalid state transitions, as applicable to the contract.
  • Failure behavior: expected errors, rejected promises, timeouts, or unavailable dependencies where the code is responsible for handling them.
  • Interactions: meaningful combinations of inputs or component behavior that a unit test alone cannot establish.

VS Code’s documented prompt example explicitly asks for edge cases. Make the request concrete by listing the cases and expected outcomes, then verify those outcomes against requirements or established behavior.

Choose the right test level

Test level Ask the assistant to cover Review for
Unit A function or component’s behavior in isolation, including relevant boundary and error cases. Meaningful inputs and outputs, appropriate isolation, and assertions on public behavior.
Integration Interactions between connected modules or services that matter to a stated workflow. Whether the test exercises the intended integration rather than mocking away the behavior it is meant to verify.
End-to-end A user-visible workflow across the application, with explicit starting conditions and expected outcomes. Whether the scenario reflects an actual requirement and can run within the project’s existing setup.

These are different scopes, not interchangeable labels. A passing unit test does not establish that connected components work together, and a broad end-to-end test may not explain which component caused a failure. Ask for the level that matches the risk and use the project’s established testing tools.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How to tell whether generated tests are useful

For every test, ask whether it would catch a realistic regression. A test that executes a line without checking a meaningful outcome may raise coverage without protecting behavior. Coverage is useful for locating code that lacks tests, but a high percentage is not evidence that requirements are correctly tested.

  • Does the test call the real target code and assert the expected observable result?
  • Would it fail for a plausible incorrect implementation?
  • Does it reflect an actual requirement rather than an assumption inferred from the code?
  • Are mocks limited to dependencies that should be isolated?
  • Are setup and cleanup correct, and are names clear about the behavior under test?

Reviewing the test itself matters because generated tests can be unusable or misaligned. In a 2024 peer-reviewed study, Khalid El Haji, Carolin Brandt, and Andy Zaidman evaluated 290 Copilot-generated tests across a sample of 53 tests from open-source Python projects. In that study setup, approximately 45.28% of generated tests passed when an existing test suite was available; without one, 92.45% were failing, broken, or empty. These results concern that tool, language, sample, and study—not a universal failure rate for today’s AI assistants.

Test quality also affects benchmark conclusions. OpenAI’s 2026 audit of 138 difficult SWE-bench Verified tasks found material issues in test design and/or problem descriptions in 59.4% of those audited tasks, including tests that were too narrow or checked behavior absent from the task description. That is a warning about benchmark evaluation quality, not a measured rate of failure for everyday AI-generated tests.

Common problems and fixes

Symptom Likely cause What to do
Generated code does not compile or imports missing modules. The prompt lacked framework, language, or project setup details. Share the relevant test file and exact framework conventions; ask for a correction that fits the project rather than a new generic example.
Tests pass but do not catch an incorrect result. The assertions are weak, or the test duplicates implementation logic. State the expected public outcome and consider a plausible wrong result; strengthen the assertion so that wrong result fails.
Tests fail because mocks or fixtures are wrong. The assistant inferred setup or dependency behavior not established by the repository context. Provide the actual fixture, setup, or dependency contract and ask for a focused repair.
Many cases are missing despite a broad request. “Test everything” did not define requirements or boundaries. List valid, invalid, boundary, failure, and interaction scenarios explicitly; compare the resulting suite with the contract.
The assistant changes expected behavior to silence a failure. The fix optimizes for a green test rather than the requirement. Decide the intended result yourself, preserve the correct assertion, and request a diagnosis of the implementation or test setup.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup

If part of your test workflow requires a website screenshot—for example, capturing a rendered page for visual review—ScreenshotNeo can return a screenshot or PDF from one request. This is a separate website-capture tool, not a test generator.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

cURL example (see the ScreenshotNeo API documentation):

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

ScreenshotNeo removes supported cookie and consent banners, newsletter popups, and chat widgets before capture; bot checks, blank pages, and failed loads are not billed. Its MCP server offers screenshot tools for AI agents. The free plan includes 1,000 shots per month with no card, and paid plans start at $5 for 3,000 shots. Sign up for free.

Frequently Asked Questions

Can AI-generated tests replace code review or QA?

No. Treat them as candidate test code that still needs human review against the intended behavior.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Should I ask an AI to maximize test coverage?

Ask it to cover specified behaviors and risks. Coverage can reveal untested code, but does not show that assertions are meaningful.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

More from the Wire

  1. Shenzhen desk3 min
    HONOR Expands Beyond Smartphones With Humanoid Robot RevealHONOR said it unveiled its first humanoid robot at MWC 2026 and named shopping assistance, workplace inspections, and supportive companionship as intended uses. Later Robotics D1 claims and a reported…
  2. Cupertino desk5 min
    Apple Unveils AirPods Max 2: The Upgrade That Should Have Happened Years AgoAirPods Max 2 adds H2-powered audio features and Apple claims up to 1.5× more effective ANC, but its design, Smart Case, and 20-hour battery rating are unchanged. Wired lossless audio…
  3. Cupertino desk4 min
    Apple’s OLED Touch MacBooks Are Coming—but the Dynamic Island Is the Real GambleApple has not announced an OLED touchscreen MacBook, but reports point to high-end models arriving in late 2026 or early 2027. The reported Mac Dynamic Island could be useful, but…
Recommended PC Tool
Recommended PC Tool
Crashes, No Sound, or Screen Glitches?Free driver scan
PC Slower Than It Used to Be?Free scan - under a minute

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.