Use unit tests to check isolated behavior and integration tests to check whether connected parts work together. For AI-generated code, choose the test level according to the behavior or boundary at risk—and treat AI-written tests as drafts until you have checked their assumptions, run them in the project, and confirmed they assert a real requirement. Passing generated tests alone does not prove the code is correct.
What each test type is meant to prove
Testing terminology varies by team. ISO’s 2025 overview of AI-system testing lists several levels, including unit/component, integration, system, system integration, and acceptance testing. Some teams use “unit” and “component” for roughly the same layer; agree on the boundary your project means before discussing coverage or test ownership. ISO/IEC TS 42119-2:2025
| Question | Unit/component test | Integration test |
|---|---|---|
| What is under test? | An isolated function or component and whether it behaves as required. | Connected components or services and whether they work together across a boundary. |
| What happens to dependencies? | External services are usually replaced with controlled mocks or stubs when those services are not the subject of the test. | The interaction being evaluated is exercised, using real or representative dependencies where feasible. |
| Typical setup and feedback | Usually quick and isolated, making this layer useful for deterministic logic and frequent runs. | Usually needs more configuration; it can expose contract, data-flow, configuration, and boundary problems. |
| What it can reveal in AI-generated code | Local logic errors, input-boundary mistakes, error handling, and transformation defects. | Incompatibilities and coordination failures that isolated tests may not reveal. |
| Important limitation | A test can pass while asserting the wrong behavior or mocking away the defect. | Environment and service variability can make tests slower or less stable, so keep the scope intentional. |
This distinction is about the question being asked, not whether a human or AI wrote the code. A unit test that mocks a dependency cannot establish that the real dependency interaction works; an integration test focused on that interaction can.
How to choose the right layer for AI-generated code
Start with requirements and observable behavior
Identify what the code must do, how success or failure will be observed, and which existing test command, framework, fixtures, and project conventions apply. For each behavior, distinguish what is specified from what remains undecided. A model should not silently turn an unspecified product decision into an expected result.
Recommended Free Tools
Use unit tests for deterministic local logic
When a component transforms inputs, applies rules, validates data, or prepares an LLM request, test that deterministic behavior in isolation. Include normal cases, invalid inputs, relevant error paths, and values on both sides of important boundaries. Substitute external services with controlled responses when the service itself is not the subject of the test.
For example, a unit test can verify that a request builder formats a prompt correctly, or that surrounding code handles a mocked timeout. It should not depend on a live network request merely to test those local decisions.
Add integration tests for important boundaries
Write an integration test when the interaction among components, APIs, tools, storage, or workflow steps is itself part of the requirement. Check the real contract or data flow that isolated tests cannot establish—for example, whether one component’s output is accepted and interpreted correctly by the next. Keep external dependencies controlled or representative when practical, but do not replace the very interaction the test claims to verify.
For systems that call nondeterministic AI services, separate deterministic handling from evaluation of the live interaction. Unit tests can check surrounding code against controlled responses; integration or system-level tests should assess actual interactions against explicit criteria suited to the application. For agentic systems, AWS recommends broader testing across prompts, tools, workflows, and AI behavior because isolated exact-match unit tests may miss behavioral failures. AWS guidance on testing agentic systems
Do these 3 things before closing this tab:
1Fix the driver behind crashes, sound loss and screen glitches2Repair Windows errors before they cause bigger problems3Scan for outdated or missing drivers - takes under a minuteHow to use AI to draft tests without inheriting bad assumptions
Microsoft’s VS Code guide puts the distinction plainly: “Adding tests to an existing project involves more than generating test code.” Its documented workflow supports the key practice: review proposed cases, fit them to the project, and inspect what actually runs. VS Code: Test existing code with AI
- Give the tool context. Provide the relevant code, agreed requirements, existing test command and conventions, and any fixtures or helpers it should reuse.
- Ask for cases before code. Request normal behavior, boundary values, invalid inputs, and relevant error cases. Ask it to identify ambiguous requirements instead of deciding them for you.
- Review and agree on the cases. Check that each expected result follows from a requirement or an intentional decision. Ask for test-only changes and explicit expected values where that suits the project.
- Inspect the assertions and dependencies. Confirm that tests check observable behavior rather than merely echoing implementation details, and that mocks have not replaced the behavior supposedly under test.
- Run the project’s tests. Use the actual test command and environment, then inspect failures, skips, and warnings instead of relying on a coding tool’s summary. Confirm that the intended code executes.
- Use coverage as a map, not a verdict. Coverage can point to untested code, but it does not establish that assertions capture requirements. Mutation testing can provide a stronger check of whether assertions detect deliberately introduced faults; it is still one measure, not proof of overall correctness.
- Keep useful checks in CI. Run automated tests on changes, especially for deterministic application logic, so regressions receive rapid feedback.
Why passing generated tests is not proof of correctness
A model can produce plausible tests that encode a mistaken interpretation, test only the implementation’s current behavior, or avoid the defect by mocking too much. Human review must establish that the expected outcome is actually required and that the assertion would fail if the relevant behavior were wrong.
Rank #4
There is also a broader challenge when testing AI systems: sometimes the expected result is difficult to define. ISO/IEC TR 29119-11:2020 identifies this as the “test oracle problem”—testers may find it difficult to determine expected results and therefore whether a test passed or failed. The guidance addresses AI-based systems generally, including black-box approaches and neural-network-specific white-box testing; it should not be confused with the narrower task of testing ordinary software merely because a code-generation model authored it. ISO lists the 2020 document as published and under review. ISO/IEC TR 29119-11:2020
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.What published results do—and do not—show
Benchmarks can illustrate the limits of generated tests, but their results belong to their stated setups. TestGenEval, an ICLR 2025 study, comprises 68,647 tests from 1,210 unique code-test file pairs. In that benchmark’s setup, its best-performing model, GPT-4o, averaged 35.2% coverage and an 18.8% mutation score. These are historical results for the paper’s evaluated setup, not a current model comparison or a general estimate of test quality. The study uses coverage and mutation score alongside pass metrics and describes test generation in large real-world projects as challenging. TestGenEval, ICLR 2025
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Best Value
NIST’s 2025 GenAI (Pilot) Code Challenge evaluates generated unit tests for elementary Python code. Its stated scope is a pilot; it does not establish performance across programming languages, large repositories, integration tests, or production systems. NIST GenAI (Pilot) Code Challenge
Those measures answer different questions: coverage indicates how much code tests execute, while mutation score checks whether tests detect certain deliberately altered behaviors. Neither tells you, by itself, whether a test encodes the right requirement or whether the whole application is correct.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




