Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

Use an AI agent as a junior test engineer: give it current framework documentation and written project rules, let it inspect the running application, and ask it to build and run one focused test before expanding coverage. Review its selectors, assertions, permissions, and code changes yourself. Agents can speed up drafting and debugging, but a passing generated test is not proof that the application works correctly.

What an AI agent can—and cannot—do in QA

An AI coding agent can turn acceptance criteria into draft test cases, inspect a live page to suggest locators, write browser steps, run a focused test, and use actual errors and screenshots to propose repairs. It can also derive boundary, negative, and regression cases from requirements, then help produce failure reports with reproduction steps and logs.

These are useful implementation patterns, not a guarantee of autonomous defect discovery. An agent can misunderstand the expected behavior, assert the wrong result, or write a test that passes while missing the bug. Treat its output as a proposal that must be checked against the user requirement and the application’s real behavior.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Good starting tasks

  • Draft one test for a clearly described user journey.
  • Explore a page and verify a proposed locator against the live DOM.
  • Run a failing test and explain the exception using its logs and screenshot.
  • Suggest negative cases from specific acceptance criteria.

Tasks that need boundaries

Do not give an agent unrestricted access to production accounts, sensitive credentials, destructive actions, or shared test data. Use isolated accounts and data, least-privilege credentials, and approval gates for operations that could affect real users or other test runs.

Prepare the agent before asking it to write tests

Put the rules the agent needs in a repository instruction file that it can read on each task. Selenium’s guidance on AI coding agents specifically recommends giving agents current references and project context; without that, they may reproduce obsolete Selenium 2 or 3 patterns. Keep the contract concise enough to stay usable, but specific enough that a new contributor could follow it.

Include a practical agent contract

  • Versions and references: identify the framework and language versions actually used by the repository, and link to the current API documentation for those versions. Ask the agent to verify unfamiliar methods against those references.
  • Commands: document how to install dependencies, start the app, run one test, run the whole suite, and collect failure artifacts.
  • Conventions: state where tests belong, how they are named, how fixtures and setup/teardown work, and which locator patterns the project uses.
  • Test environment: specify the base URL, required browser matrix, account and data setup, and how each run gets isolated state.
  • Safety: identify forbidden production actions, credential handling rules, and which operations require human approval.
  • Definition of done: require a focused test, repeated successful local runs, a reviewable diff, and a clear account of any remaining uncertainty.

Keep these instructions in version control so agent work and human work follow the same conventions. Add a rule to avoid arbitrary sleeps: the agent should wait for the state that the next step actually depends on.

Build one stable browser test at a time

  1. Choose a real user journey. Start with one workflow, such as signing in and opening an account page, and write down the observable result that proves the journey succeeded. Keep the initial test narrow enough that a failure points to a small part of the product.
  2. Let the agent inspect the running app. Provide an approved browser tool or allow a temporary exploration script against the test environment. Have the agent confirm the page structure and candidate locators in the live DOM instead of inferring them from a design mockup.
  3. Ask for one test and its reasoning. Request the test, the expected state, the locator choices, and any setup or data assumptions. Have it identify the source of each assertion in the acceptance criteria.
  4. Run that test by itself. Use the repository’s documented command, then inspect the actual exception, browser output, and screenshot if it fails. Give those artifacts to the agent for a proposed fix; do not ask it to “make it pass” without evidence.
  5. Repeat the test before expanding. A test that passes only once may be relying on a timing accident, stale data, or leftover session state. Run it repeatedly under the documented setup and investigate intermittent outcomes before adding more coverage.
  6. Review the diff and intent. Check that the test verifies the requested behavior, uses current APIs, waits for meaningful conditions, isolates state, and does not weaken an assertion simply to achieve a green run. Then review fixtures, permissions, and test data before merging.

For example, a test for a form should assert the user-visible result of a successful submission, not merely that the submit button was clicked. For an invalid submission, assert the documented validation response, not just that the page remained open. The expected behavior should come from the requirement, not the agent’s guess.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Choose Playwright or Selenium for the suite you need

Neither framework is automatically the right choice because an AI agent can generate code for it. Compare the existing project, required browsers, language, and debugging needs. Playwright documents a single API for Chromium, Firefox, and WebKit and recommends resilient locators and web-first assertions. Selenium’s guidance emphasizes WebDriver standards, current language bindings, Selenium Manager, explicit waits, stable locators, and WebDriver BiDi for browser events and network interception.

Decision point Playwright Selenium What to give the agent
Browser coverage One documented API for Chromium, Firefox, and WebKit. Cross-browser WebDriver workflows. The browser matrix your team actually supports and the command used to run each project.
Locators and waiting Resilient locators and web-first assertions are recommended; its generator prioritizes role, text, and test-id locators. Use stable locators and explicit waits for the condition needed by the next step. Preferred locator order and examples from your own application.
Debugging evidence Use the artifacts and failure output enabled in your test setup. Use the logs, screenshots, and browser evidence available in your setup; Selenium guidance also points to WebDriver BiDi for browser events and network interception. Where artifacts are saved and how the agent should report them.
Standards and project fit Choose it when its browser and language support fit the suite and team. WebDriver standards and existing bindings may be important to a suite already built around Selenium. Do not migrate frameworks as a side effect of asking for one test.

For either framework, have the agent consult the current documentation for the version in the project rather than relying on remembered examples. Keep framework selection grounded in coverage and maintainability, not the speed of the first generated test.

Keep selectors and waits from becoming flaky

Prefer locators tied to user-facing meaning

Prefer accessible roles and labels, meaningful text, stable IDs or names, and dedicated test IDs when appropriate. Avoid absolute XPath and generated CSS class names: they often encode incidental layout or implementation details rather than the control the user recognizes. Playwright’s code generator prioritizes role, text, and test-id locators, but generated selectors still need a human check for meaning and stability.

Wait for conditions, not elapsed time

Use the framework’s explicit or web-first wait for the condition the next action depends on: a result becoming visible, a button becoming enabled, or a navigation completing. Do not cover a race with an arbitrary sleep or blindly increase a timeout. Selenium’s project documentation puts the trade-off plainly: “A fixed sleep is either too short, and the test fails, or too long, and the suite crawls.”

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Control the source of intermittent failures

  • Give each run isolated accounts and test data; prevent parallel tests from modifying the same shared record.
  • Make setup and teardown explicit, and ensure a test does not depend on another test’s prior state.
  • When the failure is intermittent, preserve the exception and screenshot, then determine whether the cause is an application defect, environment problem, data collision, or test race.
  • Do not remove a meaningful assertion or add broad retries without understanding the failure. A test that conceals a real intermittent defect is not more reliable.

Use a review gate before merging agent-written tests

Review both the code and the test’s intent. A green result is useful only if the test was checking the correct thing under controlled conditions.

  • Behavior: Does each assertion correspond to an acceptance criterion or an explicitly documented invariant?
  • Selectors: Are they stable and understandable, and do they identify the intended control?
  • Synchronization: Does each wait target the required state instead of masking a race with delay?
  • Isolation: Can the test run alone and alongside other tests without shared-state collisions?
  • Safety: Are credentials handled safely, and are destructive or production actions blocked?
  • Maintenance: Does the test follow repository naming, fixtures, and ownership conventions?
  • Diff quality: Did the agent make only the requested changes, and can a reviewer understand them?

Automation tooling does not by itself make a well-designed suite. Keep ownership, setup, teardown, and conventions consistent as coverage grows.

Expand to cross-browser runs and CI deliberately

Once the focused test is repeatable, add the browsers and execution environments the product requires. Playwright supports Chromium, Firefox, and WebKit; Selenium supports cross-browser WebDriver workflows. Configure a deliberate browser matrix rather than asking an agent to add every possible combination. Run the focused test locally first, then add it to CI with controlled test data and retained failure evidence.

Parallel execution can shorten feedback time, but it also exposes collisions in shared accounts, records, and environment setup. Before increasing parallelism, make state isolation explicit and confirm the tests do not rely on ordering. When a CI run fails, feed the agent the exact command, browser, exception, logs, and screenshot; distinguish a product failure from an infrastructure or test-agent failure before changing the test.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Test the AI agent separately from the website

A browser test can fail because the product is broken, because the generated test is wrong, or because the agent’s own workflow failed. Keep those layers distinct. The OpenAI Agents SDK documents utilities for testing agent workflows, sandbox sessions, realtime sessions, and voice pipelines; those deterministic harnesses can be combined with browser-level QA so that a failure in the test agent is not mistaken for a defect in the product under test.

Troubleshoot common agent-generated test failures

Symptom Likely cause What to do
Code uses a method or API your installed framework does not have The agent used stale or mismatched documentation. Give it the project’s framework version and current documentation link; reject calls that cannot be verified in that reference.
Element lookup fails even though the page looks correct The selector may be brittle, incorrect, or based on a different page state. Inspect the live DOM and use a role, label, stable ID, name, or test ID that matches the intended element.
Test passes locally but fails intermittently or in CI A timing race, shared test data, leftover session, or environment difference may be involved. Use the actual exception and screenshot to identify the dependency; wait for the required condition and isolate state rather than adding a blind sleep.
Test passes but users still encounter the reported bug The assertion may verify the wrong outcome or a path that misses the failure. Compare the test step and assertion to the acceptance criteria, then reproduce the user-visible failure in the test environment.
One test breaks other tests or changes shared records Setup, teardown, credentials, or test data are not isolated. Use dedicated data and accounts, document cleanup, and require approval for actions that affect shared or production state.

Or skip the browser setup

For screenshot evidence, you can make one API request instead of installing and operating a browser capture stack. ScreenshotNeo is a website screenshot API and MCP server for developers; it is not a replacement for assertions and interaction tests in Playwright or Selenium. A screenshot can help an agent inspect a page or attach visual evidence, while the QA test must still verify behavior.

The API can remove cookie/consent banners, newsletter popups, and chat widgets before capture; each cleanup step can be turned off. Bot checks/CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers report the page verdict and billing status. Its MCP server provides take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients. Every plan includes every feature. The free plan includes 1,000 shots per month without a card; paid plans start at $5 for 3,000 shots.

Read the ScreenshotNeo API documentation for request options and response details. Example cURL request:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

These examples save or retrieve a screenshot; they do not run a browser test or validate an expected result. Use the response’s verdict and billing headers when deciding whether a capture is usable. ScreenshotNeo also offers PNG, JPEG, WebP, or PDF output, full-page capture with lazy images loaded, element capture by CSS selector, viewport and device settings, custom CSS or JavaScript, waits, request blocking, headers and cookies, caching, signed image links, async jobs, bulk capture, and a usage API. See ScreenshotNeo for the service and sign up free for 1,000 screenshots a month with no card.

Make the test suite more reliable as it grows

Expansion should follow evidence that the first flow is stable. Add the next journey or browser project in a reviewable change, preserve logs and screenshots for failures, and keep the rules file current when versions or conventions change. If the agent repeatedly proposes the same fragile pattern, improve the repository contract or examples rather than fixing each generated test in isolation.

Do not use unsubstantiated productivity or defect-detection percentages to set expectations. Whether an agent saves time depends on the project, test environment, and quality of review; the practical measure is whether the team gets maintainable tests and useful failure evidence without weakening coverage.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.