Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

To perform browser actions programmatically, start a browser session, navigate to a page, locate the target, act on it, wait for an observable result, and close the session. For most application interactions, use Playwright or Selenium WebDriver rather than coordinates or low-level browser commands. The example below uses Playwright with JavaScript to fill a form and verify its response.

Automate a browser interaction with Playwright

This runnable Node.js example opens a page, enters a search query, submits the form, checks the result, and closes the browser even if an assertion fails. Replace the example URL and accessible names with those used by your application.

Install the dependencies

In a new project directory, install Playwright and its Chromium browser:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
npm init -y
npm install playwright
npx playwright install chromium

Save the following as browser-action.mjs. It uses the demo search form at https://www.selenium.dev/selenium/web/web-form.html, whose textbox is labeled “Text input” and submit button is named “Submit.”

import { chromium } from 'playwright';

const browser = await chromium.launch({ headless: true });
const page = await browser.newPage();

try {
  await page.goto('https://www.selenium.dev/selenium/web/web-form.html', {
    waitUntil: 'domcontentloaded',
    timeout: 30_000
  });

  await page.getByLabel('Text input').fill('browser automation');
  await page.getByRole('button', { name: 'Submit' }).click();

  const message = page.locator('#message');
  await message.waitFor({ state: 'visible', timeout: 10_000 });
  console.log(await message.textContent());
} finally {
  await browser.close();
}

Run it with node browser-action.mjs. The script prints the result message after the page displays it. If the demo page changes, inspect its current labels and selectors and update the locators rather than relying on screen coordinates.

The browser-automation lifecycle

A reliable script makes each stage explicit. Selenium’s first-script example follows this same pattern: start a WebDriver session, navigate, locate controls, interact, read the result, then quit the driver. Selenium describes WebDriver as a way to “drive[] a browser natively”; it is an interface implemented through language bindings and browser-specific drivers. See the Selenium WebDriver documentation.

  1. Choose a control layer. Select a library or protocol that supports your language, browser, and execution environment.
  2. Start or connect to a session. Launch a local browser, or connect through your framework’s supported remote setup.
  3. Navigate. Open the target URL and wait for the page state your next step needs.
  4. Locate the target. Prefer role/name or label-based locators, then stable test IDs or selectors. Avoid coordinates when an element locator is available.
  5. Act. Click, fill, select, check, hover, drag, or send keyboard input as the task requires.
  6. Wait and verify. Wait for a meaningful element or state change, then assert or read the outcome.
  7. Clean up. Close the page, browser, or WebDriver session so repeated runs do not leave processes behind.

Choose the right browser-control approach

Playwright and Selenium cover common browser interaction and testing. The documentation does not establish a universal speed or reliability winner; choose based on language, browser support, existing project setup, and the tooling you need.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Approach Best fit Important qualification
Playwright library Common page interaction using page and locator APIs, with actions that wait for actionability conditions. Confirm current language and browser support in the Playwright Page API documentation.
Selenium WebDriver A language-neutral browser-control interface, browser-specific drivers, or local and remote sessions. Bindings and driver implementations are distinct parts of the setup; consult the official guide for your language and browser.
Chrome DevTools Protocol (CDP) Chromium/Blink-specific instrumentation, debugging, inspection, and profiling. The tip-of-tree protocol can change frequently and has no backward-compatibility guarantee. See CDP documentation.
WebDriver BiDi Bidirectional browser event streaming over WebSocket, such as network, console, and JavaScript error events through Selenium. Feature support is evolving and depends on implementation. See Selenium’s BiDi documentation.
Playwright MCP Tool-driven interactions for an agent or other MCP client, using accessibility snapshots and target refs or unique locators. This is a tool interface, not the same thing as calling the Playwright library directly. See Playwright MCP.

Target elements robustly

A locator describes which element to act on. Prefer selectors tied to user-facing semantics because they are usually clearer and less coupled to layout than CSS structure or screen coordinates.

  • Role and accessible name: Use a button role and its visible or accessible name, such as getByRole('button', { name: 'Submit' }).
  • Label: For form controls, use the associated label, as in getByLabel('Text input').
  • Test ID or stable selector: Use an application-provided test ID or stable CSS selector when semantics are unavailable.
  • Frames: If the target is in an iframe, first target its frame, then find the control inside it. Playwright provides frame locators for this purpose.

Playwright recommends locator-based actions for many interactions. A self-contained action such as locator.click() can reduce the risk that the page changes between separately checking an element and acting on it. Consult the Playwright API documentation for locator and frame methods.

Wait for outcomes instead of guessing

Fixed pauses are fragile: the same delay may be unnecessarily long on one run and too short on another. Prefer waits tied to the state you need, such as an element becoming visible, a result appearing, or a URL changing. Browser frameworks also impose action and navigation timeouts; set them deliberately for the expected environment.

After the action, verify something the user would recognize: a confirmation message, updated control state, changed URL, or loaded content. A click completing does not prove that the application accepted the action. In the example, the script waits for the result element before reading it.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Other common browser actions

Select, check, and hover

For a select menu, checkbox, or hover target, use the framework’s corresponding locator action rather than simulating a sequence of mouse coordinates. Confirm the resulting selected or checked state where it matters.

Keyboard input and drag operations

Use keyboard APIs for shortcuts or input that genuinely depends on keyboard events. Use a framework’s drag-and-drop action for draggable elements, then wait for the destination state. Applications can implement these interactions differently, so verify their outcome instead of assuming that dispatching input alone succeeded.

Take a screenshot

When debugging a failure or preserving a visual result, capture a screenshot after the relevant state is reached. A screenshot is evidence of what rendered, not a substitute for asserting that a workflow completed successfully.

Debug and handle unusual page structure

When a locator does not resolve, check whether the element is inside an iframe, whether the page has finished rendering it, and whether the accessible name or selector matches the current page. For tool-driven interaction, Playwright MCP can expose an accessibility snapshot and refs for interactive targets; CDP is another option when low-level Chromium inspection or profiling is needed.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Keep automation authorized and within the site’s rules. Browser-based collection may be technically possible but prohibited by a site’s terms, and sites may block automated access. Review applicable terms before automating scraping or collection; see Selenium’s guidance on browser automation use.

Troubleshooting common failures

Element not found or strict locator failure

Cause: The element has not appeared yet, the accessible name differs, a selector is stale, or multiple elements match. Fix: Inspect the current page and locator target, wait for the relevant state, and narrow the locator using role, name, label, or a stable test ID.

Click times out

Cause: The target is hidden, covered, disabled, or not yet actionable. Fix: Wait for the intended visible and enabled state; check for overlays or consent dialogs; use the correct frame if the control is embedded. Do not switch to a coordinate click until you understand why the locator action cannot proceed.

Script reads a blank or old result

Cause: The script reads immediately after an action, before the application updates. Fix: Wait for the expected message, URL, or state transition, then read or assert it.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Browser or driver does not start

Cause: A browser binary or driver is missing, incompatible, or unavailable in the runtime environment. Fix: Follow the current installation instructions for the framework and browser combination, and check the environment used by the script, especially in a container or remote runner.

Behavior differs across browsers

Cause: Browser engines, driver implementations, and protocol support are not identical. Fix: Verify the framework’s current browser support and test the actual target browser. For BiDi in particular, check whether the event or command you need is implemented in the chosen stack.

CDP integration breaks after an update

Cause: CDP’s tip-of-tree API can change without a compatibility guarantee. Fix: Pin and validate the browser/protocol combination you depend on, or use a higher-level API where low-level Chromium access is not necessary.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Performance, reliability, and operating cost

Reliable automation is usually about waiting for the right state and keeping sessions isolated, not adding longer arbitrary delays. Close sessions in cleanup code, use explicit timeouts, and capture enough context to diagnose failed runs. Remote sessions can fit distributed execution, while local sessions reduce the moving parts for a small script; which is preferable depends on your infrastructure and browser needs.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Browser automation consumes runtime and browser resources, so avoid launching unnecessary sessions and avoid retrying actions blindly when they may submit a form or cause another side effect. Make retries conditional on a known failure state and verify whether the first attempt took effect. No universal speed, cost, or reliability comparison between Playwright and Selenium is established by their documentation cited here.

Or skip the browser setup

If your goal is a page image or PDF rather than interacting with controls, a screenshot API avoids installing and maintaining a browser session. ScreenshotNeo is a website screenshot API and MCP server; its clean-shot options can accept cookie or consent banners and remove known consent platforms, newsletter popups, and chat widgets before capture. Each step can be turned off. Bot checks/CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, with response headers indicating the page verdict and billing status. Its MCP server provides screenshot and PDF tools for AI agents. The free plan includes 1,000 shots per month without a card; paid plans start at $5 for 3,000 shots.

One GET request returns an image or PDF; this cURL example saves a WebP screenshot of Stripe:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

See the ScreenshotNeo API documentation for request options, formats, and setup. Sign up for 1,000 free screenshots a month, with no card required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Frequently Asked Questions

Can I automate a browser without a graphical window?

Yes. The Playwright example uses headless Chromium, which runs without showing a browser window.

Is browser automation the same as scraping?

No. Automation describes controlling a browser; scraping is one possible use. Whether collection is allowed depends on the site’s terms and access controls.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.