Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

Playwright is optional. An AI browser agent can control Chromium through the Chrome DevTools Protocol (CDP), connect to a live browser with Chrome DevTools MCP, use the standards-based WebDriver BiDi protocol through Selenium, drive Chrome or Firefox with Puppeteer, or run inside an agent-focused framework such as Browser Use. Choose based on browser coverage, event handling, deployment, and how much control you need—not on whether a library is fashionable.

What replaces Playwright in an AI browser agent?

A browser agent has three separate jobs: obtain a browser session, send actions and JavaScript to that session, and observe enough page state to decide the next action. Playwright bundles those jobs behind one API, but none of them requires Playwright.

  • CDP gives direct, Chromium-specific control over pages, targets, network traffic, console output, storage, screenshots, and performance domains.
  • Chrome DevTools MCP exposes a live Chrome instance to an AI client through Model Context Protocol tools.
  • WebDriver BiDi is the W3C bidirectional protocol for cross-browser automation and event streams; Selenium supplies a mature client.
  • Puppeteer is a JavaScript driver that can use CDP or WebDriver BiDi, so an existing Puppeteer codebase can become an agent tool layer.
  • Browser Use is an agent-oriented runtime that documents both local Chrome-profile reuse and cloud browsers reachable over CDP.

The agent itself still needs a policy loop: inspect the page, choose an action, execute it, verify the result, and stop or ask for approval when an operation is irreversible. Replacing Playwright changes the transport and driver, not that control loop.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Choose the control path by requirement

Route Browser coverage Events and diagnostics Best fit Main trade-off
Chrome DevTools MCP Chrome/Chromium DOM inspection, JavaScript, screenshots, network and performance inspection through an MCP-connected live browser An AI client that should operate an already-open Chrome session Attached tabs can contain highly sensitive data; it is not a cross-browser contract
Direct CDP Chrome/Chromium Low-level control of targets, runtime, network, console, storage, performance and captures Chromium-specific services and maximum control Vendor-specific protocol and more plumbing to maintain
Selenium with WebDriver BiDi Browsers implementing the standard, subject to current driver support Bidirectional, event-driven network, console and JavaScript-error streams Cross-browser products and standards-oriented teams Feature support varies by browser and client version
Puppeteer Chrome first; Firefox support depends on protocol and current release High-level JavaScript API over CDP or WebDriver BiDi JavaScript teams with Puppeteer knowledge or Chrome-heavy workloads Protocol choice and feature parity are version-sensitive
Browser Use Local or hosted browsers, depending on deployment Agent-level abstractions; inspect which underlying protocol a feature uses Teams wanting an agent runtime instead of a hand-written driver Do not assume portability or anti-bot behavior without checking the specific implementation

Route 1: connect an AI agent through Chrome DevTools MCP

Chrome’s official “Get started with Chrome DevTools for agents” guide describes its MCP server as one that “Connects your AI agent to a live browser instance.” The agent can inspect the DOM, evaluate JavaScript, take screenshots, inspect network activity and diagnose performance while a human-visible Chrome session remains open.

When MCP is the shortest path

  • You already use an MCP-capable client such as Claude, Cursor or another MCP host.
  • The task benefits from seeing the same page a developer sees, including an authenticated tab.
  • You want browser inspection tools without designing a custom CDP message router.

Safe setup

  1. Install and configure the Chrome DevTools MCP server using the current Chrome DevTools for agents instructions.
  2. Start a dedicated Chrome profile, not your everyday profile. Keep it separate from personal mail, payment accounts and administrator sessions.
  3. Give the agent the narrowest account permissions possible. Require explicit approval before deleting data, submitting purchases, changing permissions or sending messages.
  4. Close the MCP connection when the task ends and discard the temporary profile when credentials should not persist.

An attached agent can read or modify page content and may reach cookies, local storage and authenticated tabs. Treat the MCP connection as equivalent to handing someone control of that browser.

Route 2: use CDP directly

CDP is the native debugging and automation interface for Chrome and Chromium. It exposes domains such as Runtime, Page, Network, Console and Performance. Your agent can call those methods itself or wrap them as tools for an LLM.

Start an isolated Chromium instance

google-chrome --headless=new 
  --remote-debugging-address=127.0.0.1 
  --remote-debugging-port=9222 
  --user-data-dir=/tmp/agent-chrome-profile

Use a temporary profile for automation. In a production service, run the browser in a container or sandbox with restricted filesystem and network access. Headless mode is appropriate for background jobs; a visible window is useful while developing and approving actions.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Minimal Python CDP client

Install requests and websocket-client. Open a page in the isolated browser, then run:

import itertools
import json
import requests
import websocket

# The first page target is only an example; select by URL or title in production.
targets = requests.get("http://127.0.0.1:9222/json/list", timeout=5).json()
page = next(t for t in targets if t.get("type") == "page")
ws = websocket.create_connection(page["webSocketDebuggerUrl"], timeout=10)
counter = itertools.count(1)

def cdp(method, params=None):
    request_id = next(counter)
    ws.send(json.dumps({"id": request_id, "method": method, "params": params or {}}))
    while True:
        message = json.loads(ws.recv())
        if message.get("id") == request_id:
            if "error" in message:
                raise RuntimeError(message["error"])
            return message.get("result", {})

cdp("Runtime.enable")
result = cdp("Runtime.evaluate", {
    "expression": "document.title",
    "returnByValue": True
})
print(result["result"]["value"])
ws.close()

For an agent, expose narrowly scoped functions such as get_dom_summary, click_css, fill_input and evaluate_readonly instead of allowing arbitrary expressions. Subscribe to CDP events to stream network failures or console errors back into the agent’s observation step. CDP’s Chromium scope is its advantage for Chrome diagnostics and its limitation for a browser matrix.

Route 3: WebDriver BiDi through Selenium

Selenium describes WebDriver BiDi as the W3C standard bidirectional protocol for browser automation. MDN describes it as event-driven communication between the automation client and browser. Unlike classic request-and-response WebDriver, BiDi keeps a WebSocket open so the client can receive network requests, console messages and JavaScript errors as they occur.

Why BiDi is the standards-first choice

  • Use it when the product must support more than Chromium and you want a protocol designed for multiple browser vendors.
  • Use event subscriptions to let the agent react to failed requests, console exceptions or page lifecycle changes instead of polling.
  • Keep browser-specific fallbacks where a feature is not yet implemented consistently across current drivers.

Selenium’s BiDi APIs and browser support are version-sensitive. Pin Selenium and browser-driver versions in CI, run a smoke test for every event you depend on, and consult the current client documentation before upgrading.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Agent design with BiDi

  1. Create a driver with a dedicated profile and explicit headless or visible mode.
  2. Open a page and subscribe to only the event domains your policy needs, such as network or log events.
  3. Convert raw events into compact observations: URL, status, console level, selector and timestamp.
  4. Give the model structured actions with allow-lists for domains, selectors and mutation types.
  5. Unsubscribe and quit the driver in a finally block so credentials and browser processes are not left behind.

Route 4: Puppeteer without Playwright

Google’s automation guidance describes Puppeteer as a JavaScript library that controls Chrome through CDP or WebDriver BiDi, and its guidance demonstrates Firefox automation with BiDi as well. Puppeteer is therefore a practical alternative when your agent service is already Node.js-based.

CDP-oriented Puppeteer example

import puppeteer from 'puppeteer';

const browser = await puppeteer.launch({
  headless: true,
  userDataDir: '/tmp/agent-puppeteer-profile',
  args: ['--no-sandbox'] // use only in an appropriately isolated container
});

try {
  const page = await browser.newPage();
  await page.goto('https://example.com', {waitUntil: 'networkidle2', timeout: 30000});
  const observation = await page.evaluate(() => ({
    title: document.title,
    url: location.href,
    headings: [...document.querySelectorAll('h1,h2')].map(e => e.innerText.trim()).slice(0, 20)
  }));
  console.log(JSON.stringify(observation));
} finally {
  await browser.close();
}

When a current Puppeteer release supports your required BiDi feature, select that protocol explicitly in the launch options documented for that release and test the same workflow in Firefox and Chrome. Do not assume that a CDP-only method, an extension API or a timing behavior has identical BiDi support.

Route 5: an agent-oriented runtime such as Browser Use

Browser Use documents reusing a local Chrome profile and connecting to hosted browsers through CDP. This can remove much of the action-planning and session-management code you would otherwise write. It is useful when the product requirement is “let an agent complete a browser task” rather than “build a portable browser-driver abstraction.”

Before adopting it, identify the protocol behind each required feature, where the browser runs, how credentials are injected, and what logs contain. A hosted CDP connection can simplify scaling but adds a trust boundary and network dependency; a local profile simplifies access to an existing session but increases the impact of a compromised agent.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Build the agent loop without a Playwright dependency

  1. Observe: collect URL, title, visible text, relevant accessibility or DOM structure, recent console errors and network failures.
  2. Plan: ask the model for one bounded action with a reason and a verification condition, not an unrestricted script.
  3. Act: execute a typed action such as navigate, click a selector, fill a field, press a key or capture a screenshot.
  4. Verify: check a URL change, selector state, response status or page text. Retry only when the failure is transient and the action is idempotent.
  5. Escalate: pause for human approval before irreversible or high-impact operations.

Use stable selectors and explicit waits. A fixed sleep can hide a race; waiting for a selector, a lifecycle event or a network condition gives the agent evidence that the page is ready. Redact tokens, cookies and personal data from model-visible logs, and cap page text and screenshot dimensions to control context size.

Reliability, performance and deployment decisions

Session isolation

Give each job a fresh profile or a narrowly scoped persistent profile. Never point an agent at a personal profile containing unrelated authenticated tabs. Store credentials outside prompts, rotate short-lived tokens, and limit outbound network access where possible.

Headless versus visible

Headless browsers are efficient for workers and scheduled jobs. Visible mode is better for debugging, consent flows and human approval. The Chrome configuration documentation supports headless operation, but rendering, extensions and window-dependent behavior should be tested in the mode you deploy.

Retries and observability

  • Set navigation, connection and action timeouts separately.
  • Capture the URL, action, protocol error, console errors and a redacted screenshot on failure.
  • Retry navigation or a read operation with bounded exponential backoff; do not blindly retry payments, deletes or form submissions.
  • Pin browser, driver, client and MCP versions, then run a smoke suite after each update.

There is no authoritative speed, reliability or cost percentage that applies across these protocols. Browser version, page complexity, network conditions, model latency and hosting dominate the result, so measure your own workflow rather than selecting a protocol from an unqualified benchmark.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Common failures and fixes

The agent cannot connect

Confirm that Chrome was launched with remote debugging, that the port is reachable from the worker, and that a firewall or container network is not blocking it. For MCP, verify the MCP host loaded the server configuration and that the browser target is still open.

A target or page is missing

CDP exposes multiple targets, including extensions and workers. Query the target list and select a type of page whose URL matches your job; do not assume the first target is correct.

Events never arrive in BiDi

Check that the client and browser driver both implement the event domain, that the subscription was created before navigation, and that your callback is being serviced by the client’s event loop. Test one event in isolation before adding several subscriptions.

Actions run before the page is ready

Replace arbitrary delays with a selector wait, a lifecycle/network-idle condition or an application-specific readiness marker. Capture the DOM and console output at the failure point to distinguish a race from a page error.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Login works locally but fails in automation

Use a dedicated profile and confirm the account is allowed to automate. Expect MFA, bot checks and device-bound credentials to require an explicit human step; do not attempt to bypass them. Keep the agent paused until approval is provided.

Browser updates break the workflow

Pin compatible browser, driver and client versions, maintain a small cross-browser smoke test, and isolate protocol-specific code behind your own interface. This lets you move from CDP to BiDi for one capability without rewriting the agent policy layer.

Or skip the browser setup

If the task is to obtain a clean website image rather than operate a full interactive session, ScreenshotNeo is the first screenshot API to try: it removes consent banners, newsletter popups and chat widgets before capture, and only clean shots are billed.

One GET request returns PNG, JPEG, WebP or PDF. The API reports whether the result was a clean page, a bot check/CAPTCHA, a blank page, a timeout, a failed load or a cache hit through the X-Page-Verdict and X-Billed headers. Those non-clean outcomes and cache hits cost nothing.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

cURL

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

See the ScreenshotNeo documentation for request options. The service also provides full-page and element captures, device and viewport controls, dark mode, custom CSS and JavaScript, waits, request blocking, cookies and headers, timezone and geolocation, resizing, selectable PDF output, caching with your TTL, signed image links, asynchronous webhooks, bulk capture for up to 100 URLs per call, a usage API, an OpenAPI specification and an MCP server with take_screenshot, get_page_info and capture_pdf tools.

Every plan includes every feature: 1,000 screenshots per month are free with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account and use the browser agent for interaction while delegating clean, reliable captures to the API.

Bottom line

Use Chrome DevTools MCP for the most direct live-Chrome agent connection, CDP for deep Chromium control, WebDriver BiDi through Selenium for a standards-oriented cross-browser system, Puppeteer for a JavaScript-first driver, and Browser Use when you want an agent runtime. Secure the browser session first, keep protocol code behind a small interface, and verify every model-chosen action.

Frequently Asked Questions

Can an AI agent control an already logged-in browser without Playwright?

Yes. Chrome DevTools MCP can connect to a live Chrome instance, and direct CDP can attach to a debugging target. Use a dedicated profile and least-privilege account because the agent can access that session’s page data and storage.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Is WebDriver BiDi a drop-in replacement for every CDP feature?

No. BiDi is the cross-browser standard, while CDP exposes Chromium-specific domains and diagnostics. Check current browser and client support for each capability you need and keep a fallback where parity is incomplete.

Which option should a JavaScript team with Puppeteer code choose?

Keep Puppeteer and select CDP or WebDriver BiDi per workflow. CDP is usually the simpler Chrome-first path; BiDi is preferable when the same agent must run across browsers and the required features are implemented by your target versions.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.