Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
Use agent-browser through its own CLI; you do not add agent-browser commands to a Playwright Page object. Install the CLI, download a supported Chrome build, open a URL, take a snapshot, and act on the snapshot’s references or on selectors. The default agent-browser daemon is implemented with Node.js and Playwright, but the project says an end user does not need a separate Playwright installation (or Node.js installed just for the daemon). Playwright’s separate playwright-cli is a different tool with different installation and commands.
What “with Playwright” means
There are two practical meanings:
- Playwright inside agent-browser: the default Node.js daemon uses Playwright internally. You operate the documented
agent-browserCLI; you do not write a Playwright test or install Playwright separately for the basic workflow. The project README describes this as “No Playwright or Node.js required for the daemon.” See the agent-browser README and command reference. - Playwright’s own coding-agent CLI: Microsoft documents
playwright-clias a separate command-line product. It has its own package, prerequisites and command surface. Its setup is documented at playwright.dev/docs/getting-started-cli.
This article covers the first meaning: controlling a browser with agent-browser while understanding that Playwright is its default implementation.
Install agent-browser and a browser
Global installation
npm install -g agent-browser
agent-browser install
agent-browser install downloads Chrome for Testing, the browser build used by the normal local workflow.
Project-local installation
npm install agent-browser
agent-browser install
Use the project-local form when you want the dependency and its version recorded with an application rather than shared globally.
#1 Best Overall
Other distribution routes and Linux dependencies
The project also documents Homebrew on macOS and Cargo distribution. On Linux systems that need browser libraries, run:
agent-browser install --with-deps
These are ordinary CLI-use requirements. They are different from the requirements for building agent-browser from source, which the project lists as Node.js 24+, pnpm 11+ and Rust.
Your first agent-browser session
Run this sequence from a shell:
agent-browser open https://example.com
agent-browser snapshot
agent-browser click @e2
agent-browser snapshot
agent-browser close
- Open: starts a browser session and navigates to the URL.
- Snapshot: prints a compact representation of the page, including references such as
@e2. - Act: pass a current reference to
click, or use a selector. - Refresh: take another snapshot after navigation or a significant state change.
- Close: ends the session and releases the browser process.
The exact reference is page-dependent; never assume that @e2 will identify the same element on another page. If a click, modal dismissal or navigation changes the document, obtain a fresh snapshot before using old references.
Quick wins for a faster PC:
Repair Windows errors before they cause bigger problemsFix Now →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Clear out junk files and repair common Windows errorsFree Scan →Target elements with references, selectors and roles
Snapshot references
References are efficient for an agent loop:
agent-browser snapshot
agent-browser click @e4
agent-browser snapshot
Use references only while the page state that produced them is still valid. The project specifically recommends taking a new snapshot after dismissing an element that was covering the target.
CSS selectors
Selectors are useful when the markup has a stable identifier:
agent-browser click "#submit"
agent-browser fill "input[name=email]" "[email protected]"
agent-browser get text "main"
Quote selectors in the shell so characters such as #, brackets and spaces are passed unchanged.
Rank #2
Role-based finding
For accessible controls, use the documented role finder:
The Tool Desk
Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →agent-browser find role button click --name "Submit"
Role and name queries can be more resilient than a generated CSS path, but they still depend on the page exposing the expected accessible name.
A useful form-and-navigation workflow
agent-browser open https://example.com/login
agent-browser snapshot
agent-browser fill "input[name=email]" "[email protected]"
agent-browser fill "input[name=password]" "$PASSWORD"
agent-browser find role button click --name "Sign in"
agent-browser snapshot
agent-browser get text "main"
agent-browser screenshot login-result.png
agent-browser close
Keep credentials out of command history where possible (for example, read them from environment variables in a wrapper script). If the click causes a navigation or reveals validation errors, the second snapshot is the authoritative page state. get text reads content, while screenshot creates visual evidence.
Commands you will use beyond the first example
Pages and tabs
The CLI supports tab commands for working with multiple pages. Use the command reference in the official repository for the current subcommands and flags rather than relying on an older blog post.
Attach to an existing browser
connect can attach to a browser over CDP. This is useful when Chrome is managed by another process or hosted elsewhere, but the endpoint, authentication and lifecycle belong to that browser environment.
Free tools Windows power users keep installed
One-click scans. No signup required.
Waits and dynamic pages
For pages that render asynchronously, wait for the relevant state before acting. Prefer a selector or other documented wait condition over an arbitrary sleep when the page provides a reliable readiness signal. Then take a snapshot and continue. Exact wait syntax and flags are version-sensitive, so check the command reference for the installed release.
Snapshots are the control loop
An agent-browser automation loop is usually:
- Navigate or perform an action.
- Capture a snapshot.
- Choose a reference, selector or role/name target.
- Perform one action.
- Capture a new snapshot whenever the DOM, URL, dialog state or overlay may have changed.
This avoids a common failure: a reference points to an element from an earlier DOM. A page transition, cookie dialog, client-side re-render or dismissed overlay can all invalidate it. Treat references as short-lived handles, not permanent selectors.
When you actually want Playwright APIs
If your goal is a test script using browser.newPage(), fixtures, assertions, tracing or Playwright’s language APIs, use Playwright directly rather than assuming agent-browser exposes a bridge. The material documented for agent-browser supports its CLI workflow; it does not establish a supported API that embeds agent-browser commands into a user’s Playwright Page or Browser object.
If your goal is an AI-oriented command workflow supplied by Playwright itself, install and use playwright-cli according to Microsoft’s guide:
Windows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallOutdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchnpm install -g @playwright/cli
playwright-cli open https://example.com
playwright-cli snapshot
Do not mix playwright-cli commands with agent-browser commands in the same script without deliberately managing two separate tools. Playwright documents Node.js 20+ as a prerequisite for its CLI.
Runtime modes and browser compatibility
The normal daemon is Playwright-based. A project changelog entry dated March 3, 2026 describes an experimental native Rust daemon that uses direct CDP/WebDriver instead. The entry records important differences at that point: Firefox and WebKit are not supported in native mode, Playwright trace format and HAR export are unavailable, and network routing uses CDP Fetch rather than Playwright’s route API. Close the session before switching modes. Because this is an experimental, version-sensitive feature, verify the current status in the agent-browser changelog before depending on it.
Local versus remote browser execution
Local Chrome is the simplest path: installation, browser binaries and the daemon run on your machine. If a local browser is unsuitable—for example, a restricted CI runner—you can use a documented remote-browser route such as Browserbase. That is an optional deployment choice; credentials and provider setup are separate from agent-browser’s basic local installation.
Rank #4
Browser versions, updates and reliability
Keep the agent-browser package and its downloaded browser aligned. Playwright notes that browser binaries are tied to Playwright versions; after updating a Playwright-based tool, rerun the relevant browser installation when required. Pin the package version in a project, exercise the critical navigation path in CI, and retain snapshots or screenshots when diagnosing a failure.
For dependable automation:
- Use a stable URL and explicit readiness condition.
- Prefer role/name or stable IDs over brittle positional selectors.
- Refresh snapshots after navigation, dialogs, consent actions and major re-renders.
- Close sessions in cleanup code so crashed jobs do not leave orphaned browsers.
- Separate authentication data from logs and shell history.
Troubleshooting
“Command not found: agent-browser”
The package is not on your PATH, or you installed it locally but are invoking it as a global command. Confirm the installation location and use the project’s local executable mechanism, or install globally.
Browser executable or launch failure
Run agent-browser install. On Linux, retry with agent-browser install --with-deps so documented system libraries are installed. In a locked-down environment, use a supported remote browser instead.
Clicking a reference does nothing or reports a missing target
The snapshot reference is stale, the element is covered, or the page has not finished rendering. Take a new snapshot, wait for the target state, dismiss the covering element, then snapshot again before clicking.
Selector matches nothing
Check the current URL and snapshot, verify spelling and quoting, and confirm that the element is not inside a different page or frame context. A dynamic application may render the control only after an earlier action.
Text or screenshot shows an unexpected page
Capture a fresh snapshot and inspect the URL and visible state. Redirects, authentication failures, bot checks and client-side error screens can all produce a valid browser response that is not the content you expected.
Native mode loses a Playwright feature
That can be expected: the March 3, 2026 changelog description lists capability differences for the experimental Rust daemon. Return to the default daemon when you need Playwright traces, HAR export, Firefox/WebKit support or Playwright route behavior, subject to the current documentation.
Or skip the browser setup
For a plain website image or PDF, ScreenshotNeo is a direct HTTP alternative: it accepts a URL and returns a PNG, JPEG, WebP or PDF. It removes cookie/consent banners, newsletter popups and chat widgets before capture; bot checks, blank pages, timeouts, failed loads and cache hits are not billed. It also provides an MCP server with take_screenshot, get_page_info and capture_pdf for Claude, Cursor and other MCP clients.
Use the documented options and examples at screenshotneo.com/docs/. A one-call capture with cURL is:
Do these 3 things before closing this tab:
1Clear out junk files and repair common Windows errors2Scan for outdated or missing drivers - takes under a minute3Repair Windows errors before they cause bigger problemscurl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python:
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
The service reports page and billing outcomes in X-Page-Verdict and X-Billed headers. The Free plan includes 1,000 shots per month with no card; paid plans start at $5 for 3,000 shots. Every feature is on every plan. Sign up free for ScreenshotNeo.
Frequently Asked Questions
Does installing agent-browser install Playwright’s test runner?
No. The default daemon uses Playwright internally, but the documented agent-browser CLI is a separate interface. Use Playwright directly when you need Playwright test APIs.
Can I reuse an agent-browser snapshot reference tomorrow?
No. References describe a particular page state and should be reacquired after navigation or a meaningful DOM change.
Which CLI should I choose for an AI coding workflow?
Choose agent-browser for its documented snapshot-and-reference workflow; choose Playwright’s separate playwright-cli when you specifically want Microsoft’s coding-agent CLI.
Recommended Free Tools
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

