Do these 3 things before closing this tab:
1Fix the driver behind crashes, sound loss and screen glitches2Repair Windows errors before they cause bigger problems3Scan for outdated or missing drivers - takes under a minuteSome links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
Agent Mode in Vercel Labs’ agent-browser is a repeatable inspect–act–inspect loop. Open a page, ask for an interactive JSON snapshot, let the agent choose an element reference, perform an action such as click or fill, then take a fresh snapshot after the page changes. The current snapshot—not an old guess about the DOM—drives the next action.
This guide shows the documented command flow, installation options, local and hosted execution choices, JSON handling, reliability practices, and fixes for common failures. The project documentation is a mutable repository, so check the installed release before treating a flag or integration as permanent.
What Agent Mode actually does
Agent Mode is the agent-oriented way to drive the agent-browser CLI. Instead of hard-coding a long selector script, the agent reads a structured representation of the current page, identifies a target, and acts on the reference exposed by that snapshot.
- Open the URL.
- Inspect an interactive snapshot in JSON.
- Choose the target from the returned element references.
- Act with a command such as
click @e2orfill @e3 "input text". - Inspect again whenever the action may have changed the page.
The snapshot is the hand-off format between the browser and an AI agent. JSON output makes it practical for an orchestration layer to parse roles, labels, text, visibility, and references without scraping terminal prose.
#1 Best Overall
Why references must be refreshed
A reference belongs to the page state that produced it. Navigation, a modal opening, an AJAX update, or a re-render can change the interactive tree. Request a new snapshot after those events and use the new references; do not assume that @e2 still means the same control.
Selectors and semantic locators
References are not the only targeting method. The CLI also documents conventional CSS selectors and semantic locators based on role, label, text, placeholder, and other attributes. Use a semantic locator when the page exposes a stable accessible name; use a snapshot reference when the agent needs to discover the target dynamically.
Install the CLI and its browser
The project documents several installation routes. Choose one that fits your workstation or build system, then run the browser installer before the first session.
Global npm installation
npm install -g agent-browser
agent-browser install
Local project installation
npm install agent-browser
npx agent-browser install
A local install keeps the CLI version with your project. Invoke it through your package runner or an npm script so other contributors and CI use the same dependency declaration.
Homebrew and Cargo
The repository also documents Homebrew and Cargo installation paths. After installing the CLI through either route, run agent-browser install to download Chrome for Testing on first use.
Existing browsers and Linux dependencies
Existing Chrome, Brave, Playwright, and Puppeteer installations are detected automatically according to the project documentation. On Linux, a machine may still need system libraries. The documented installer variant is:
agent-browser install --with-deps
Building from source
Building the project yourself has separate prerequisites: Node.js 24 or newer, pnpm 11 or newer, and Rust. Those requirements apply to a source build, not simply to using a packaged CLI.
Quick wins for a faster PC:
Scan for outdated or missing drivers - takes under a minuteDriver Scan →Repair Windows errors before they cause bigger problemsFix Now →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Your first Agent Mode session
The smallest useful loop can be run directly in a shell:
agent-browser open example.com
agent-browser snapshot -i --json
agent-browser click @e2
agent-browser fill @e3 "input text"
agent-browser snapshot -i --json
Replace @e2 and @e3 with references that actually appear in the first snapshot. The example is intentionally two-stage: after the click or fill, inspect the changed page before selecting the next control.
Rank #2
Read the first snapshot as an agent
- Look for an interactive element whose role and accessible name match the task.
- Confirm that it is visible and relevant to the current state.
- Record its reference exactly as returned.
- Act once, then obtain a new snapshot.
If the task is “search for invoices,” an agent should find the search textbox by its label or role, fill it, submit, and then inspect the results page for the next target. It should not carry a reference from the pre-search page into the results view.
When to chain commands
Command chaining is useful when intermediate output is irrelevant—for example, opening a known URL and immediately requesting a snapshot. Run commands separately when output must be parsed to decide the next action. A separate command boundary also makes retries and logging easier to reason about.
Choosing a local browser or a hosted session
Local execution
Local mode is appropriate when your workstation, runner, or container can install and launch the browser. It keeps the browser process near the CLI and is usually the simplest path for development and a conventional CI machine.
The project describes a CLI-and-daemon architecture: the CLI communicates with a Rust daemon over CDP, and the daemon persists between commands. Chrome is the default engine, with a documented Lightpanda engine option. These implementation details can change between releases, so verify them against the version you install.
Separate sessions
The project documents separate browser sessions with distinct browser instances and state. Use separate sessions when two jobs must not share cookies, local storage, or navigation state. Give each worker its own session identity and close it when the job is complete.
Hosted or remote execution
When a serverless function, locked-down CI runner, or deployment environment cannot run a local browser, the repository documents integrations with Browserless, Browserbase, Browser Use, and Kernel. These are documented integration paths, not a guarantee of current availability, pricing, or service quality. Check each provider’s current terms and configure its required environment variables or provider flag before running a job.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Choose hosted execution only after checking three things: whether local browser installation is possible, whether the workflow needs a remote/serverless session, and whether the provider’s current operational and commercial terms fit the workload.
Make the loop reliable
Wait for the state you need
Do not treat a successful click as proof that the next control is ready. Take a fresh snapshot and verify that the expected role, label, or text is present and visible. If the application updates asynchronously, let the agent wait for the resulting state before acting again.
Prefer stable meaning over brittle markup
Use role, label, text, or placeholder locators when they describe the user-visible control. CSS selectors remain useful for a component with a stable contract, but generated class names and positional selectors are more likely to break after a redesign.
Rank #3
Keep actions small
One meaningful action per inspect cycle makes failures diagnosable. If a multi-step command fails, split it at the point where the page can change so the agent can re-read the state.
Capture machine-readable output
Use --json for snapshots and other commands that an orchestrator must parse. Keep raw output in logs during development; it shows whether the problem is an incorrect reference, an unexpected page state, or a browser startup failure.
Close finished work
For a local quick start, the documented flow ends by closing the browser after the final assertion or extraction. In long-running workers, explicitly close or recycle sessions so stale state does not leak into the next job.
Common failures and precise fixes
“Browser executable not found”
Cause: Chrome for Testing has not been downloaded, or the runner lacks required Linux libraries.
Fix: Run agent-browser install; on Linux try agent-browser install --with-deps. Confirm that the process has permission to write the browser cache and launch a sandboxed browser.
Recommended Free Tools
A reference no longer works
Cause: The page re-rendered, navigated, or opened a modal after the snapshot.
Fix: Request agent-browser snapshot -i --json again and select a reference from the new response. Do not blindly retry the old @eN value.
The agent cannot identify a control
Cause: The control may be outside the interactive snapshot, hidden, unlabeled, or represented differently than expected.
Fix: Inspect the JSON for its role, visible text, and label. Try the documented semantic or CSS locator forms, and verify that a consent dialog, overlay, or disabled state is not covering it.
Rank #4
Actions run against the wrong page
Cause: Multiple commands or sessions share state, or a redirect completed between steps.
Fix: Use separate browser sessions for independent jobs, snapshot immediately after navigation, and include the current URL or page title in your agent’s decision context.
Local execution is impossible in CI
Cause: The runner cannot install a browser, lacks a display or required kernel features, or disallows long-lived browser processes.
Fix: Evaluate a documented remote integration such as Browserless, Browserbase, Browser Use, or Kernel. Verify the provider’s current setup and terms directly before committing production traffic.
Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minuteWindows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallSource build fails before the CLI runs
Cause: The source-build toolchain is older than the repository’s stated requirements.
Fix: Use Node.js 24+, pnpm 11+, and Rust for a source build, or install a published package instead.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Performance, isolation, and cost decisions
The biggest practical cost in an Agent Mode workflow is often unnecessary browser work: launching too many instances, repeating snapshots after every command when no decision is needed, or re-running a page that has not changed. Chain only deterministic commands whose output is not used; keep decision points separate.
Persistent daemons can reduce startup overhead across commands, while separate sessions prevent state contamination. Balance those properties: reuse a session within one coherent task, but isolate unrelated jobs and users.
The Tool Desk
Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Hosted execution moves browser maintenance to a provider but introduces provider-specific latency, limits, authentication, and pricing. The project documentation establishes the integration paths, not a universal performance or cost figure. Measure your own navigation, snapshot, and action timings under the provider and region you select.
Best Value
Or skip the browser setup
If your goal is a clean image or PDF rather than interactive browser control, ScreenshotNeo gives you a single HTTP request. It accepts cookie and consent banners as a visitor, removes more than 60 known consent platforms plus newsletter popups and chat widgets, and lets each cleanup step be turned off. Only clean shots are billed: bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and the response identifies the result with X-Page-Verdict and X-Billed headers.
Use the API documentation at https://screenshotneo.com/docs/ for the complete option list. A basic WebP capture is:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
Node.js
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
ScreenshotNeo also supports full-page captures with lazy images loaded, CSS-element capture, dark mode, 12 device presets and arbitrary viewports, retina scale, PDF paper sizes and page ranges, HTML/CSS rendering, custom JavaScript and CSS, pre-capture clicks, selector hiding, waits for selectors, delays or network idle, blocking ads/trackers/requests/resource types, custom headers, cookies, user agents and Authorization, timezone and geolocation, transparent backgrounds, resizing, selectable cache TTLs, signed public-image links, asynchronous jobs with signed webhooks, bulk capture of up to 100 URLs per call, a usage API, an OpenAPI specification, and compatible parameter names used by other screenshot APIs.
Free tools Windows power users keep installed
One-click scans. No signup required.
Its MCP server exposes take_screenshot, get_page_info, and capture_pdf to Claude, Cursor, and other MCP clients. Every feature is on every plan:
| Plan | Allowance and price |
|---|---|
| Free | 1,000 shots/month, no card |
| Starter | $5 for 3,000 shots |
| Growth | $15 for 15,000 shots |
| Pro | $39 for 60,000 shots |
| Scale | $99 for 250,000 shots |
| Business | $249 for 1,000,000 shots |
Yearly billing gives two months free. Create a free ScreenshotNeo account to get 1,000 screenshots a month with no card; paid plans start at $5 for 3,000.
FAQ
Is Agent Mode a separate executable?
No. It is the agent-oriented usage pattern documented for the agent-browser CLI: JSON snapshots provide the state and element references, and subsequent CLI commands act on that state.
Can I use a browser other than Chrome?
The project describes Chrome as the default engine and documents a Lightpanda engine option. Confirm the exact flag and support in the release you install.
The Tool Desk
Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Are hosted integrations included with the CLI?
The repository documents integration paths, but it does not establish provider pricing, service levels, or continuing availability. Those details belong to each provider’s current documentation.
Is the repository documentation a versioned API contract?
No. It is a mutable project repository. Pin your CLI dependency where reproducibility matters and verify commands against that pinned release.
Frequently Asked Questions
Does Agent Mode require an AI model built into the CLI?
The documented pattern treats the CLI as the browser-control layer; an external agent or orchestrator interprets JSON snapshots and chooses the next command.
Can one worker safely automate unrelated accounts at once?
Use separate browser sessions with distinct instances and state, rather than sharing one session across unrelated jobs.
Free tools Windows power users keep installed
One-click scans. No signup required.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

