Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The fairest way to evaluate a browser agent is a layered test, not a single leaderboard. Start with benchmarks that match the agent’s operating surface—WebArena for realistic self-hosted websites, WebVoyager for live sites, WorkArena for ServiceNow knowledge work, and OSWorld or OSWorld 2.0 when the agent must control a complete desktop. Then add a private holdout set drawn from production traces.

Score the intended end state with a programmatic check, and report pass rate together with actions, latency, cost, retries, interventions and safety incidents. Freeze the model, prompt, tools, browser image, website state, step limit and reset procedure so that another team can reproduce the run.

Why one browser-agent leaderboard is not enough

Browser automation is not one task. An agent that fills a form on a live news site is solving a different problem from an agent that changes a record in ServiceNow or moves files between desktop applications. A leaderboard number combines task difficulty, website state, action interface and evaluator design. Treating the number as a universal ranking can produce the wrong production choice.

Use public benchmarks to establish a common reference, then test a private set that represents your own users, permissions, data and failure costs. Report public and private results separately.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Choose the benchmark that matches your product

Benchmark Operating surface What it reveals Important qualification
WebArena Realistic workflows on self-hosted websites Multi-step browser reasoning, navigation and state changes in reproducible environments Not a live-web test; results are not directly comparable with WebVoyager
WebVoyager Browsing on live websites Robustness to current, changing web pages Tasks are generally simpler than WebArena tasks
WorkArena ServiceNow enterprise workflows Knowledge-work actions, forms, records and permissions in an enterprise application The suite contains 33 enterprise tasks
OSWorld Full operating systems, desktop applications, web apps and file I/O Cross-application control and execution over a complete desktop The original study describes 369 tasks and programmatic execution scripts
OSWorld 2.0 Long-horizon computer-use workflows Stateful profiles, authentic artifacts, safety behavior and performance over long trajectories The 2026 release contains 108 workflows and adds comparisons by turns, actions, output tokens and cost

Map tests to risk tiers

Define your production task distribution before selecting a score. A low-risk tier might read public pages; a medium-risk tier might create or edit records; a high-risk tier might send money, publish content or expose personal data. Put each tier on the benchmark whose interface and side effects resemble it, and reserve a private task for behavior that no public suite captures.

Do not mix scores without explaining the environment

An OSWorld percentage measures full-desktop control, while WebArena and WebVoyager measure browser workflows. Even the two browser suites differ in self-hosted versus live sites and in task difficulty. A comparison is meaningful only when models saw the same task instances, interface and limits.

Freeze the evaluation harness

Write a versioned manifest for every run. Change one variable at a time; otherwise a score increase cannot be attributed to the model.

  • Model: exact model name, release or checkpoint and serving configuration.
  • Instructions: system prompt, task wording, examples and any planning policy.
  • Tool schema: action names, argument types, accessibility-tree or pixel observations, coordinate system and screenshot resolution.
  • Environment: browser and operating-system image, extensions, viewport, locale, timezone, network policy and website versions.
  • State: account permissions, seed data, cookies, files and starting URL. Use isolated credentials and prevent unintended side effects.
  • Limits: maximum turns or actions, wall-clock timeout, token or compute budget and retry policy.
  • Reset: deterministic setup and teardown scripts that return every task to the same initial state.
  • Randomness: record seeds and any sampling parameters; run enough repetitions to expose variance.

Store the complete trajectory: observations, actions, timestamps, tool errors, retries, model messages and final state. Redact secrets before retaining or publishing logs.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Make the end state the primary score

Count a task as passed only when the intended outcome is verified programmatically. For example, check that the expected record exists with the correct fields, that a file has the required checksum, or that an order reached the specified status. A judge reading the transcript can provide a diagnostic label, but should not silently replace an execution check when one is available.

Use partial credit only as a diagnostic

Record milestones such as reaching the right page, selecting the correct account or entering valid values. These signals explain where an agent fails, but partial progress must not turn an incorrect final state into a success.

Report the metrics that expose brittleness

Metric Definition Why it matters
Task success Verified end states divided by attempted tasks The headline measure of useful automation
Actions or turns Tool actions until completion or termination Shows efficiency and how much opportunity exists for an error
Wall-clock latency Elapsed time from first observation to verified outcome Reveals user-visible delay; publish median and tail values
Token or compute cost Model and tool consumption per task Connects quality to operating expense
Retry rate Tasks requiring a model or tool retry Separates apparently successful but fragile runs
Human intervention Runs requiring takeover, approval or correction Measures how much automation is actually delivered
Safety incidents Unauthorized, harmful or policy-violating actions A high success rate is unacceptable if risky actions are hidden

Publish aggregate results and per-task results. Include confidence intervals for pass rates, preferably with a method suited to binary outcomes such as a Wilson interval or a bootstrap over tasks. For repeated trials, state whether the interval treats trials or task identities as the sampling unit.

What published results really say

Published numbers illustrate why context matters:

  • OpenAI reported 38.1% on OSWorld, 58.1% on WebArena and 87.0% on WebVoyager for its Computer-Using Agent (CUA) in 2025. The same report notes that WebVoyager tasks are generally simpler than WebArena tasks, so the percentages are not an apples-to-apples ranking.
  • The original OSWorld study reported human success above 72.36% and best-model success of 12.24% in 2024 across 369 tasks. Its task examples include detailed initial states and custom execution-based evaluation scripts.
  • WebArena reported 78.24% human success versus 14.41% for the best GPT-4 agent in the 2023 study, demonstrating a substantial gap on realistic, reproducible web work.
  • WorkArena’s 2024 publication evaluates 33 enterprise tasks and reports that current agents remain considerably short of full task automation.
  • OSWorld 2.0, released in 2026, adds 108 long-horizon workflows, authentic artifacts, stateful user profiles, safety reports and comparisons by turns, actions, output tokens and cost.

These figures are historical measurements under each project’s own harness. They are useful baselines, not guarantees for your browser version, websites or prompts.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A repeatable evaluation runbook

  1. Describe the production distribution. List task types, frequency, permissions, data sensitivity, expected duration and the consequence of an incorrect action. Assign each type a risk tier.
  2. Select coverage. Map browser-only work to WebArena or WebVoyager, ServiceNow work to WorkArena, desktop and cross-application work to OSWorld, and long-horizon or safety-sensitive work to OSWorld 2.0. Add private tasks for missing workflows.
  3. Build setup and teardown. Provision isolated accounts, seed records and files, pin browser and OS images, and verify the starting state before each trial.
  4. Standardize the interface. Give every model the same observation format, tool definitions, viewport, timeout, action cap and retry rules. If you test pixels and accessibility trees, treat them as separate conditions.
  5. Run repeated trials. Execute identical task instances for every model. Keep failed, timed-out and safety-stopped runs in the denominator; do not replace them with a rerun unless your protocol says why.
  6. Verify outcomes. Run the task’s execution checker, then classify partial milestones, tool failures and safety events. Human review is for diagnosis or genuinely non-programmatic outputs.
  7. Analyze tails and causes. Report median and high-percentile latency, action counts, intervention rate and a failure taxonomy, not just the mean score.
  8. Publish the manifest. Include model versions, prompts, tools, browser and website versions, task instances, step caps, seeds, exclusions, reset procedure and confidence intervals.
  9. Re-run after changes. A model update, browser update, website redesign or benchmark revision invalidates a direct comparison. Keep old results as historical records and label the new run.

Build a failure taxonomy

Use mutually understandable labels so engineering work follows from the report. Useful categories include wrong-page navigation, missed or misread state, invalid input, selector or coordinate error, tool/API failure, timeout, exceeded action cap, authentication or permission failure, unsafe action, and evaluator defect. Allow one primary and several contributing labels. Review full trajectories for a sample of each category to catch mislabeled end states.

Safety is part of capability

Define forbidden actions before running the agent: sending external messages, changing permissions, deleting data, making purchases or exposing secrets may require an explicit approval gate. Test whether the model recognizes ambiguous instructions, respects least privilege and stops when a requested action conflicts with policy. Report safety stops and near misses separately from ordinary task failures; otherwise a model can appear more successful simply by taking dangerous shortcuts.

Capture reproducible visual evidence

For browser tasks, save screenshots at key milestones so a failure can be inspected alongside the machine-verifiable state. A do-it-yourself setup can use a pinned browser image, fixed viewport and a scripted capture after each action. Keep screenshots tied to task ID, trial ID and timestamp; avoid capturing credentials or personal data in shared artifacts.

Or skip the browser setup:

ScreenshotNeo provides a website screenshot API and MCP server that can supply consistent visual artifacts without maintaining capture code. Its capture options include full-page screenshots with lazy images loaded, CSS-selector element capture, dark mode, 12 device presets or a custom viewport, retina scale, PDF output with paper size, margins, landscape and page ranges, custom CSS and JavaScript, clicks before capture, waits for selectors, delays or network idle, blocking ads, trackers, requests or resource types, custom headers, cookies, user agents and Authorization, timezone and geolocation, transparent backgrounds, resizing, configurable-TTL caching, signed links, asynchronous jobs with signed webhooks, bulk capture of up to 100 URLs per call, a usage API and an OpenAPI specification. Parameter names used by other screenshot APIs also work.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Before capture, it accepts cookie or consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups and chat widgets; each cleanup step can be disabled. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads and cache hits are not billed, and each response identifies the result with X-Page-Verdict and X-Billed headers. Its MCP server exposes take_screenshot, get_page_info and capture_pdf tools for Claude, Cursor and other MCP clients.

See the ScreenshotNeo documentation for authentication and options. A one-call capture is:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

ScreenshotNeo has a Free plan with 1,000 shots per month and no card; paid plans start at $5 for 3,000 shots. The complete plan schedule is:

Plan Price Included shots
Free $0 1,000 per month
Starter $5 3,000
Growth $15 15,000
Pro $39 60,000
Scale $99 250,000
Business $249 1,000,000

Every feature is available on every plan, and yearly billing gives two months free. Start with 1,000 free screenshots a month with no card. Cookie banners, popups and chat widgets are removed before the shot; bot checks, blank pages and failed loads are never billed; and the MCP server lets AI agents take screenshots.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Common evaluation failures and fixes

The score changes between runs

Check for live-site drift, randomized data, unpinned browser versions, leftover cookies or nondeterministic model sampling. Pin what you can, reset state, record seeds and use repeated trials.

A high pass rate hides poor user experience

Inspect action counts, retries, intervention rate and tail latency. Require the end-state checker and publish per-task results so easy tasks cannot mask failures on critical ones.

Models appear incomparable

Verify that they received identical task instances, observations, tools, step caps and timeouts. State explicitly when one result comes from WebVoyager and another from WebArena or OSWorld.

The evaluator marks an incorrect result as success

Replace page-text or screenshot-only judging with a checker that reads the authoritative database, file, application state or artifact. Add adversarial cases that try to satisfy the checker without completing the user’s intent.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Safety incidents are missing from the report

Instrument destructive and external actions, add approval gates, and log blocked, attempted and completed unsafe actions as separate outcomes.

FAQ

Should a private benchmark replace public suites?

No. Public suites provide an external reference; private tasks provide relevance and protect against overfitting. Use both and label them separately.

What if a task has no deterministic end state?

Define an auditable artifact or rubric before testing, use independent reviewers, and report reviewer agreement and unresolved cases instead of silently converting judgments into binary success.

How should long tasks be compared?

Report success by horizon bands and include turns, actions, output tokens, cost and safety events. A single average can hide failures that occur only after many decisions.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Frequently Asked Questions

Should a private benchmark replace public suites?

No. Public suites provide an external reference; private tasks provide relevance and protect against overfitting. Use both and label them separately.

What if a task has no deterministic end state?

Define an auditable artifact or rubric before testing, use independent reviewers, and report reviewer agreement and unresolved cases instead of silently converting judgments into binary success.

How should long tasks be compared?

Report success by horizon bands and include turns, actions, output tokens, cost and safety events. A single average can hide failures that occur only after many decisions.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.