October DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsWindows FixRecommendedWindows errors stealing your time? Find the fix fastScan stability, cleanup and performance issues.Fix NowOctober DealsAmazon USDeal season is back - check today's better picksAmazon US: current deals, useful picks and tech finds.See Picks×
Skip to content
browser automation

Scalable Web Scraping with Playwright and Browserless: 2026 Guide

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Scale Playwright scraping by treating browsers as scarce, short-lived resources. Put every URL in a bounded queue, keep independent jobs in separate BrowserContexts, connect to Browserless with the protocol that matches your required features, and close every context and browser in a finally block. Your safe concurrency is the lowest of your application limit, the Browserless plan limit, and the request rate the target site permits.

When a browser is actually necessary

A browser is justified when the page needs JavaScript execution, client-side rendering, authenticated session state, clicks, scrolling, or other browser interaction. It costs more resources than an ordinary HTTP request because each job uses a browser process, pages, and session state. Use direct HTTP and an HTML parser for static endpoints; reserve Playwright for the parts that genuinely require rendering. No universal throughput figure applies: target behavior, page weight, JavaScript, network distance, and your plan all change the result.

The operating model: queue, isolate, release

Bound concurrency in your application

Playwright Test’s workers setting controls test worker processes. A production scraper needs its own queue and semaphore; changing the test setting does not protect a scraper from opening unbounded remote sessions.

import asyncio
from playwright.async_api import async_playwright

CONCURRENCY = 5  # Choose below your account and application ceiling

async def scrape_one(browser, url):
    context = await browser.new_context()
    try:
        page = await context.new_page()
        await page.goto(url, wait_until="domcontentloaded", timeout=45_000)
        return {"url": url, "title": await page.title()}
    finally:
        await context.close()

async def main(urls):
    semaphore = asyncio.Semaphore(CONCURRENCY)
    async with async_playwright() as pw:
        browser = await pw.chromium.connect(BROWSERLESS_WS_ENDPOINT)
        async def bounded(url):
            async with semaphore:
                try:
                    return await scrape_one(browser, url)
                except Exception as exc:
                    return {"url": url, "error": type(exc).__name__}
        try:
            return await asyncio.gather(*(bounded(url) for url in urls))
        finally:
            await browser.close()

# Set BROWSERLESS_WS_ENDPOINT from a secret, not source control.

The ceiling should be configurable. Start conservatively, observe queueing and failures, then increase only when the account, runner, and target workload can handle it.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use a context for each independent identity

A BrowserContext isolates cookies and storage while sharing the connected browser. Create one per account, tenant, or job when state must not leak between tasks. Playwright describes contexts as fast and cheap to create. Close each context after its pages finish, even when navigation or extraction fails.

Keep sessions short

Long-lived pages consume concurrency and make recovery harder. Navigate, extract, and release. If a workflow must be long, split it into explicit stages and record state outside the browser.

Connect Playwright to Browserless deliberately

Browserless supplies remote WebSocket browser endpoints; you connect to an existing browser rather than launching one on the local machine. Keep the token and complete endpoint in environment variables or a secret manager. Anyone who obtains a browser-server WebSocket path may be able to control the OS user behind it.

Native Playwright connection

Use Playwright’s native connect when you need Playwright-protocol features. Put the current Browserless endpoint, including its token, in BROWSERLESS_WS_ENDPOINT; endpoint paths and regional hosts can change, so copy the value shown for your account.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
import os
from playwright.async_api import async_playwright

async def run(url):
    endpoint = os.environ["BROWSERLESS_WS_ENDPOINT"]
    async with async_playwright() as pw:
        browser = await pw.chromium.connect(endpoint)
        try:
            context = await browser.new_context()
            try:
                page = await context.new_page()
                response = await page.goto(url, wait_until="domcontentloaded", timeout=45_000)
                return {
                    "final_url": page.url,
                    "status": response.status if response else None,
                    "title": await page.title(),
                    "html": await page.content()
                }
            finally:
                await context.close()
        finally:
            await browser.close()

CDP connection

Use connectOverCDP with Browserless’s CDP endpoint when you need Chrome DevTools Protocol and Browserless helper integrations. Native Playwright and CDP endpoints have different paths and feature support. Check the current Browserless feature matrix before relying on Firefox, WebKit, routing, API request contexts, extensions, or vendor-specific helpers.

import { chromium } from 'playwright';

const endpoint = process.env.BROWSERLESS_CDP_ENDPOINT;
const browser = await chromium.connectOverCDP(endpoint);
try {
  const context = await browser.newContext();
  try {
    const page = await context.newPage();
    const response = await page.goto(process.argv[2], {
      waitUntil: 'domcontentloaded', timeout: 45_000
    });
    console.log(JSON.stringify({
      url: page.url(),
      status: response?.status() ?? null,
      title: await page.title()
    }));
  } finally { await context.close(); }
} finally { await browser.close(); }

Choose a documented region close to your job runner when latency matters. Regional names and availability are account-dependent and can change.

Understand Browserless capacity and pricing limits

Browserless defines concurrency as simultaneous browser sessions. When the limit is full, new work can queue. Its pressure information distinguishes running, queued, and maximum values. A queue absorbs bursts; it does not add capacity. Sustained queue growth means you should reduce demand, add capacity, or redesign the workload.

Plan example Concurrent browsers Maximum session duration
Free 2 2 minutes
Prototyping (monthly / yearly) 5 / 10 15 minutes
Starter (monthly / yearly) 30 / 40 30 minutes
Scale (monthly / yearly) 80 / 100 60 minutes

These are Browserless figures documented in 2026, not throughput benchmarks. Pricing, quotas, session maximums, endpoint paths, and supported features are volatile; verify the live plan before sizing. Browserless also documents self-hosted defaults of 10 concurrent sessions and a queue length of 10, configurable with environment variables.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Make extraction resilient

Wait for the page you need

Prefer domcontentloaded, a specific selector, or a bounded delay that matches the application. Waiting for every connection to become idle can hang on analytics, ads, or streaming requests. Set navigation and action timeouts, and fail with a structured reason.

Retry only transient failures

Record URL, final URL, status, duration, attempt, and failure class. Retry timeouts, temporary connection failures, and documented service errors with bounded exponential backoff. Do not retry authentication errors, persistent 4xx responses, or selector-not-found errors indefinitely.

Close on every path

Use nested try/finally blocks for context and browser cleanup. Browserless specifically recommends closing sessions so they do not occupy concurrency after a job has ended.

Respect the target

Concurrency is not permission. Follow the site’s terms, robots guidance where applicable, authentication rules, and rate limits. A successful proxy connection or browser render does not establish that a scrape is allowed.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Proxy configuration without false promises

Playwright supports HTTP(S) and SOCKSv5 proxies at browser or context scope, with optional credentials and bypass hosts. Apply a proxy only for an authorized, documented requirement. It may change network location, but it does not guarantee access, prevent bot checks, or override a target site’s terms.

const browser = await chromium.launch({
  proxy: {
    server: process.env.PROXY_SERVER,
    username: process.env.PROXY_USER,
    password: process.env.PROXY_PASSWORD,
    bypass: 'localhost,internal.example'
  }
});

Observe the system before increasing parallelism

  • Track active sessions, queue depth, wait time, job duration, navigation status, timeout count, retry count, and extraction failures.
  • Separate Browserless capacity errors from target-site responses and application bugs.
  • Alert on sustained queue growth, rising session duration, and cleanup failures.
  • Use a request identifier in logs, but never log tokens, cookies, authorization headers, or complete signed endpoints.

Local Playwright or Browserless?

Decision axis Local Playwright Browserless
Browser installation and updates Your team owns images, browsers, patches, and scaling. The hosted service manages browser pools and isolation, according to its service documentation.
Concurrency Limited by your machines and scheduler. Limited by the account or self-hosted configuration; excess work can queue.
Geography and latency Controlled by where your runners operate. Choose an available regional endpoint near the workload.
Protocol Full local Playwright control. Select native Playwright or CDP; support differs.
Cost Infrastructure and engineering cost. Plan quota and session limits; verify current pricing.

Use local browsers when you need maximal infrastructure control and already operate reliable browser images. Use Browserless when outsourcing browser pool operations and account-level capacity is preferable. Compare total cost at your measured workload, not a claimed universal requests-per-second number.

Common failures and fixes

Connection rejected or times out

Check that the endpoint protocol matches the method (connect versus connectOverCDP), the token is current, outbound WebSockets are allowed, and the selected region is available.

Jobs remain queued

Inspect running, queued, and maximum pressure values. Lower application concurrency, shorten sessions, or move to a plan with more concurrency. Do not compensate with unlimited retries.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Pages are blank or incomplete

Wait for a stable selector or the application’s loading event, confirm the final URL, and capture console or network errors. A longer global timeout is not a substitute for the correct readiness condition.

Data leaks between jobs

Create a fresh context per independent identity, avoid shared mutable cookies, and close contexts after extraction.

Sessions consume capacity after errors

Put context and browser closure in finally; verify cleanup metrics and terminate orphaned jobs in the scheduler.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup

If the deliverable is a clean website image or PDF rather than extracted data, ScreenshotNeo is a simpler API option. It accepts consent banners before capture, removes more than 60 known consent platforms plus newsletter popups and chat widgets, and bills only clean shots: bot checks, CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed. Its MCP server exposes take_screenshot, get_page_info, and capture_pdf to Claude, Cursor, and other MCP clients.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

See the ScreenshotNeo API documentation for all 63 options, including full-page and selector capture, device presets, custom CSS and JavaScript, waits, blocking, headers, cookies, geolocation, resizing, caching, signed links, webhooks, bulk capture, and PDF controls. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.

FAQ

Can Browserless guarantee a target site’s availability?

No. It provides browser capacity and connection protocols; the target can still block, challenge, rate-limit, or change its page.

Should I share one context across all URLs?

Only when those URLs intentionally belong to the same session. Separate contexts are safer for independent identities and tenants.

Does a larger plan guarantee faster scraping?

No. More concurrency removes one capacity constraint, but page behavior, network latency, rendering cost, and target-site limits remain.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Frequently Asked Questions

What is the first scaling change to make?

Add a bounded queue and explicit concurrency limit before adding more browser sessions.

Which Browserless protocol should I choose?

Use native Playwright connect for Playwright-protocol features and connectOverCDP for CDP and Browserless helper integrations; verify the current feature matrix.

Where should the Browserless token live?

In an environment variable or secret manager, never in source code, client bundles, or ordinary logs.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Read next

Recommended PC Tool
Recommended PC Tool
PC Slower Than It Used to Be?Free scan - under a minute
Outdated Drivers Are Slowing You DownFree scan - exact matches

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.