Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

Use a headless browser when the data you need appears only after JavaScript runs, a user interaction changes the page, or browser requests fetch the useful content. For permitted targets, Playwright is a practical starting point: it can launch bundled Chromium, Chromium’s headless shell, a newer Chromium headless mode, or an installed Chrome or Edge channel. Start with the default bundled browser, then validate a specific channel when browser compatibility matters. A browser is not a permission bypass: check the site’s terms, authentication requirements and applicable law separately.

What headless browser scraping actually does

A headless browser runs a normal browser engine without displaying its window. It loads HTML, executes JavaScript, applies CSS, maintains cookies and storage, follows redirects and can perform actions such as clicking, typing and scrolling. Your scraper reads the resulting DOM or observes the requests the page makes.

This differs from an HTTP client that downloads a response and parses it. A simple request is usually faster and easier when the server response already contains the data. Choose browser automation only when browser behavior is part of the data path.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

When a browser is justified

  • The initial HTML is an app shell and records arrive through JavaScript.
  • Content appears after a click, login flow, pagination action or client-side route change.
  • Lazy-loaded images or sections require scrolling or waiting.
  • You need to reproduce a permitted user journey and capture the rendered result.
  • Network inspection is needed to diagnose whether data arrives through XHR or fetch.

When it is unnecessary

If a documented, authorized endpoint returns the needed fields directly, use an HTTP client and parser instead. It consumes fewer resources and has fewer timing failures. Do not infer that an endpoint observed in browser traffic is an authorized or stable public API; confirm permission and terms independently.

Install Playwright and choose a runtime

  1. Install the library in a new project: npm init -y followed by npm install playwright.
  2. Install Playwright’s bundled browsers with npx playwright install chromium.
  3. Save the script below as scrape.js and run it with node scrape.js.

Playwright uses open-source Chromium builds by default and also ships a separate Chromium headless shell. Its documentation describes an opt-in newer headless mode through the chromium channel and warns that the shell and newer mode can behave differently. Installed branded Chrome and Edge channels are supported but are not installed by Playwright. Select a runtime based on the compatibility you need, and test the target rather than assuming modes are equivalent.

Runtime choice Use it when Important qualification
Bundled Chromium You want a reproducible default for development and CI. Validate against the target’s browser-sensitive behavior.
Chromium headless shell Your deployment favors the dedicated headless binary. Playwright documents behavior differences from newer headless mode.
New headless mode via channel: 'chromium' High-fidelity web-app or extension testing requires the newer implementation. It is opt-in; verify behavior on your pages.
Installed Chrome or Edge channel The target must match a branded browser version or enterprise setup. The branded browser must already be installed.

A complete Playwright scraping example

This example visits a permitted page, waits for a meaningful selector, extracts links and writes a small JSON file. Replace the URL and selector with values appropriate to your target.

const { chromium } = require('playwright');
const fs = require('node:fs/promises');

(async () => {
  const browser = await chromium.launch({ headless: true });
  const page = await browser.newPage({
    viewport: { width: 1440, height: 900 },
    userAgent: 'ExampleResearchBot/1.0 (contact: [email protected])'
  });

  page.on('console', message => console.log('[console]', message.type(), message.text()));
  page.on('pageerror', error => console.error('[page error]', error.message));

  await page.goto('https://example.com/catalog', {
    waitUntil: 'domcontentloaded',
    timeout: 30000
  });
  await page.locator('[data-testid="product-list"]').waitFor({ timeout: 15000 });

  const records = await page.locator('[data-testid="product"]').evaluateAll(nodes =>
    nodes.map(node => ({
      name: node.querySelector('.name')?.textContent?.trim() ?? null,
      price: node.querySelector('.price')?.textContent?.trim() ?? null,
      url: node.querySelector('a')?.href ?? null
    }))
  );

  await fs.writeFile('records.json', JSON.stringify(records, null, 2));
  await browser.close();
})();

Waiting correctly

domcontentloaded means the document was parsed; it does not mean an application’s data is ready. Prefer a specific readiness condition such as locator(...).waitFor(). A short, justified delay can handle an animation, but fixed sleeps are less reliable than waiting for a selector or a state change. For pages whose activity settles predictably, you can wait for network idle, but analytics, ads and live connections may prevent that state.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Interactions and lazy content

Use locators for actions, for example await page.getByRole('button', { name: 'Next' }).click(), then wait for the next page’s distinctive selector. Scroll only as far as needed for lazy loading and deduplicate records by a stable identifier. Close dialogs only when the site presents them as part of a permitted user flow; do not use automation to defeat access controls.

Inspect browser network activity

Playwright can monitor and modify HTTP and HTTPS traffic, including XHR and fetch requests. Logging requests helps answer “where did this data come from?” before you decide whether DOM extraction or another permitted integration is appropriate.

const { chromium } = require('playwright');

(async () => {
  const browser = await chromium.launch();
  const page = await browser.newPage();

  page.on('request', request => {
    const type = request.resourceType();
    if (type === 'xhr' || type === 'fetch') {
      console.log('REQUEST', request.method(), request.url());
    }
  });
  page.on('response', async response => {
    const request = response.request();
    if (request.resourceType() === 'xhr' || request.resourceType() === 'fetch') {
      console.log('RESPONSE', response.status(), response.url());
    }
  });

  await page.goto('https://example.com/app', { waitUntil: 'domcontentloaded' });
  await page.waitForTimeout(3000); // replace with a page-specific readiness check
  await browser.close();
})();

Request logs are diagnostic evidence, not authorization. An observed URL may require a session, carry personal data, change without notice or be disallowed by the site’s terms. Keep credentials and response bodies out of logs unless you have a documented reason and appropriate handling.

Browser options: useful controls, not permission

Headless and browser channels

The BrowserType launch option headless defaults to true. Set headless: false while debugging to see the page, then return to headless operation in automation. To test the newer Chromium headless implementation, use chromium.launch({ channel: 'chromium', headless: true }) after installing the required browser. A branded channel can be selected similarly when Chrome or Edge is installed.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Proxy configuration

const browser = await chromium.launch({
  headless: true,
  proxy: { server: 'http://proxy.example:8080' }
});

Playwright supports HTTP and SOCKS proxies. A proxy changes the network route; it does not grant authorization, override contractual restrictions or guarantee that a target will respond. Use only infrastructure you control or are allowed to use.

Robots.txt, permission and security are different

RFC 9309 defines robots.txt as requested crawler instructions and states: “These rules are not a form of access authorization.” Google likewise explains that robots.txt cannot enforce crawler behavior, that crawlers may interpret syntax differently and that a disallowed URL can still be indexed when other pages link to it. Robots.txt is therefore one signal in an access policy, not a security boundary.

  • Permission: Read the site’s terms, obtain approval where required and respect account or API agreements.
  • Technical controls: Passwords, authorization checks and private networks protect content; do not attempt to bypass them.
  • Crawler instructions: Honor robots.txt and published rate guidance as part of responsible operation, while recognizing their limits.
  • Data handling: Minimize collection, protect credentials and consider privacy obligations before storing personal information.

The cited standards and Google guidance do not determine the legal status of scraping in every jurisdiction. When the activity is consequential, obtain qualified legal advice.

Make a scraper reliable and efficient

Control concurrency

Each browser process is comparatively heavy. Reuse one browser, create isolated contexts for separate sessions and cap concurrent pages. Start conservatively, measure CPU, memory and target response times, then increase concurrency only when the site and your infrastructure can support it.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use bounded waits and retries

Set navigation and selector timeouts. Retry transient navigation failures with exponential backoff and a small maximum attempt count; do not retry authentication failures or deliberate access denials. Record URL, browser channel, elapsed time, status and failure reason for each job.

Cache and deduplicate

Cache pages only when freshness requirements allow it. Store a canonical URL and stable record key, and design pagination jobs to resume after interruption. Avoid retaining full HTML or response bodies when extracted fields are sufficient.

Keep an audit trail

Record the policy basis for the job, the account used, timestamps and the selectors or schema version. This makes changes and complaints diagnosable without collecting unnecessary content.

Troubleshooting common failures

Symptom Likely cause Fix
Empty list after navigation Data is rendered later or a selector changed. Log XHR/fetch traffic, wait for a stable application selector and verify the locator in headed mode.
Timeout on networkidle Analytics, streaming or long polling keeps connections open. Wait for the specific element or response that means the page is ready.
Works headed, fails headless Different browser mode, timing or target behavior. Compare bundled, shell, new headless and required branded channels; test each explicitly.
403, CAPTCHA or login page The target requires authorization or rejects the request. Stop, obtain permission or use the site’s supported API. A proxy is not a permission solution.
Browser executable missing Playwright browsers were not installed in the environment. Run npx playwright install chromium during setup or deployment.
Memory grows over time Pages or contexts are not closed, or concurrency is too high. Close pages, reuse the browser, cap workers and monitor process memory.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server when your goal is a clean visual capture rather than extracting structured records. It accepts consent banners before capture and removes more than 60 known consent platforms, newsletter popups and chat widgets. Bot checks, blank pages, timeouts, failed loads and cache hits are not billed; response headers identify the page verdict and billing status. Its MCP server exposes take_screenshot, get_page_info and capture_pdf for Claude, Cursor and other MCP clients.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

One GET request returns PNG, JPEG, WebP or PDF. The API supports full-page and element captures, device presets, custom viewports, dark mode, retina scale, PDF paper and page controls, custom CSS or JavaScript, clicks, waits, blocking rules, headers, cookies, user agents, authorization, timezone, geolocation, transparent backgrounds, resizing, TTL caching, signed links, asynchronous webhooks, bulk capture of up to 100 URLs per call, usage reporting and an OpenAPI specification. Parameters used by other screenshot APIs also work, easing migration.

See the ScreenshotNeo documentation for current options. Example cURL:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

The Free plan includes 1,000 screenshots per month with no card. Paid plans start at $5 for 3,000 shots; every feature is on every plan. Create a free ScreenshotNeo account.

FAQ

Is headless scraping invisible?

No. Sites can observe requests, browser characteristics, authentication and usage patterns. Headless means no visible window, not anonymity.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Should I parse HTML or use browser automation?

Parse direct responses when they contain the needed data; use a browser when rendering or interaction is essential.

Can robots.txt make a private page secure?

No. Use authentication and technical access controls for private content.

Is Playwright the only headless browser option?

No. This guide uses Playwright because its documented Chromium, channel and network-monitoring controls fit the examples; other tools may also suit your language and target.

Frequently Asked Questions

How do I know whether a page needs JavaScript?

Fetch the HTML once and compare it with the content visible after the page loads. If the records are absent from the response and appear after browser execution, investigate with a permitted browser run.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Can I use a proxy to avoid a block?

A proxy changes routing but does not provide permission or justify bypassing a restriction. Resolve the access issue with the site owner or an approved integration.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.