Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Playwright lets you drive Chromium, Firefox, and WebKit with the same browser, context, page, locator, and event APIs. A reliable scraper normally follows this sequence: launch a browser, create an isolated context, open a page, wait for a condition that proves the page is ready, locate data with user-facing locators, validate the result, and close everything in a finally block. The examples below use the standalone Playwright library in JavaScript rather than Playwright Test fixtures. Check the examples against the Playwright version installed in your project; the official documentation does not expose one stable version number for every page.

Only collect information you are allowed to access. Playwright does not grant permission to scrape a site, bypass authentication, defeat a CAPTCHA, or ignore a site’s terms and applicable law.

Install Playwright and run a first page

Create a project, install the library, and install at least one browser engine:

mkdir pw-scraper
cd pw-scraper
npm init -y
npm install playwright
npx playwright install chromium

This standalone script navigates to a benign example page, reads its title, and saves a screenshot. The browser is always closed, even when navigation or extraction fails.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
const { chromium } = require('playwright');

(async () => {
  const browser = await chromium.launch();
  try {
    const context = await browser.newContext();
    const page = await context.newPage();
    await page.goto('https://example.com', { waitUntil: 'domcontentloaded' });
    console.log(await page.title());
    await page.screenshot({ path: 'example.png', fullPage: true });
    await context.close();
  } finally {
    await browser.close();
  }
})();

The Page API documents the browser-to-context-to-page workflow and the basic page.screenshot() call: Playwright Page API.

Choose locators that survive page changes

Locators are the central piece of Playwright’s auto-waiting and retry-ability, according to the official locator guide. Prefer a locator that describes what a user sees or what the site explicitly promises, rather than a long CSS chain tied to today’s DOM.

Role and accessible-name locators

const heading = page.getByRole('heading', { name: 'Latest articles' });
await heading.waitFor();

const articleCards = page.getByRole('article');
const texts = await articleCards.evaluateAll(cards =>
  cards.map(card => card.textContent?.trim() ?? '')
);
console.log(texts);

Useful built-in choices include getByRole, getByText, getByLabel, getByPlaceholder, getByAltText, getByTitle, and getByTestId. For buttons, links, headings, and form controls, a role plus accessible name is usually clearer than a class name.

Filter a repeated card before acting

When every product card has a similar button, first narrow the parent locator, then find the child control:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
const product = page.getByRole('listitem').filter({ hasText: 'Coffee grinder' });
await product.getByRole('button', { name: 'Add to cart' }).click();

This expresses the intended relationship and avoids clicking the first matching button elsewhere on the page. CSS and XPath remain available when a semantic locator or explicit test contract is unsuitable, but selectors such as div:nth-child(3) > span.item are coupled to structure that can change.

Extract only the fields you need

const rows = page.getByRole('row');
const records = await rows.evaluateAll(items => items.map(row => {
  const cells = [...row.querySelectorAll('th,td')]
    .map(cell => cell.textContent?.trim() ?? '');
  return { name: cells[0] ?? '', value: cells[1] ?? '' };
}));

for (const record of records) {
  if (!record.name || !record.value) {
    throw new Error(`Invalid row: ${JSON.stringify(record)}`);
  }
}
console.log(JSON.stringify(records, null, 2));

evaluateAll() runs a DOM operation over the elements currently matched by a locator. Keep the mapping focused, normalize whitespace, and validate required fields after extraction. The output depends on the target page’s markup; no selector is universal.

Wait for the page’s real readiness condition

Navigation completion is not the same as data readiness. A single-page application may render its list after the initial document loads. Wait for a meaningful condition instead of adding an arbitrary sleep.

Wait for a heading or list

await page.goto('https://example.com/catalog', { waitUntil: 'domcontentloaded' });
await page.getByRole('heading', { name: 'Catalog' }).waitFor();
const items = page.getByRole('listitem');
await items.first().waitFor();
const names = await items.allTextContents();

Wait for a selector or application state

await page.waitForSelector('[data-ready="true"]');
await page.waitForFunction(() => window.catalogLoaded === true);

Use a condition tied to the page’s contract: a heading, a row, a “loaded” marker, or a known application state. If the page’s list changes while you collect it, do not call locator.all() immediately and assume it waits. The Locator API documentation warns that all() returns current matches without waiting, so changing lists can produce unpredictable results. Wait for the list’s readiness condition first, then read it.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Navigation timeout and failure handling

page.setDefaultNavigationTimeout(45_000);
page.setDefaultTimeout(15_000);

try {
  await page.goto(targetUrl, { waitUntil: 'domcontentloaded' });
} catch (error) {
  console.error(`Navigation failed for ${targetUrl}:`, error.message);
  throw error;
}

A timeout does not prove that a site is down: it can mean slow resources, a redirect loop, a consent wall, authentication, or a bot check. Record the URL and error, then decide whether to retry under the site’s permitted access rules.

Scrape a paginated list without losing your place

For a conventional “Next” link, collect one page at a time and stop when the control is disabled or absent. Scope each extraction to the current page and set a maximum page count as a safety guard.

const results = [];
const maxPages = 20;

for (let pageNumber = 1; pageNumber <= maxPages; pageNumber++) {
  await page.getByRole('article').first().waitFor();
  const pageRecords = await page.getByRole('article').evaluateAll(cards =>
    cards.map(card => ({
      title: card.querySelector('h2,h3')?.textContent?.trim() ?? '',
      text: card.textContent?.trim() ?? ''
    }))
  );
  results.push(...pageRecords);

  const next = page.getByRole('link', { name: /next/i });
  if (await next.count() === 0 || await next.isDisabled().catch(() => false)) break;
  await Promise.all([
    page.waitForLoadState('domcontentloaded'),
    next.click()
  ]);
}

console.log(`Collected ${results.length} records`);

Some sites use a “Load more” button or infinite scrolling instead. Click the control and wait for the count to increase, or scroll only as far as the page’s documented behavior requires. Deduplicate records by a stable URL or identifier before writing output.

Keep users and sessions isolated with BrowserContexts

A BrowserContext is an isolated, incognito-like profile. Cookies, local storage, permissions, and other session state are separated, and contexts are designed to be fast and inexpensive to create. The browser-context documentation also shows how separate contexts model multiple users in one workflow.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
const browser = await chromium.launch();
try {
  const guest = await browser.newContext();
  const member = await browser.newContext();

  const guestPage = await guest.newPage();
  const memberPage = await member.newPage();

  await guestPage.goto('https://example.com');
  await memberPage.goto('https://example.com/account');

  // Cookies or local storage created in one context are not shared with the other.
  await guest.close();
  await member.close();
} finally {
  await browser.close();
}

Use one context when a workflow intentionally shares a login. Create separate contexts when testing guest/member behavior, processing independent accounts, or preventing one job’s cookies from affecting another. Isolation is a way to organize permitted sessions; it is not a way to bypass access controls.

Capture full-page, element, and in-memory screenshots

Save a full-page image

await page.screenshot({ path: 'catalog.png', fullPage: true });

Capture one element

const card = page.getByRole('article').first();
await card.screenshot({ path: 'first-card.png' });

Keep the image in memory

const buffer = await page.screenshot({ type: 'png' });
require('node:fs').writeFileSync('catalog-buffer.png', buffer);

The stable Page API covers these basic workflows. Playwright’s next-version screenshots guide is forward-looking; verify any option described there against the installed release before depending on it. A full-page image can be tall and expensive to process, while an element shot is smaller but requires a stable locator and a rendered element.

Wait for downloads and save them before closing the context

Start waiting for the download before clicking. The page emits its download event when the download starts; then save the completed file with saveAs().

const downloadPromise = page.waitForEvent('download');
await page.getByText('Download file').click();
const download = await downloadPromise;
const fileName = download.suggestedFilename();
await download.saveAs(`/absolute/path/output/${fileName}`);

Files associated with a browser context are deleted when that context closes, so save the file before closing the context. In production, validate the suggested filename and resolve it beneath an intended output directory rather than allowing path separators supplied by a remote page. See the Download API. The next downloads page is also forward-looking; prefer the stable API page unless you have verified your version.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Combine navigation, extraction, and download in one script

This complete example visits a page, extracts visible article data, takes a screenshot, and saves a download if the page exposes the expected control.

const { chromium } = require('playwright');
const fs = require('node:fs/promises');

async function run(targetUrl) {
  const browser = await chromium.launch();
  const context = await browser.newContext({ viewport: { width: 1440, height: 900 } });
  try {
    const page = await context.newPage();
    await page.goto(targetUrl, { waitUntil: 'domcontentloaded', timeout: 45_000 });
    await page.getByRole('article').first().waitFor({ timeout: 15_000 });

    const articles = await page.getByRole('article').evaluateAll(nodes =>
      nodes.map(node => ({
        title: node.querySelector('h2,h3')?.textContent?.trim() ?? '',
        text: node.textContent?.replace(/s+/g, ' ').trim() ?? ''
      }))
    );
    if (articles.some(article => !article.title)) throw new Error('A record has no title');
    await fs.writeFile('articles.json', JSON.stringify(articles, null, 2));
    await page.screenshot({ path: 'page.png', fullPage: true });

    const downloadLink = page.getByRole('link', { name: /download/i });
    if (await downloadLink.count()) {
      const pending = page.waitForEvent('download');
      await downloadLink.click();
      const download = await pending;
      const safeName = download.suggestedFilename().replace(/[^a-zA-Z0-9._-]/g, '_');
      await download.saveAs(`downloads/${safeName}`);
    }
  } finally {
    await context.close();
    await browser.close();
  }
}

run('https://example.com').catch(error => {
  console.error(error);
  process.exitCode = 1;
});

Common failures and practical fixes

“Locator resolved to multiple elements”

Your locator is ambiguous. Add an accessible name, use filter({ hasText: ... }), or select a deliberately scoped parent before the child control. Do not blindly add first() unless the first match is genuinely the required one.

“Timeout exceeded” while waiting

Check the URL, redirects, authentication state, consent UI, and whether the expected role or accessible name exists. Inspect the page with a headed browser or save its HTML for diagnosis. Replace a guessed sleep with a condition that represents readiness, and increase the timeout only when the slower operation is expected.

An empty or partial collection

The list may still be rendering, may be virtualized, or may be inside an iframe. Wait for a specific row or application marker before reading it. If it changes while being read, avoid immediate all(); collect after the list stabilizes and validate the record count and required fields.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Clicks do not start the download

Start waitForEvent('download') before the click, ensure the locator targets the actual download control, and check whether a new tab or popup is opened instead. Save the download before closing its context.

Data differs between runs

Record the URL, timestamp, viewport, locale, and relevant response or page errors. Dynamic content, personalization, rate limits, and changing markup can all affect output. Keep concurrency within the target’s permitted limits and retry only errors that are safe to retry.

Performance, reliability, and operating choices

  • Reuse a browser process carefully: launching one browser and creating contexts for independent jobs avoids accidental cookie sharing while keeping lifecycle management explicit.
  • Choose the smallest result: extract fields rather than entire HTML documents, and capture an element instead of a full page when that is all you need.
  • Make readiness observable: wait for a role, selector, count, or application state that proves the data exists; do not treat a fixed delay as a guarantee.
  • Validate at the boundary: reject missing identifiers, malformed URLs, unexpected empty pages, and duplicate records before they reach downstream systems.
  • Preserve diagnostics: keep structured logs for navigation failures and selected screenshots or HTML snapshots when permitted. Do not store credentials or personal data unnecessarily.
  • Use bounded work: cap pages, records, retries, and output size so a changed site cannot create an unbounded job.

The reviewed Playwright documentation does not provide a universal speed or success-rate benchmark. The right browser, context, locator, and waiting strategy depends on the target page and your permitted workload.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup

For a server-side screenshot rather than a local browser workflow, ScreenshotNeo accepts one GET request and returns a PNG, JPEG, WebP, or PDF. Its cleanup steps accept the cookie or consent banner as a visitor and remove more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be turned off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and the response identifies the result with X-Page-Verdict and X-Billed headers. It also provides an MCP server for Claude, Cursor, and other MCP clients, with take_screenshot, get_page_info, and capture_pdf tools.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

See the ScreenshotNeo API documentation for authentication and options. This is a one-call WebP example:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

The equivalent Python and Node.js calls are:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

ScreenshotNeo includes full-page capture with lazy images loaded, CSS-selector element capture, dark mode, 12 device presets plus custom viewports, retina scale, PDF paper and page-range controls, custom CSS and JavaScript, pre-capture clicks, hidden selectors, waits for a selector, delay or network idle, request and resource blocking, headers, cookies, user agent, Authorization, timezone, geolocation, transparent backgrounds, resizing, configurable-TTL caching, signed image links, asynchronous jobs with signed webhooks, bulk capture of up to 100 URLs per call, a usage API, an OpenAPI specification, and compatibility with the parameter names used by other screenshot APIs. Every feature is on every plan: Free includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account to try the 1,000 monthly screenshots without a card.

Frequently asked questions

Should I use Playwright Test or the standalone library?

Use the standalone library when your program is a scraper, data pipeline, or one-off automation job. Use Playwright Test when you also want its test runner, fixtures, assertions, and reporting. The examples here intentionally use the library API.

Can Playwright scrape content rendered inside an iframe?

Yes, when you can legally access it: obtain a frame locator or frame reference and use locators within that frame. The page locator and the iframe’s document are separate scopes.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Why does a screenshot show a consent banner?

Playwright captures what its page rendered. Locate and handle a permitted consent flow before capture, or hide a known selector only when doing so accurately represents your intended result. A screenshot service such as ScreenshotNeo can perform its documented consent and widget cleanup before capture.

How should I schedule large scraping jobs?

Queue bounded jobs, isolate sessions with contexts, limit concurrency for the target, persist checkpoints, and make retries idempotent. Add monitoring for empty results and markup changes instead of assuming every run succeeded.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.