Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

Playwright is useful for Python scraping when the data appears only after browser rendering or interaction. Install the Python package and browser binaries, open a page, locate content with resilient locators, wait for a meaningful condition, validate what you extracted, and save structured results. For static HTML, an HTTP client and parser are usually simpler and faster; use Playwright when you genuinely need a browser.

What Playwright does in a scraping workflow

Playwright automates real browser engines. Its Python package supports synchronous and asynchronous APIs, and the official installation flow provides Chromium, Firefox and WebKit binaries. A BrowserContext is an isolated browser session; a Page represents a tab or popup inside that context. You use the page to navigate, interact with controls and read the rendered DOM.

That makes Playwright a good fit for pages that load records with JavaScript, require a click before showing data, use client-side pagination, or render different content after selecting a location or account state. It is unnecessary overhead for a page whose complete data is already in the initial HTML response.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Check the target site’s terms, access requirements and applicable rules before collecting data. This tutorial does not establish permission for any particular site, nor does it replace advice specific to your jurisdiction or use case.

Install Playwright for Python

  1. Create and activate a virtual environment if the project does not already have one.
  2. Install the Python package:
python -m pip install playwright

Installing the package does not by itself install browser executables. Download them with:

playwright install

The command installs the browsers supported by the Playwright Python distribution, including Chromium, Firefox and WebKit. In a deployment image you can install only the engine you intend to launch, but keeping the default installation is the least surprising setup while developing.

A first synchronous scraper

The synchronous API is easy to follow for a sequential script. This example opens a page, reads its title and a heading, then closes resources even if an exception occurs.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
from playwright.sync_api import sync_playwright

URL = "https://example.com"

with sync_playwright() as p:
    browser = p.chromium.launch(headless=True)
    context = browser.new_context()
    page = context.new_page()

    response = page.goto(URL, wait_until="domcontentloaded", timeout=30_000)
    if response is None:
        raise RuntimeError("The navigation did not produce a response")

    print("HTTP status:", response.status)
    print("Title:", page.title())
    print("Heading:", page.locator("h1").inner_text())

    context.close()
    browser.close()

page.goto() navigates the tab and returns a response when one is available. A page can still be rendered from a cache or a non-document navigation, so treat a missing response as a condition to handle rather than assuming it is always present. Replace the example URL and selectors only for a site you are allowed to access.

Choose locators that survive markup changes

Locators are the central piece of Playwright’s auto-waiting and retry behavior. Prefer a locator based on what a user sees or on an explicit contract with the application: role, label, visible text, placeholder, alternative text, title or test ID. Scope it to the record or component you want before reading fields.

# A user-facing role and accessible name
submit = page.get_by_role("button", name="Search")
submit.click()

# A label associated with an input
query = page.get_by_label("Search products")
query.fill("keyboard")

# A stable application contract
cards = page.get_by_test_id("product-card")

for card in cards.all():
    name = card.get_by_role("heading").inner_text()
    price = card.get_by_test_id("price").inner_text()
    print({"name": name, "price": price})

When a page contains repeated records, first locate the record container and then locate its fields. This prevents a price or title elsewhere on the page from being paired with the wrong item. CSS selectors and XPath remain available for cases where the page exposes no better contract, but long positional selectors such as “the third div inside the second list” are fragile.

For attributes, use methods such as get_attribute("href"). For many matching elements, count() lets you check whether the page contains the expected number before iterating.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Wait for the data, not an arbitrary delay

Playwright auto-waits for many actions and locator operations. Use that behavior rather than adding a fixed sleep whenever a page uses JavaScript. Wait for the content your scraper actually needs:

results = page.get_by_test_id("search-result")
results.first.wait_for(state="visible", timeout=20_000)

if results.count() == 0:
    raise RuntimeError("The page reported no results")

rows = [item.inner_text() for item in results.all()]

A locator wait proves only the condition you stated. The first result being visible does not prove that every later result has finished loading. If the page exposes a “loaded” marker, a result count, or a next-page control, wait for that observable signal and validate the final count.

The Page API discourages networkidle as a generic readiness test and discourages fixed timeout waits in production. Network activity can continue because of analytics, advertising or long-lived connections even after the records you need are ready. A timeout is a diagnostic signal: inspect the selector, URL, response, consent dialog and page state before increasing it.

Complete example: search, extract and save JSON

This pattern searches a site, waits for its result cards, extracts a URL and text, rejects incomplete records, and writes UTF-8 JSON. The selectors are intentionally explicit placeholders; adapt them to the permitted target site’s markup.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
import json
from playwright.sync_api import sync_playwright, TimeoutError as PlaywrightTimeoutError

START_URL = "https://example.com/catalog"
OUTPUT = "products.json"

with sync_playwright() as p:
    browser = p.chromium.launch(headless=True)
    context = browser.new_context()
    page = context.new_page()
    try:
        page.goto(START_URL, wait_until="domcontentloaded", timeout=30_000)
        page.get_by_label("Search products").fill("keyboard")
        page.get_by_role("button", name="Search").click()

        cards = page.get_by_test_id("product-card")
        cards.first.wait_for(state="visible", timeout=20_000)

        records = []
        for card in cards.all():
            title = card.get_by_role("heading").inner_text().strip()
            link = card.get_by_role("link").get_attribute("href")
            price = card.get_by_test_id("price").inner_text().strip()
            if not title or not link:
                continue
            records.append({"title": title, "url": link, "price": price})

        if not records:
            raise RuntimeError("No complete records were extracted")
        with open(OUTPUT, "w", encoding="utf-8") as fh:
            json.dump(records, fh, ensure_ascii=False, indent=2)
    except PlaywrightTimeoutError as exc:
        raise RuntimeError("Expected content did not appear; inspect the selector and page state") from exc
    finally:
        context.close()
        browser.close()

Validate duplicates, missing fields and unexpected record counts before publishing or loading data downstream. Extraction code should fail loudly when a redesign changes the page instead of silently producing an empty file.

Async Python for concurrent workflows

Use the async API when the surrounding application already runs asyncio, or when you need to coordinate multiple independent pages. Do not mix sync calls into an active event loop.

import asyncio
from playwright.async_api import async_playwright

async def scrape(url: str):
    async with async_playwright() as p:
        browser = await p.chromium.launch(headless=True)
        page = await browser.new_page()
        try:
            await page.goto(url, wait_until="domcontentloaded", timeout=30_000)
            heading = page.locator("h1")
            await heading.wait_for(state="visible", timeout=15_000)
            return {"title": await page.title(), "heading": await heading.inner_text()}
        finally:
            await browser.close()

print(asyncio.run(scrape("https://example.com")))

On Windows, Playwright’s driver subprocess requires the ProactorEventLoop rather than SelectorEventLoop. Playwright’s API is not thread-safe; in a multi-threaded application, create a separate Playwright instance for each thread.

Handling interaction, pagination and dynamic records

Consent and modal dialogs

A consent banner can cover a button or change what is visible. Locate its accept or reject control by role or label, click it when appropriate for your use, then wait for the banner to disappear. Treat an unexpected modal as a distinct failure case rather than repeatedly retrying the obscured click.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Pagination

Extract one page, identify whether a next control is enabled, click it, and wait for a condition that changes—such as a new first-record URL or a page-number label. Keep a set of visited page URLs or record keys to prevent loops. Stop at a deliberate maximum and record the stopping reason.

Infinite scroll

Scroll only when the page’s loading behavior requires it. After each scroll, wait for the record count to increase or for a loading indicator to disappear. If the count does not change after a bounded number of attempts, stop and log that the end or a failure was reached.

Downloads and popups

Use Playwright’s event context when an action opens a new page or starts a download. Capture the resulting object, check its suggested filename or URL, and close the originating context in a finally block.

Browser engine and API choices

Choice Use it when Trade-off
Sync Python A straightforward, sequential script is the goal. Blocking control flow; not suitable inside an already-running async loop.
Async Python Your application uses asyncio or coordinates concurrent tasks. Every browser operation must be awaited and event-loop configuration matters.
Chromium The target environment is Chromium-based or you need to match it. It is not a universal proxy for Firefox or WebKit behavior.
Firefox or WebKit You must automate or verify that specific engine. Markup, timing or browser behavior can differ; no universal fastest engine is established here.

Performance, reliability and cost considerations

  • Reuse a browser and create isolated contexts for related jobs instead of launching a new browser process for every URL.
  • Close pages, contexts and browsers deterministically so memory and child processes do not accumulate.
  • Limit concurrency to what the target and your machine can handle; more tabs do not automatically mean more useful throughput.
  • Use bounded navigation and locator timeouts, log the URL and failed condition, and retry only failures that are plausibly transient.
  • Prefer a direct HTTP request and parser when the required data is in the response HTML. A browser consumes more CPU, memory and startup time.
  • Cache results where your use permits it, and avoid requesting the same page unnecessarily.

Playwright itself does not grant access to private data or bypass a site’s controls. Authentication, rate limits, bot checks and terms remain properties of the target service and your authorization.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Troubleshooting common failures

“Executable doesn’t exist”

The package is installed but browser binaries are missing. Run playwright install in the same environment used to run the script, and ensure the deployment image includes the downloaded browsers.

Timeout while waiting for a locator

Confirm that navigation reached the expected URL, inspect the rendered page, and verify the locator’s role, accessible name or test ID. Check for a consent dialog, login redirect or error page before changing the timeout.

Click is intercepted or the element is not visible

A modal, overlay or disabled state may be blocking the action. Locate and handle the overlay, wait for the control to become visible and enabled, and avoid forcing a click unless you have verified that a real user could not interact normally.

Results are empty although the page looks loaded

Your wait may target the shell rather than the records. Wait for the record locator or a documented loading marker, then check whether the data is inside an iframe or appears only after an interaction.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Windows async errors or thread failures

Use the ProactorEventLoop on Windows, and do not share a Playwright instance across threads. Create one instance per thread or keep the workflow on a single event loop.

The scraper breaks after a redesign

Prefer role, label and test-ID locators, scope them to record containers, and add validation for counts and required fields. Keep selector changes versioned with the target site’s markup assumptions.

Or skip the browser setup

If your requirement is simply a rendered screenshot or PDF rather than structured field extraction, ScreenshotNeo provides a website screenshot API and MCP server. It accepts consent banners before capture and removes more than 60 known consent platforms, newsletter popups and chat widgets; each step can be disabled. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads and cache hits are not billed, and response headers identify the page verdict and billing result.

One GET request returns PNG, JPEG, WebP or PDF. See the ScreenshotNeo API documentation for all options.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

ScreenshotNeo also offers an MCP server with take_screenshot, get_page_info and capture_pdf tools for Claude, Cursor and other MCP clients. Every feature is on every plan: 1,000 screenshots per month are free with no card; paid plans start at $5 for 3,000 shots, with yearly billing providing two months free.

Create a free ScreenshotNeo account to use the 1,000-shot monthly allowance without a card.

Practical checklist

  • Confirm the target site’s policies and your authorization.
  • Install both playwright and its browser binaries.
  • Choose sync or async to match the surrounding application.
  • Use user-facing or contract-based locators scoped to each record.
  • Wait for the data condition, not a guessed sleep or generic network quietness.
  • Validate counts, required fields and duplicates before saving.
  • Bound retries, concurrency and pagination, and close all resources.

Frequently Asked Questions

Do I need Playwright for every Python scraper?

No. Use a normal HTTP client and parser when the required data is present in the initial HTML. Choose Playwright when rendering or interaction is part of the data path.

Which Playwright browser should I install?

Install the engine that matches the environment you must automate or verify. Chromium, Firefox and WebKit are all supported; the available material does not establish a universally best or fastest engine.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Why does a locator timeout even though the page loaded?

The selector may be wrong, the content may require another interaction, a consent or login layer may be blocking it, or the page may have navigated to an error or redirect URL. Inspect those conditions before increasing the timeout.

Can Playwright bypass a CAPTCHA or access private information?

No. Playwright is an automation library; access controls, authentication, bot checks, rate limits and site policies still apply.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.