Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use a real browser to discover and synchronize an SPA, then call its data endpoint directly when that endpoint is stable and permitted. A reliable Python scraper normally follows this sequence: launch Playwright, wait for application state rather than initial HTML, inspect the XHR or fetch request that carries the data, and move the repeatable extraction to HTTP requests or Scrapy. Keep browser automation for login, interaction, client-side computation, and endpoint discovery.

The practical architecture: browser first, API second

A single-page application (SPA) often returns a small HTML shell. JavaScript then makes XHR or fetch calls, renders cards, and changes the URL without a full reload. Parsing the response from requests.get() alone can therefore produce an empty page.

Start with a browser when you do not yet know how the application works. Playwright’s Python package can drive Chromium, Firefox, or WebKit, and browsers run headlessly by default. Once you identify a stable, permitted request containing the records you need, reproduce that request with ordinary Python HTTP code. This hybrid design avoids paying the startup and rendering cost for every page while retaining a browser for the parts that genuinely require one.

Install Python and Playwright

  1. Create and activate a virtual environment for the scraper.
  2. Install the Python package: pip install playwright.
  3. Download the browser binaries: playwright install. You can install only a required engine, such as Chromium, if your deployment image must be smaller.
  4. Confirm that your target permits the collection. Read its robots.txt, terms of service, authentication rules, and published rate limits before sending requests.

Use a Playwright browser context to make cookies, locale, proxy, permissions, and JavaScript behavior explicit. Do not rely on a developer’s personal browser profile in production.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Reconnaissance: discover what “loaded” means

Open the target in a headed browser while developing so you can see consent dialogs, login redirects, infinite scrolling, and “Load more” controls. Record the interaction that reveals the data and the DOM element that proves the operation completed.

Useful readiness signals

  • A selector for the first real record, such as [data-testid="product"].
  • A URL transition after a search or filter.
  • A particular JSON response from an XHR or fetch request.
  • A count, table row, or status element whose text changes from “Loading” to a completed state.

A fixed sleep can work accidentally on a fast run and fail on a slow one. Treat readiness as application state. Register a response or request listener before clicking the control that triggers it, so the event cannot be missed.

A complete Playwright scraper

The following asynchronous script waits for rendered records, extracts their text and links, and reports browser failures separately from an HTTP error returned by the site. Replace the URL and selectors with values found during reconnaissance.

import asyncio
import json
from playwright.async_api import async_playwright, TimeoutError as PlaywrightTimeoutError

URL = "https://example.com/catalog"

async def main():
    async with async_playwright() as pw:
        browser = await pw.chromium.launch(headless=True)
        context = await browser.new_context(locale="en-US")
        page = await context.new_page()
        page.set_default_timeout(15_000)

        try:
            response = await page.goto(URL, wait_until="domcontentloaded", timeout=60_000)
            if response is not None and not response.ok:
                raise RuntimeError(f"navigation returned HTTP {response.status}")

            cards = page.locator('[data-testid="product"]')
            await cards.first.wait_for(state="visible")
            records = await cards.evaluate_all("""
                nodes => nodes.map(node => ({
                    title: node.querySelector('[data-testid="title"]')?.textContent?.trim() || null,
                    price: node.querySelector('[data-testid="price"]')?.textContent?.trim() || null,
                    url: node.querySelector('a')?.href || null
                }))
            """)
            print(json.dumps(records, ensure_ascii=False, indent=2))
        except PlaywrightTimeoutError as exc:
            raise RuntimeError("content did not reach the expected state before the timeout") from exc
        finally:
            await context.close()
            await browser.close()

if __name__ == "__main__":
    asyncio.run(main())

page.goto() completing does not prove that the application succeeded. A 404 or 500 is still a completed HTTP response, so inspect its status. The same rule applies to later API responses.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Wait for XHR and fetch instead of guessing with sleeps

Playwright can monitor HTTP and HTTPS traffic, including requests made by XHR and fetch. Capture the request and response metadata while performing the action that reveals the data.

from playwright.async_api import async_playwright

async def inspect_request():
    async with async_playwright() as pw:
        browser = await pw.chromium.launch(headless=True)
        page = await browser.new_page()
        captured = []

        def on_response(response):
            request = response.request
            if request.resource_type in {"xhr", "fetch"}:
                captured.append({
                    "url": response.url,
                    "method": request.method,
                    "status": response.status,
                    "resource_type": request.resource_type,
                })

        page.on("response", on_response)
        await page.goto("https://example.com/catalog", wait_until="domcontentloaded")
        await page.get_by_role("button", name="Load more").click()
        await page.locator('[data-testid="product"]').last.wait_for(state="visible")
        print(captured)
        await browser.close()

For a known endpoint, synchronize directly with the action:

async with page.expect_response(
    lambda r: "/api/products" in r.url and r.request.method == "GET"
) as response_info:
    await page.get_by_role("button", name="Load more").click()
api_response = await response_info.value
if not api_response.ok:
    raise RuntimeError(f"API failed with HTTP {api_response.status}")
payload = await api_response.json()

Inspect the captured request in your browser’s developer tools or through Playwright. Preserve its method, query parameters, request body, relevant headers, cookies, and pagination cursor. A request may be a POST even when the page operation looks like a simple search.

Move stable data extraction to direct HTTP

Scrapy’s documentation recommends reproducing the additional request that contains the desired data when that request is appropriate. Direct HTTP is usually easier to retry, paginate, test, and run in parallel than a full browser.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
import requests

API = "https://example.com/api/products"
session = requests.Session()
session.headers.update({
    "Accept": "application/json",
    "User-Agent": "catalog-research/1.0"
})

all_items = []
for page_number in range(1, 11):
    response = session.get(
        API,
        params={"q": "laptop", "page": page_number, "per_page": 50},
        timeout=30,
    )
    response.raise_for_status()
    data = response.json()
    items = data.get("items", [])
    all_items.extend(items)
    if not data.get("next_page") or not items:
        break

print(f"collected {len(all_items)} records")

Do not copy browser-only headers blindly. Keep the headers required by the endpoint, remove transient telemetry where possible, and use the authentication method the site authorizes. Validate the JSON shape before writing records so a login page or an error object cannot silently become your dataset.

When the browser must remain in the loop

  • The endpoint requires a login flow, a short-lived token, or a cookie established by JavaScript.
  • Clicking, scrolling, or client-side computation changes which records are requested.
  • The application assembles data from several calls and no single stable endpoint contains the required fields.
  • The endpoint is intentionally unavailable to your permitted client, while the visible browser workflow is allowed.

In those cases, use Playwright to authenticate and interact, then either parse the rendered DOM or capture the response inside that authenticated context.

Pagination, scrolling, and stateful interactions

Cursor or page pagination

Prefer the server’s cursor or documented page token over incrementing a number until a request happens to return no rows. Persist the last successful cursor, stop when the API explicitly signals completion, and make a maximum-page guard part of the job.

Infinite scrolling

Scroll only when scrolling is the operation that requests more data. After each scroll, wait for either a new card count or the specific response that carries the next batch. Stop when the count stops increasing and no “load more” request is observed.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Filters and URL state

Set filters through the same controls a user would use, then capture the resulting URL and request parameters. If the URL changes, wait for that URL pattern before extracting results; a visible filter label alone may update before the data does.

Playwright, Selenium, or direct requests?

Approach JavaScript fidelity Network/API visibility Startup and parallelism Best fit
Playwright High; drives Chromium, Firefox, and WebKit Built-in request, response, routing, and waiting APIs Browser startup is costly; contexts allow controlled concurrency Modern SPAs, discovery, interaction, authenticated flows
Selenium High when paired with a real browser driver Available, but synchronization and network inspection generally require more driver-specific setup Browser and driver overhead; mature grid options for parallel runs Existing Selenium suites, organizations standardized on WebDriver
Direct requests or Scrapy None; does not execute page JavaScript Complete control once you know the data endpoint Low overhead and straightforward parallelism Stable, permitted JSON or HTML endpoints

Choose based on the page’s behavior, not brand preference. A direct request is preferable when the data-bearing endpoint is stable and allowed; browser automation is preferable when rendering or interaction is essential.

Reliability and performance controls

  • Timeouts: Set navigation and locator timeouts explicitly. Use a larger timeout for a known slow operation rather than a global, indefinite wait.
  • Retries: Retry idempotent GET requests with bounded backoff. Do not blindly replay a state-changing POST.
  • Status and schema checks: Record status codes, content type, and a small schema fingerprint. Alert when a field disappears or changes type.
  • Deterministic boundaries: Save the cursor or page token, deduplicate by a stable identifier, and cap the number of pages per run.
  • Resource use: Reuse one browser process and create isolated contexts; close contexts after each job. Block unnecessary images or third-party resources only when doing so does not alter the application’s data path.
  • Observability: Log URL, elapsed time, retry count, response status, and parser outcome without logging passwords, session cookies, or unnecessary personal data.

Common failures and fixes

The HTML contains no records

Cause: records arrive after JavaScript runs. Fix: wait for a record selector, inspect XHR/fetch traffic, and use the data endpoint if it is stable and permitted.

Timeout waiting for a selector

Cause: the selector is wrong, a consent or login screen blocked the app, or the request failed. Fix: inspect the post-JavaScript DOM, capture failed requests, handle the required dialog or authentication, and wait for a meaningful state rather than a guessed class name.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The response is 200 but contains an error object

Cause: an application-level error, expired session, or rate-limit response is encoded inside a successful HTTP response. Fix: validate required keys and error fields before parsing items; refresh the authorized session or slow the request rate.

Navigation “succeeds” on a 404 or 500

Cause: navigation completion is not the same as application success. Fix: inspect response.status and fail the job for unexpected status codes.

Direct requests return a login page

Cause: the browser established cookies or tokens that your HTTP client lacks. Fix: use the documented authentication flow, or keep the extraction inside an authorized Playwright context. Never bypass access controls.

Pagination duplicates or skips records

Cause: an unstable offset, changing sort order, or an incorrectly reused cursor. Fix: use the server’s cursor, request a deterministic sort, persist progress, and deduplicate by the site’s stable identifier.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Compliance and data minimization

Review robots.txt and the site’s terms before collection. Honor explicit rate limits and access restrictions, avoid bypassing authentication or technical controls, and collect only the personal data necessary for your stated purpose. Store credentials and cookies securely, set a retention period, and provide a removal path when your use case requires one. A technically successful scraper can still violate a site’s rules or privacy obligations.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup

If your goal is a clean image or PDF of a page rather than structured records, ScreenshotNeo provides a single website-screenshot API request. It accepts consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets. Only clean shots are billed: bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and the response identifies the result with X-Page-Verdict and X-Billed headers.

See the full parameter reference in the ScreenshotNeo documentation. This call returns a WebP image for the example URL:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

ScreenshotNeo also offers full-page captures with lazy images loaded, CSS-selector element captures, dark mode, device presets and custom viewports, retina scale, PDF paper and margin controls, custom CSS and JavaScript, click-before-capture actions, selector hiding, waits for selectors, delays or network idle, request and resource blocking, custom headers and cookies, user-agent and authorization values, timezone and geolocation, transparent backgrounds, resizing, configurable-TTL caching, signed image links, asynchronous jobs with signed webhooks, bulk capture for up to 100 URLs per call, a usage API, and an OpenAPI specification. Its parameter names are compatible with those used by other screenshot APIs, which can simplify migration. An MCP server exposes take_screenshot, get_page_info, and capture_pdf to Claude, Cursor, and other MCP clients.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Plan Included shots Price
Free 1,000 per month $0, no card
Starter 3,000 $5
Growth 15,000 $15
Pro 60,000 $39
Scale 250,000 $99
Business 1,000,000 $249

Every feature is included on every plan, and yearly billing provides two months free. Start with 1,000 free screenshots a month; no card is required.

FAQ

Can I scrape an SPA without executing JavaScript?

Sometimes. If the page’s data endpoint is public, stable, and permitted for your use, call it directly. Otherwise, the browser is needed to establish state or perform the interaction that produces the data.

How do I know whether an endpoint is safe to reproduce?

Check that it consistently returns the fields you need, uses an authorized authentication method, has predictable pagination, and is allowed by the site’s terms and access policy. Reconfirm these conditions when the application changes.

Should I run headed browsers in production?

Usually no. Develop headed so you can observe dialogs and state transitions; Playwright runs headlessly by default for unattended jobs. Switch to headed mode only when diagnosing a visual or environment-specific failure.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What should I archive for a scraper audit?

Keep the target URL, collection timestamp, policy review, request status, parser version, and a minimal record of failures. Do not archive session secrets or personal data that the project does not need.

Why capture both request and response events?

The request reveals method, parameters, and headers; the response reveals status, timing, and payload availability. Together they distinguish a missing trigger from an endpoint that returned an application error.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.