Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

Start without a browser. Fetch the page with a normal HTTP client, inspect the HTML and embedded state, then identify the network request that supplies the data. Reproduce that request and parse its response when possible. Use Playwright or another headless browser only when the request is impractical to reproduce, interaction is required, or the rendered DOM itself is your output.

The reliable decision process

  1. Fetch the initial response. Request the URL without JavaScript and inspect the returned HTML.
  2. Look for data already delivered. Check normal elements, script tags containing JSON-like state, and links to data files.
  3. Inspect network requests. In browser developer tools, open Network, reload the page, filter to Fetch/XHR, and identify the request that returns the records.
  4. Replay the request. Match its method, URL, query or form body, and only the headers, cookies or tokens that are actually required.
  5. Render as a fallback. Choose a headless browser when reproducing the request is difficult, interaction changes the result, or the rendered DOM or screenshot is the thing you need.

This order matters: a page that looks empty until JavaScript runs may still receive complete data through a simple JSON request. Rendering the entire page in that case adds browser startup, transfer and parsing work without adding information.

First pass: fetch and inspect without rendering

Check the response body

Use an ordinary HTTP client and save the response before writing selectors. If the records are in HTML, select them directly. If they are inside a script element, extract the script text and parse the structured portion rather than trying to interpret the whole page as visible text.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
import requests
from bs4 import BeautifulSoup

url = "https://example.com/catalog"
r = requests.get(url, timeout=30, headers={"User-Agent": "MyResearchBot/1.0"})
r.raise_for_status()

soup = BeautifulSoup(r.text, "html.parser")
for card in soup.select("article.product"):
    name = card.select_one("h2")
    price = card.select_one(".price")
    print({
        "name": name.get_text(" ", strip=True) if name else None,
        "price": price.get_text(" ", strip=True) if price else None,
    })

Save a sample response and verify what it contains. A server-rendered page can be parsed with selectors; a page whose source contains only a shell may require the next step.

Extract embedded state carefully

Many applications place initial state in a script element so the front end can hydrate quickly. Locate the specific script, isolate its JSON text, and parse it with a JSON parser. Do not use unrestricted regular expressions to “parse JSON”; they fail on escaped quotes, nested objects and braces inside strings. Some sites wrap the data in an assignment, so remove only the known prefix and suffix before parsing.

Find the request that actually contains the data

Use browser developer tools

  1. Open the page in a desktop browser and press F12 (or choose Inspect).
  2. Select Network, enable recording, clear the log and reload.
  3. Filter by Fetch/XHR. Trigger the page action that reveals more records, such as a search, filter or “next page” button.
  4. Open likely requests and inspect the Response or Preview panel. Identify the response containing the records, not merely an analytics or configuration call.
  5. Record the method, full URL, query parameters, request body, relevant headers, cookies and pagination fields. The browser’s Copy as cURL command is a useful starting point.

Replay the smallest faithful request. A JSON response should be parsed as JSON; HTML or XML should be handled with selectors. Authentication, anti-forgery tokens, cursor values and an Origin or Referer header may be required, but copying every browser header makes a scraper fragile. Remove headers one at a time and keep only those proven necessary.

Example: reproduce a JSON endpoint

import requests

endpoint = "https://example.com/api/products"
params = {"q": "laptop", "page": 1, "limit": 50}
headers = {"Accept": "application/json", "User-Agent": "MyResearchBot/1.0"}

r = requests.get(endpoint, params=params, headers=headers, timeout=30)
r.raise_for_status()
payload = r.json()

for item in payload.get("items", []):
    print(item.get("name"), item.get("price"))

Adapt the field names to the endpoint you observed. Confirm pagination semantics: some APIs use a page number, others a cursor returned in the previous response. Keep the response URL and status code in logs so a later schema change is diagnosable.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

When a headless browser is the right tool

Use browser automation when the required request is difficult to reproduce, the workflow depends on clicks or typed input, or the desired value exists only after scripts mutate the DOM. Playwright provides navigation and page-event APIs; scrapy-playwright lets a Scrapy crawl hand selected requests to Playwright while retaining Scrapy’s scheduling and item pipeline.

Minimal Playwright example

from playwright.sync_api import sync_playwright

with sync_playwright() as p:
    browser = p.chromium.launch(headless=True)
    page = browser.new_page()
    page.goto("https://example.com/catalog", wait_until="networkidle")
    page.locator("button.load-more").click()
    page.wait_for_selector("article.product")
    rows = page.locator("article.product").all()
    for row in rows:
        print(row.inner_text())
    browser.close()

Replace a fixed sleep with a meaningful condition: a selector, a URL change, a response event or a count of loaded records. “Network idle” is useful but not universal; pages with analytics or streaming connections may never become idle.

Scrapy plus scrapy-playwright

This combination is useful when most URLs are ordinary Scrapy requests and only selected pages need a browser. Set the request metadata that enables Playwright, extract from the rendered response, and close pages when you open them manually. The integration serializes the rendered DOM into the response body. Consequently, a JSON document can appear inside a pre element rather than as a normal JSON response; inspect the actual body before calling a JSON parser.

Choose the approach by task

Approach Use it when Trade-off
HTTP request plus parser Data is in the initial HTML, embedded state or a reproducible endpoint Requires understanding the request and response format, but avoids browser overhead
Headless browser Interaction, browser-only output or hard-to-reproduce requests are essential Adds browser processes, automation complexity and rendered-page failure modes
Scrapy plus scrapy-playwright A crawl needs Scrapy’s workflow with browser handling on selected requests Requires integration settings and awareness that responses are rendered DOM

Operational details that prevent brittle scrapers

Pagination and state

  • Record whether pagination is page-based, offset-based or cursor-based.
  • Stop when the endpoint returns no items or a documented end cursor; do not assume a fixed page count.
  • Persist the last successful cursor so a restart does not begin at the first page.

Waiting and retries

Set explicit connect and read timeouts. Retry transient network failures with bounded backoff, but do not blindly retry validation errors, authentication failures or a stable 404. For browser jobs, capture the URL, console errors and a screenshot or HTML sample when a selector is missing.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Response validation

Check status, content type and a required field before accepting a response. A login page, bot challenge or maintenance page can return status 200 while containing none of the expected records.

Robots and access rules

Check the site’s robots.txt and terms before crawling. Scrapy includes robots.txt middleware and a ROBOTSTXT_OBEY setting; configure the user agent used for robots matching deliberately. These controls are technical crawl guidance, not a legal determination. Rate-limit requests, identify your client where appropriate and avoid collecting data you do not need.

Common failures and fixes

“The HTML is empty”

Cause: the server sent an application shell and JavaScript later requested records. Fix: inspect Fetch/XHR traffic and reproduce the data request; render only if that request cannot be used directly.

“The endpoint works in the browser but returns 401 or 403”

Cause: a session cookie, short-lived token, anti-forgery value or required header is missing. Fix: compare the copied request with your client, obtain credentials through the site’s supported flow, and avoid hard-coding expired tokens.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

“The selector finds nothing in Playwright”

Cause: the page has not reached the relevant state, the content is inside an iframe, or the selector changed. Fix: wait for a state-specific selector, inspect frames, and verify the live DOM rather than the original source.

“JSON parsing fails after using scrapy-playwright”

Cause: the integration returned serialized rendered HTML, with JSON displayed in a pre element. Fix: inspect the response body and extract the displayed text, or call the underlying JSON endpoint directly.

“Requests intermittently disappear”

Cause: the target may be overloaded, buggy or blocking traffic. Fix: slow the crawl, log status and timing, retry transient errors with limits, and verify whether the same URL succeeds manually.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup

ScreenshotNeo provides a website screenshot API and MCP server when your deliverable is a clean image or PDF rather than extracted records. One GET request returns PNG, JPEG, WebP or PDF output. It accepts cookie and consent banners before capture and removes more than 60 known consent platforms, newsletter popups and chat widgets; each step can be disabled. Bot checks, blank pages, timeouts, failed loads and cache hits are not billed, and response headers identify the page verdict and billing status.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

For a screenshot, use the documented endpoint and options at ScreenshotNeo’s documentation:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

It also supports full-page and element capture, device and retina settings, dark mode, custom CSS and JavaScript, waits, request blocking, cookies and headers, geolocation, PDFs, caching with a chosen TTL, signed links, asynchronous webhooks, bulk capture of up to 100 URLs per call, a usage API and an MCP server with take_screenshot, get_page_info and capture_pdf for AI clients. The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000. Sign up free.

Frequently Asked Questions

Do I need Playwright for every JavaScript site?

No. First determine whether the page’s data is present in the initial response, embedded state or a reproducible network request. Playwright is a fallback for interaction, browser-only output or requests that are impractical to replay.

How can I tell whether I found the correct API request?

Change a visible filter or page, observe which request changes, then verify that its response contains the records and the expected pagination fields. A request that only returns configuration or analytics is not the extraction source.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Is scraping a site allowed?

Technical methods do not answer that question. Check the target’s robots.txt, terms and applicable rules, identify your crawler where appropriate, rate-limit it and collect only what you are permitted to use.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.