Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

If a value appears in a browser but not in requests.get(...).text, the browser is probably obtaining it after the initial HTML response. Find that data source first; render the page with Playwright only when the source cannot reasonably be reproduced or when you need browser interaction and the rendered DOM.

This guide shows how to diagnose the difference, extract embedded or network-delivered data, render JavaScript with Python, wait for application state instead of arbitrary delays, and recover from common failures.

Why dynamic content is missing from a Python response

An HTTP client receives the server response. It does not execute the JavaScript that a browser runs afterward. A page can therefore return a small HTML shell while scripts fetch products, comments, prices or search results from another endpoint.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A visible element is not proof that its value was present in the first response. The value may be:

  • Already in the original HTML but hidden by CSS.
  • Embedded in a <script> element as JSON or JavaScript data.
  • Returned by an XHR or fetch request after navigation.
  • Created only after an interaction, such as a click, scroll, login or form submission.

Scrapy’s current documentation recommends finding the data source and extracting it directly when possible. That normally produces more structured input with less browser overhead than rendering an entire page.

Diagnose the page before choosing a scraper

Compare source HTML with the browser DOM

  1. Fetch the URL with requests and save response.text.
  2. Use the browser’s “View source” and developer-tools Elements panel. “View source” represents the response; Elements represents the current DOM.
  3. Search both for a distinctive value, a CSS class, or the element’s label.
  4. Inspect script tags for JSON, serialized state, or configuration containing the data.
import requests

url = "https://example.com/catalog"
r = requests.get(url, timeout=30)
r.raise_for_status()
print("status:", r.status_code)
print("value in response:", "Example product" in r.text)
print(r.text[:500])

If the value is in the response, parse that response instead of launching a browser. If it is absent, open the Network panel, reload the page, and filter for Fetch/XHR requests. Identify the request whose response contains the required records.

Record the complete data request

When reproducing a request, compare its method, URL, query string, body, headers, cookies and form parameters. A copied URL without the request body or an authorization header may return an empty or different result. “Copy as cURL” in browser developer tools is a useful starting point; translate the relevant parts into Python and remove incidental headers one at a time.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Choose the least complex working approach

Approach Choose it when Main trade-off
Initial HTML or embedded data The values are in the response or a script payload Lowest overhead, but the response shape must remain parseable.
Reproduced data request Network inspection reveals a structured endpoint Usually less rendering and parsing work; request details and permitted access must be understood.
Playwright with Python JavaScript execution, interaction, or the rendered DOM is essential Highest browser fidelity, with greater runtime, memory and page-change sensitivity.
Scrapy plus browser integration A crawler needs Scrapy facilities as well as browser rendering Preserves more crawler components but adds integration and compatibility setup.

This is a qualitative decision, not a benchmark. Prefer the smallest method that reliably returns the fields you need.

Extract embedded JSON without rendering

Many applications place a JSON object in a script tag. Parse it as JSON when it is valid JSON; do not treat a regular expression as a general JavaScript parser. JavaScript object literals can contain single quotes, comments, trailing commas and expressions that require a JavaScript-aware parser.

import json
import requests
from bs4 import BeautifulSoup

r = requests.get("https://example.com", timeout=30)
r.raise_for_status()
soup = BeautifulSoup(r.text, "html.parser")

for script in soup.find_all("script"):
    text = script.string or script.get_text()
    if '"products"' not in text:
        continue
    try:
        state = json.loads(text)
    except json.JSONDecodeError:
        continue
    products = state.get("products", [])
    for product in products:
        print(product.get("name"))
    break

Use a stable script identifier or a documented state property when available. Validate that the expected keys exist and log a useful error when the site changes its serialization.

Reproduce the browser’s data request with Python

Suppose the Network panel shows a POST request returning JSON. Recreate only the material request properties:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
import requests

endpoint = "https://example.com/api/search"
payload = {"query": "laptops", "page": 1}
headers = {
    "Accept": "application/json",
    "Content-Type": "application/json",
    "User-Agent": "my-research-client/1.0",
}

r = requests.post(endpoint, json=payload, headers=headers, timeout=30)
r.raise_for_status()
data = r.json()
for row in data.get("results", []):
    print(row.get("title"))

If the browser sends form data, use data= rather than json=. If it sends query parameters, use params=. Preserve authentication and session cookies only when you are authorized to use them. Check pagination, rate limits and whether the endpoint is intended for your access.

Render JavaScript with Playwright Python

Install and launch

The example assumes a current Playwright for Python installation. Check the version-specific installation instructions for your environment before pinning this in production.

python -m pip install playwright
python -m playwright install chromium

Wait for the result, not an elapsed time

Playwright performs actionability checks before actions, and locators resolve against the current DOM. Use a locator for the result and an assertion or state condition. Fixed sleeps can pass on a fast run and fail on a slow one; Playwright’s documentation discourages timeout waits for production readiness checks. The networkidle state is also discouraged as a generic test of readiness because modern pages may keep connections open or perform work after it occurs.

from playwright.sync_api import sync_playwright, TimeoutError as PlaywrightTimeoutError

URL = "https://example.com/catalog"

with sync_playwright() as p:
    browser = p.chromium.launch(headless=True)
    page = browser.new_page()
    try:
        response = page.goto(URL, wait_until="domcontentloaded", timeout=60_000)
        if response is None:
            raise RuntimeError("Navigation produced no main response")
        if response.status >= 400:
            raise RuntimeError(f"HTTP status from main document: {response.status}")

        cards = page.locator("[data-testid='product-card']")
        cards.first.wait_for(state="visible", timeout=30_000)
        rows = cards.evaluate_all("els => els.map(e => ({name: e.querySelector('h2')?.textContent?.trim(), price: e.querySelector('.price')?.textContent?.trim()}))")
        print(rows)
    except PlaywrightTimeoutError:
        print("The expected product card did not become visible")
    finally:
        browser.close()

page.goto() does not throw merely because the server returns a valid HTTP error such as 404 or 500, so inspect the response status yourself. A response can also be successful while the application later displays an error; assert the page state you actually need.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Interact with hydrated controls

A button can be visible before its JavaScript event listener is attached. After clicking or filling, verify a URL change, a result count, a new DOM state or another observable outcome.

from playwright.sync_api import sync_playwright

with sync_playwright() as p:
    browser = p.chromium.launch(headless=True)
    page = browser.new_page()
    page.goto("https://example.com/search", wait_until="domcontentloaded")

    search = page.get_by_role("textbox", name="Search")
    search.fill("python")
    page.get_by_role("button", name="Search").click()

    results = page.locator("[data-testid='search-result']")
    results.first.wait_for(state="visible", timeout=30_000)
    print(results.all_text_contents())
    browser.close()

If filled text disappears, the application may have re-rendered during hydration. Fill again after the functional state is present and assert the resulting data rather than assuming the click succeeded.

Capture a stable snapshot after population

Do not call locator.all() once at startup and assume that list represents the final page. Resolve the locator after the target state is reached. For long lists, scroll deliberately and wait for a count or sentinel element that proves another page of results loaded.

Scrapy projects and browser integration

Playwright can be used alongside Scrapy, but directly embedding a browser in a spider can bypass Scrapy components. Prefer a maintained Scrapy integration when you need both crawling and browser rendering, and verify compatibility with the Scrapy and Playwright versions installed in your project. For a small job, a standalone Playwright script is often easier to operate.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Reliability, performance and responsible access

Make failures diagnosable

  • Set explicit navigation and locator timeouts and include the URL and selector in errors.
  • Save the HTML, screenshot or response body when a run fails.
  • Log HTTP status separately from browser exceptions.
  • Use a bounded retry policy for transient network failures, not for selector bugs.
  • Pin browser binaries in deployment and test after upgrades.

Reduce cost and runtime

  • Use the structured endpoint instead of a browser when it supplies complete data.
  • Block unnecessary images, fonts or analytics only when doing so does not change the data you need.
  • Reuse a browser process for multiple pages, but isolate contexts when cookies or identities must not leak.
  • Paginate at the data endpoint where possible rather than repeatedly scrolling a rendered page.

Check authorization

Rendering a page does not establish permission to collect its data. Review the site’s terms, robots guidance, account rules and applicable law for your location and use case. Avoid bypassing access controls, bot checks or authentication that you are not authorized to bypass.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Common errors and fixes

“The HTML is empty”

Confirm whether you inspected response source or the post-JavaScript DOM. If the source contains only an application shell, locate the XHR/fetch response or use Playwright.

“The selector times out”

Check the selector in the current DOM, confirm the correct frame, and verify that navigation reached the expected URL. Wait for a meaningful result condition rather than increasing a sleep.

“The click does nothing”

Check for hydration, overlays, disabled state and a required consent action. After clicking, assert the URL, result count or changed content.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

“The request works in DevTools but not in requests”

Compare method, URL, body, headers, cookies and form parameters. Remove copied browser headers gradually; retain only those required by the endpoint and your authorized session.

“The page shows an error but goto did not throw”

Inspect the main response status and assert an application-level error message or expected result. HTTP navigation completion is not a guarantee that the application succeeded.

Or skip the browser setup

ScreenshotNeo provides a one-call website screenshot API and MCP server. It accepts consent banners before capture and removes more than 60 known consent platforms, newsletter popups and chat widgets. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads and cache hits are not billed, and response headers identify the page verdict and billing result. Its MCP tools—take_screenshot, get_page_info and capture_pdf—let Claude, Cursor and other MCP clients work with pages without you building browser orchestration.

For screenshots rather than structured scraping, call the API directly. See the ScreenshotNeo documentation for request options.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

ScreenshotNeo supports full-page and element captures, device presets, custom viewports, dark mode, retina scale, PDF output, custom CSS and JavaScript, clicks, selector waits, delays, request blocking, headers, cookies, user agents, authorization, timezone, geolocation, transparent backgrounds, resizing, chosen cache TTLs, signed image links, asynchronous webhooks and bulk capture of up to 100 URLs per call. Every feature is available on every plan. The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.

Frequently Asked Questions

Can I scrape a JavaScript site with requests alone?

Yes, when the needed values are in the initial HTML, embedded JSON, or a network endpoint you can reproduce. JavaScript execution is not required in those cases.

Should I use synchronous or asynchronous Playwright?

Use the API style that matches your application. The synchronization rules are the same: wait for a meaningful locator or state and verify the result after interactions.

Why does a page work manually but fail headlessly?

Compare viewport, user agent, cookies, geolocation, timing and authentication state. Also check whether the site presents a bot challenge or requires an interaction your script has not performed.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How do I know whether a selector is stable?

Prefer semantic roles, labels, or deliberate data attributes over generated class names. Keep an assertion that proves the extracted value is the intended one.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.