Quick wins for a faster PC:
Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Repair Windows errors before they cause bigger problemsFix Now →Scan for outdated or missing drivers - takes under a minuteDriver Scan →Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
If a value appears in a browser but not in requests.get(...).text, the browser is probably obtaining it after the initial HTML response. Find that data source first; render the page with Playwright only when the source cannot reasonably be reproduced or when you need browser interaction and the rendered DOM.
This guide shows how to diagnose the difference, extract embedded or network-delivered data, render JavaScript with Python, wait for application state instead of arbitrary delays, and recover from common failures.
Why dynamic content is missing from a Python response
An HTTP client receives the server response. It does not execute the JavaScript that a browser runs afterward. A page can therefore return a small HTML shell while scripts fetch products, comments, prices or search results from another endpoint.
A visible element is not proof that its value was present in the first response. The value may be:
#1 Best Overall
- Already in the original HTML but hidden by CSS.
- Embedded in a
<script>element as JSON or JavaScript data. - Returned by an XHR or
fetchrequest after navigation. - Created only after an interaction, such as a click, scroll, login or form submission.
Scrapy’s current documentation recommends finding the data source and extracting it directly when possible. That normally produces more structured input with less browser overhead than rendering an entire page.
Diagnose the page before choosing a scraper
Compare source HTML with the browser DOM
- Fetch the URL with
requestsand saveresponse.text. - Use the browser’s “View source” and developer-tools Elements panel. “View source” represents the response; Elements represents the current DOM.
- Search both for a distinctive value, a CSS class, or the element’s label.
- Inspect script tags for JSON, serialized state, or configuration containing the data.
import requests
url = "https://example.com/catalog"
r = requests.get(url, timeout=30)
r.raise_for_status()
print("status:", r.status_code)
print("value in response:", "Example product" in r.text)
print(r.text[:500])
If the value is in the response, parse that response instead of launching a browser. If it is absent, open the Network panel, reload the page, and filter for Fetch/XHR requests. Identify the request whose response contains the required records.
Record the complete data request
When reproducing a request, compare its method, URL, query string, body, headers, cookies and form parameters. A copied URL without the request body or an authorization header may return an empty or different result. “Copy as cURL” in browser developer tools is a useful starting point; translate the relevant parts into Python and remove incidental headers one at a time.
Choose the least complex working approach
| Approach | Choose it when | Main trade-off |
|---|---|---|
| Initial HTML or embedded data | The values are in the response or a script payload | Lowest overhead, but the response shape must remain parseable. |
| Reproduced data request | Network inspection reveals a structured endpoint | Usually less rendering and parsing work; request details and permitted access must be understood. |
| Playwright with Python | JavaScript execution, interaction, or the rendered DOM is essential | Highest browser fidelity, with greater runtime, memory and page-change sensitivity. |
| Scrapy plus browser integration | A crawler needs Scrapy facilities as well as browser rendering | Preserves more crawler components but adds integration and compatibility setup. |
This is a qualitative decision, not a benchmark. Prefer the smallest method that reliably returns the fields you need.
Extract embedded JSON without rendering
Many applications place a JSON object in a script tag. Parse it as JSON when it is valid JSON; do not treat a regular expression as a general JavaScript parser. JavaScript object literals can contain single quotes, comments, trailing commas and expressions that require a JavaScript-aware parser.
Rank #2
import json
import requests
from bs4 import BeautifulSoup
r = requests.get("https://example.com", timeout=30)
r.raise_for_status()
soup = BeautifulSoup(r.text, "html.parser")
for script in soup.find_all("script"):
text = script.string or script.get_text()
if '"products"' not in text:
continue
try:
state = json.loads(text)
except json.JSONDecodeError:
continue
products = state.get("products", [])
for product in products:
print(product.get("name"))
break
Use a stable script identifier or a documented state property when available. Validate that the expected keys exist and log a useful error when the site changes its serialization.
Reproduce the browser’s data request with Python
Suppose the Network panel shows a POST request returning JSON. Recreate only the material request properties:
Do these 3 things before closing this tab:
1Scan for outdated or missing drivers - takes under a minute2Repair Windows errors before they cause bigger problems3Fix the driver behind crashes, sound loss and screen glitchesimport requests
endpoint = "https://example.com/api/search"
payload = {"query": "laptops", "page": 1}
headers = {
"Accept": "application/json",
"Content-Type": "application/json",
"User-Agent": "my-research-client/1.0",
}
r = requests.post(endpoint, json=payload, headers=headers, timeout=30)
r.raise_for_status()
data = r.json()
for row in data.get("results", []):
print(row.get("title"))
If the browser sends form data, use data= rather than json=. If it sends query parameters, use params=. Preserve authentication and session cookies only when you are authorized to use them. Check pagination, rate limits and whether the endpoint is intended for your access.
Render JavaScript with Playwright Python
Install and launch
The example assumes a current Playwright for Python installation. Check the version-specific installation instructions for your environment before pinning this in production.
python -m pip install playwright
python -m playwright install chromium
Wait for the result, not an elapsed time
Playwright performs actionability checks before actions, and locators resolve against the current DOM. Use a locator for the result and an assertion or state condition. Fixed sleeps can pass on a fast run and fail on a slow one; Playwright’s documentation discourages timeout waits for production readiness checks. The networkidle state is also discouraged as a generic test of readiness because modern pages may keep connections open or perform work after it occurs.
from playwright.sync_api import sync_playwright, TimeoutError as PlaywrightTimeoutError
URL = "https://example.com/catalog"
with sync_playwright() as p:
browser = p.chromium.launch(headless=True)
page = browser.new_page()
try:
response = page.goto(URL, wait_until="domcontentloaded", timeout=60_000)
if response is None:
raise RuntimeError("Navigation produced no main response")
if response.status >= 400:
raise RuntimeError(f"HTTP status from main document: {response.status}")
cards = page.locator("[data-testid='product-card']")
cards.first.wait_for(state="visible", timeout=30_000)
rows = cards.evaluate_all("els => els.map(e => ({name: e.querySelector('h2')?.textContent?.trim(), price: e.querySelector('.price')?.textContent?.trim()}))")
print(rows)
except PlaywrightTimeoutError:
print("The expected product card did not become visible")
finally:
browser.close()
page.goto() does not throw merely because the server returns a valid HTTP error such as 404 or 500, so inspect the response status yourself. A response can also be successful while the application later displays an error; assert the page state you actually need.
The Tool Desk
Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Interact with hydrated controls
A button can be visible before its JavaScript event listener is attached. After clicking or filling, verify a URL change, a result count, a new DOM state or another observable outcome.
from playwright.sync_api import sync_playwright
with sync_playwright() as p:
browser = p.chromium.launch(headless=True)
page = browser.new_page()
page.goto("https://example.com/search", wait_until="domcontentloaded")
search = page.get_by_role("textbox", name="Search")
search.fill("python")
page.get_by_role("button", name="Search").click()
results = page.locator("[data-testid='search-result']")
results.first.wait_for(state="visible", timeout=30_000)
print(results.all_text_contents())
browser.close()
If filled text disappears, the application may have re-rendered during hydration. Fill again after the functional state is present and assert the resulting data rather than assuming the click succeeded.
Capture a stable snapshot after population
Do not call locator.all() once at startup and assume that list represents the final page. Resolve the locator after the target state is reached. For long lists, scroll deliberately and wait for a count or sentinel element that proves another page of results loaded.
Scrapy projects and browser integration
Playwright can be used alongside Scrapy, but directly embedding a browser in a spider can bypass Scrapy components. Prefer a maintained Scrapy integration when you need both crawling and browser rendering, and verify compatibility with the Scrapy and Playwright versions installed in your project. For a small job, a standalone Playwright script is often easier to operate.
Reliability, performance and responsible access
Make failures diagnosable
- Set explicit navigation and locator timeouts and include the URL and selector in errors.
- Save the HTML, screenshot or response body when a run fails.
- Log HTTP status separately from browser exceptions.
- Use a bounded retry policy for transient network failures, not for selector bugs.
- Pin browser binaries in deployment and test after upgrades.
Reduce cost and runtime
- Use the structured endpoint instead of a browser when it supplies complete data.
- Block unnecessary images, fonts or analytics only when doing so does not change the data you need.
- Reuse a browser process for multiple pages, but isolate contexts when cookies or identities must not leak.
- Paginate at the data endpoint where possible rather than repeatedly scrolling a rendered page.
Check authorization
Rendering a page does not establish permission to collect its data. Review the site’s terms, robots guidance, account rules and applicable law for your location and use case. Avoid bypassing access controls, bot checks or authentication that you are not authorized to bypass.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Common errors and fixes
“The HTML is empty”
Confirm whether you inspected response source or the post-JavaScript DOM. If the source contains only an application shell, locate the XHR/fetch response or use Playwright.
“The selector times out”
Check the selector in the current DOM, confirm the correct frame, and verify that navigation reached the expected URL. Wait for a meaningful result condition rather than increasing a sleep.
“The click does nothing”
Check for hydration, overlays, disabled state and a required consent action. After clicking, assert the URL, result count or changed content.
“The request works in DevTools but not in requests”
Compare method, URL, body, headers, cookies and form parameters. Remove copied browser headers gradually; retain only those required by the endpoint and your authorized session.
Best Value
“The page shows an error but goto did not throw”
Inspect the main response status and assert an application-level error message or expected result. HTTP navigation completion is not a guarantee that the application succeeded.
Or skip the browser setup
ScreenshotNeo provides a one-call website screenshot API and MCP server. It accepts consent banners before capture and removes more than 60 known consent platforms, newsletter popups and chat widgets. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads and cache hits are not billed, and response headers identify the page verdict and billing result. Its MCP tools—take_screenshot, get_page_info and capture_pdf—let Claude, Cursor and other MCP clients work with pages without you building browser orchestration.
For screenshots rather than structured scraping, call the API directly. See the ScreenshotNeo documentation for request options.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
ScreenshotNeo supports full-page and element captures, device presets, custom viewports, dark mode, retina scale, PDF output, custom CSS and JavaScript, clicks, selector waits, delays, request blocking, headers, cookies, user agents, authorization, timezone, geolocation, transparent backgrounds, resizing, chosen cache TTLs, signed image links, asynchronous webhooks and bulk capture of up to 100 URLs per call. Every feature is available on every plan. The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.
Frequently Asked Questions
Can I scrape a JavaScript site with requests alone?
Yes, when the needed values are in the initial HTML, embedded JSON, or a network endpoint you can reproduce. JavaScript execution is not required in those cases.
Should I use synchronous or asynchronous Playwright?
Use the API style that matches your application. The synchronization rules are the same: wait for a meaningful locator or state and verify the result after interactions.
Why does a page work manually but fail headlessly?
Compare viewport, user agent, cookies, geolocation, timing and authentication state. Also check whether the site presents a bot challenge or requires an interaction your script has not performed.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
How do I know whether a selector is stable?
Prefer semantic roles, labels, or deliberate data attributes over generated class names. Keep an assertion that proves the extracted value is the intended one.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

