Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
Use a real browser when JavaScript creates the content you need. A plain requests.get() call only receives the server’s initial response; it does not run the scripts that fill a product list, dashboard, table or “load more” view. In Python, launch Chromium with Playwright (or drive a WebDriver browser with Selenium), reproduce the user action, wait for a condition tied to the target data, and then either parse the rendered DOM or capture the JSON response that populated it. The examples below show both approaches, validation, troubleshooting and an API alternative.
Decide whether you need a browser
Start by inspecting the initial response. If the records, links or fields are already present in the HTML, use an HTTP client and an HTML parser; that is simpler, faster and easier to deploy. If the response contains an empty root element, a loading shell, or JavaScript bundles that fetch the records after navigation, a browser (or the underlying data request) is required.
Quick diagnostic
- Fetch the URL with
requestsand save the response. - Search that response for a distinctive value visible in the browser.
- If the value is absent, open the browser’s developer tools and inspect the Network panel for XHR or
fetchresponses. The data may be available as JSON even though it is not in the initial HTML.
Do not treat an empty parse as proof that a page has no data. It can mean that JavaScript has not run, the selector is wrong, the page requires a click, authentication or pagination, or the records arrive through a different response.
The Tool Desk
Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Install a Python browser automation stack
Playwright provides a Python API, browser contexts and request/response hooks in one package. Install it and its Chromium browser:
#1 Best Overall
python -m pip install playwright beautifulsoup4
python -m playwright install chromium
Selenium is also a valid choice when your team already operates WebDriver, a Selenium Grid or a browser fleet. The extraction principles—explicit readiness, interaction, validation and respectful rate limits—are the same. The runnable examples here use Playwright because its locators and network waits make the synchronization visible.
Capture the rendered DOM with Playwright
The following script navigates, performs a “Load more” action, waits for an actual result element, then sends the resulting HTML to BeautifulSoup. The URL, button name and CSS selector are illustrative: adapt them to the site you are allowed to collect.
from playwright.sync_api import sync_playwright, TimeoutError as PlaywrightTimeoutError
from bs4 import BeautifulSoup
URL = "https://example.com/results"
with sync_playwright() as p:
browser = p.chromium.launch(headless=True)
page = browser.new_page()
page.set_default_timeout(15_000)
try:
page.goto(URL, wait_until="domcontentloaded", timeout=45_000)
page.get_by_role("button", name="Load more").click()
page.locator("article.result").first.wait_for(state="visible")
html = page.content()
soup = BeautifulSoup(html, "html.parser")
rows = [
node.get_text(" ", strip=True)
for node in soup.select("article.result")
]
if not rows:
raise RuntimeError("The page became ready but no result nodes were found")
print(f"Parsed {len(rows)} results")
for row in rows:
print(row)
except PlaywrightTimeoutError as exc:
raise RuntimeError("The readiness condition was not reached") from exc
finally:
browser.close()
Why the wait is selector-based
domcontentloaded tells you that the initial document was parsed, not that the application finished rendering. A load event is also not a promise that later API calls and incremental rendering are complete. Waiting for a locator, assertion or other condition tied to the records you need is more reliable than sleeping for an arbitrary number of seconds.
Playwright exposes load, domcontentloaded, networkidle and commit navigation states. networkidle can be useful as a diagnostic, but ongoing analytics, polling or long-lived connections can make it a poor definition of “data ready.” Use a content-specific condition whenever possible.
Rank #2
Interact before extracting
Reproduce the path a user follows: fill a form, choose a filter, click a tab, accept a required consent dialog, or scroll an infinite list. Prefer role, label and text locators over brittle generated class names. If a list is virtualized, scrolling may be necessary before all records exist in the DOM; alternatively, capture the pagination or API response described below.
Capture the JSON response instead of the DOM
When the browser fetches a structured payload, parsing that response is usually less sensitive to visual markup changes. Wait for the matching response around the action that triggers it, then validate its shape.
from playwright.sync_api import sync_playwright
with sync_playwright() as p:
browser = p.chromium.launch(headless=True)
page = browser.new_page()
page.set_default_timeout(15_000)
page.goto("https://example.com/results", wait_until="domcontentloaded")
with page.expect_response("**/api/results") as response_info:
page.get_by_role("button", name="Load more").click()
response = response_info.value
if not response.ok:
raise RuntimeError(f"API returned HTTP {response.status}")
payload = response.json()
records = payload.get("results")
if not isinstance(records, list):
raise ValueError("Unexpected response schema: results is not a list")
for record in records:
print(record)
browser.close()
Confirm the endpoint, authentication, pagination parameters and schema for each site. A wildcard such as **/api/results is convenient for an example; in production, match the stable path and inspect query parameters so that an unrelated request cannot satisfy your wait.
Recommended Free Tools
Discovering the right request
- Open developer tools, select Network, and filter to Fetch/XHR.
- Perform the action that reveals the data.
- Inspect response content, status, request method, query parameters and required headers or cookies.
- Reproduce the action in Playwright and wait for that response.
Do not copy a private token into source control. Use an authenticated browser context or environment variables, and collect only what the site and applicable law permit.
Parse and validate safely
Extract only needed fields
Whether the source is DOM or JSON, select the smallest useful set of nodes or keys. Normalize whitespace and text, convert numbers and dates deliberately, and retain an identifier that lets you detect duplicates.
def clean_text(value):
return " ".join(str(value).split())
items = []
for node in soup.select("article.result"):
title = node.select_one("h2")
link = node.select_one("a[href]")
if not title or not link:
continue
items.append({
"title": clean_text(title.get_text(" ", strip=True)),
"url": link["href"],
})
if not items:
raise RuntimeError("No complete records; inspect readiness and selectors")
Fail loudly, not silently
- Check the HTTP status and final URL after redirects.
- Assert that a known heading, record count or JSON key exists.
- Log the selector, URL, timing and a short diagnostic (not credentials or private content).
- Detect a login page, bot challenge, consent wall or error message before parsing.
- Keep an expected-field check so a layout change raises an alert instead of producing an empty file.
Playwright or Selenium?
Both automate a real browser from Python. Choose based on the surrounding system rather than an unsupported claim that one is universally faster.
| Need | Playwright | Selenium |
|---|---|---|
| Modern locators and auto-waiting | Strong fit; locator operations wait for actionable states. | Available, but synchronization is commonly assembled with explicit waits. |
| Navigation and readiness | Navigation states plus locator/assertion waits in one API. | WebDriver navigation and explicit wait patterns. |
| Request/response capture | Built-in request and response monitoring and matching. | Possible through WebDriver features and ecosystem tools; implementation depends on browser and setup. |
| Existing grid or team expertise | Adopt when its browser contexts and API fit your deployment. | Prefer when your organization already runs WebDriver/Grid infrastructure. |
| Browser coverage | Ships managed browser binaries for supported engines. | Useful when you must connect to an established driver and browser matrix. |
Whichever you use, make readiness and validation explicit. A library choice cannot compensate for an incorrect selector, an unhandled login flow or a site that denies automated access.
Reliability, performance and operating costs
Use the narrowest readiness condition
Waiting for one target locator avoids both premature extraction and unnecessary idle time. Set navigation and operation timeouts appropriate to the site; keep a shorter diagnostic timeout for selectors so a changed layout is noticed quickly. A fixed sleep is only a last-resort workaround for an application with no observable condition.
Reuse contexts carefully
Launching a browser is expensive. For a batch, keep one browser process and create isolated contexts or pages as needed. Close pages and contexts promptly, cap concurrency to a level the destination permits, and add bounded retries with backoff for transient navigation failures. Do not retry authentication failures, access denials or deterministic selector errors indefinitely.
Reduce page weight when allowed
Block unnecessary images, ads or analytics only when doing so does not change the data path you need. A resource that appears “unrelated” may contain the API call or a script that renders the records. Measure memory and wall time on your own deployment; no universal benchmark applies across sites.
Pagination and infinite scroll
For numbered pages, loop until the next control is disabled or the response reports no more records, recording a stable ID to prevent duplicates. For infinite scroll, scroll in increments and wait for the count or last item to change. A direct paginated JSON request discovered in Network is generally more deterministic than scraping a virtualized viewport.
Compliance
Respect terms of service, robots guidance, access controls, privacy obligations and rate limits. Browser automation documentation explains how to operate a browser; it does not grant permission to collect a particular site’s data. Store only the fields you need and protect cookies, authorization headers and downloaded content.
Best Value
Common failures and fixes
| Symptom | Likely cause | Fix |
|---|---|---|
requests returns an empty shell |
Data is created after JavaScript runs. | Use Playwright/Selenium or identify and request the data endpoint. |
| Timeout waiting for a locator | Wrong selector, wrong state, slow request, consent wall or login. | Inspect the page, verify the exact role/text or CSS, handle the prerequisite flow, and wait for a meaningful condition. |
| HTML contains a loading spinner | Extraction happened after navigation but before rendering. | Wait for the first real record or an application-specific ready marker. |
| Response wait never resolves | Endpoint pattern or triggering action is incorrect. | Record Network traffic, match the method/path and place expect_response around the click or form submission. |
| JSON key is missing | Schema changed, an error payload was returned, or pagination differs. | Check status and content type, log a redacted payload sample, and validate the current schema. |
| Works headed but fails headless | Timing, viewport, browser differences or an access-control challenge. | Set a realistic viewport, replace sleeps with conditions, capture a trace/screenshot for diagnosis, and do not attempt to bypass a challenge. |
| Duplicate or missing records | Infinite-scroll virtualization, repeated retries or unstable pagination. | Deduplicate by a stable ID, track page/cursor state and verify expected counts. |
| Browser crashes or memory grows | Too many concurrent pages or unclosed contexts. | Bound concurrency, close resources in finally, and process batches. |
Or skip the browser setup
If your goal is a clean screenshot or PDF rather than custom Python parsing, ScreenshotNeo makes one HTTP request to capture a URL. It accepts cookie and consent banners before capture and removes more than 60 known consent platforms, newsletter popups and chat widgets; each step can be disabled. Only clean shots are billed: bot checks or CAPTCHAs, blank pages, timeouts, failed loads and cache hits cost nothing, and the response identifies the result with X-Page-Verdict and X-Billed headers.
For developers who still need structured extraction, keep Playwright in your Python pipeline. For visual capture, the API supports full-page shots with lazy images loaded, CSS-selector element capture, dark mode, device presets or custom viewports, retina scale, PDF paper and page options, custom CSS/JavaScript, clicks, selector or network-idle waits, request/resource blocking, headers, cookies, user agents, authorization, timezone, geolocation, transparent backgrounds, resizing, configurable-TTL caching, signed image links, asynchronous jobs with signed webhooks, bulk capture of up to 100 URLs per call, a usage API and an OpenAPI specification. Parameter names used by other screenshot APIs also work to ease migration.
cURL:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python:
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`HTTP ${res.status}`);
const buffer = Buffer.from(await res.arrayBuffer());
See the ScreenshotNeo documentation for option names and response details. Its MCP server gives Claude, Cursor and other MCP clients take_screenshot, get_page_info and capture_pdf tools. The Free plan includes 1,000 shots per month with no card; paid plans start at $5 for 3,000 shots. Sign up free for ScreenshotNeo.
Free tools Windows power users keep installed
One-click scans. No signup required.
FAQ
Can BeautifulSoup execute JavaScript?
No. BeautifulSoup parses HTML you provide; it does not run scripts. Supply it with HTML obtained after browser rendering, as in the Playwright example.
Should I save the browser’s final HTML or the API payload?
Save the payload when it contains the complete records you need and its schema is stable enough for your use case. Save rendered HTML when the data exists only after client-side transformations or when you need the exact displayed text.
How do I handle a page that requires login?
Automate an authorized login or load an approved authenticated browser state, keep credentials out of logs and source control, and check that the destination permits automated access. Do not defeat access controls.
Is networkidle always the best wait?
No. Polling, analytics and persistent connections can prevent a true idle period, and an idle network does not guarantee that the target records are present. A locator or assertion tied to the data is normally a better readiness signal.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

