Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
If a Python scraper returns an empty page while your browser shows data, first find out how that data reaches the browser. It may already be in the initial HTML, arrive in a separate JSON or HTML request, or depend on browser rendering and interaction. Reproduce the data request when practical; use Playwright or Selenium when the browser itself is necessary. This diagnostic approach is usually simpler to maintain than launching a browser for every page.
What makes a website dynamic?
A page can look dynamic for several different technical reasons. A server may return a mostly complete HTML document; JavaScript may request records after the initial page loads; or the visible content may depend on interaction such as scrolling, clicking, or choosing a filter. These cases call for different scraping methods. “Dynamic” by itself does not mean “use a headless browser.”
The useful distinction is between what your Python process receives from the server and what the browser eventually displays. Start by checking the initial response. Then, if the records are missing, observe the browser’s network requests to see whether the page fetches them separately. Scrapy’s guidance recommends finding the source request for the data and reproducing it when possible: Selecting dynamically-loaded content.
The Tool Desk
Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →First diagnose why the data is missing
Inspect the initial HTTP response
Make one request to the page and examine its status, headers, and body. This small check can tell you whether the response contains the target content, whether you received an error or redirect, and whether the server returned a different page than expected.
#1 Best Overall
import requests
url = "https://example.com/products"
response = requests.get(url, timeout=30)
print("status:", response.status_code)
print("content type:", response.headers.get("content-type"))
print(response.text[:2000])
Replace the example URL with a page you are allowed to access. If the response body contains the values you need, parse that response directly. If it contains HTML without the records, inspect whether the page embeds data in a script block or points to a separate data source. If it is an error, challenge, or unrelated response, do not assume that installing a browser will fix the underlying access or permission issue.
Find the request that supplies the records
- Open the page in a browser and open its developer tools’ Network panel.
- Reload the page and wait until the desired records appear.
- Look for requests made as the page renders or when you perform the interaction that reveals the records. Check likely fetch/XHR requests, their URLs, methods, query parameters, request bodies, and response formats.
- Open a candidate response and verify that it actually contains the fields you want. A request with a plausible name is not enough; confirm the data and its relationship to the page.
- Where permitted, reproduce that request from Python. Scrapy notes that matching the method and URL may be sufficient in some cases, but the request body, headers, or form parameters can also matter.
Use only the parameters and headers needed to make the permitted request. A browser-observed endpoint is not permission to access it, and its behavior may change. Review the site’s terms and applicable rules before collecting data.
Parse the response format you actually received
For JSON, decode JSON and work with its keys rather than trying to scrape rendered text. For HTML, use an HTML parser and selectors appropriate to the response. Keep fetching separate from extraction so you can check whether a failure comes from a changed request or a changed page structure.
Rank #2
import requests
api_url = "https://example.com/api/products"
response = requests.get(api_url, timeout=30)
response.raise_for_status()
data = response.json()
for item in data["products"]:
print(item.get("name"), item.get("price"))
The endpoint and JSON keys above are illustrative: replace them with the actual permitted request and schema you observed. Check for pagination, nested fields, and missing values rather than assuming one response contains every record.
Choose the least complex method that meets the requirement
| Approach | Best fit | Main trade-offs |
|---|---|---|
| HTTP client plus HTML or JSON parsing | The desired data is already in the response or comes from a reproducible endpoint. | Low browser overhead; your code handles request errors, pagination, parsing, and validation. |
| Scrapy | You are crawling multiple pages or need a reusable crawling pipeline. | Provides a framework for crawling and extraction, but you may still need to locate and reproduce browser-observed data requests. |
| Playwright | The result genuinely depends on browser rendering, interaction, or inspection of a browser-visible page. | Requires browser installation and execution. Explicit readiness checks matter because data can arrive after navigation. |
| Selenium WebDriver | Browser automation is needed and Selenium fits your project or team’s existing expertise. | A valid browser-automation alternative; compare it with Playwright based on your requirements rather than assuming a universal winner. |
For a small extraction from a known endpoint, a direct request may be all you need. For a multi-page crawl, Scrapy can organize the work. When the outcome depends on browser behavior, Playwright or Selenium may be appropriate. Scrapy’s dynamic-content guide discusses finding the data source, while the official Selenium WebDriver documentation describes Selenium’s browser automation approach.
Use Playwright when browser behavior is necessary
Playwright’s Python library supports Chromium, Firefox, and WebKit, with synchronous and asynchronous APIs. Installing its Python package and installing browser binaries are separate steps. For the synchronous API, install the package and then install the browsers:
python -m pip install playwright
python -m playwright install
Playwright’s installation instructions are at Getting started – Library. The following runnable example uses the synchronous API. Change the URL and locator to match the page and a specific element that indicates the data you need has appeared.
Recommended Free Tools
from playwright.sync_api import sync_playwright
url = "https://example.com/products"
with sync_playwright() as p:
browser = p.chromium.launch()
page = browser.new_page()
page.goto(url, wait_until="domcontentloaded", timeout=30_000)
# Wait for evidence that the target data is present, not just for navigation.
products = page.locator(".product-card")
products.first.wait_for(state="visible", timeout=15_000)
records = []
for product in products.all():
records.append({
"name": product.locator(".product-name").inner_text(),
"price": product.locator(".product-price").inner_text(),
})
print(records)
browser.close()
The CSS selectors are examples; inspect the target page and replace them with selectors that match its markup. This example waits for the first card, then reads the matching cards. If the page adds cards progressively or the list changes while you read it, first wait for a site-specific completion condition or otherwise establish that the set is stable.
Wait for the condition that matters
A navigation reaching the load event does not prove that late JavaScript data has arrived. Modern pages can fetch records after that event. Playwright’s navigation documentation explains its navigation and load states: Navigations. Prefer waiting for a target locator, a known response, or a page-specific state that demonstrates readiness.
Playwright locator actions auto-wait for actionability, but that does not mean every collection has finished changing. In particular, locator.all() returns the matches present immediately; its result can be unpredictable if the list is still changing. See the Locator API. Avoid arbitrary long sleeps as your only readiness check: they can waste time on fast pages and still fail on slow ones.
When interaction is part of the data path
If records appear only after a click, filter change, or scroll, automate that action and then wait for the resulting state or response. If the interaction merely triggers a data request, reproducing that request directly may remain simpler; if the browser state or rendered result itself is required, keep the interaction in the browser workflow.
Common errors and practical fixes
- The response is empty or missing the records: Check the response status and body first. If the records are absent, use the Network panel to find the request that supplies them. Parse that response directly when practical.
- Playwright times out waiting for a selector: Confirm that the selector matches the current page, that the relevant action has occurred, and that the page is the expected response rather than an error or consent screen. Use a locator tied to the actual data, not a generic page-load event.
- The browser shows content but Python does not: Compare the browser’s request with your Python request. Check method, URL, query parameters, request body, and required headers. Reproduce only what is necessary and allowed.
- Some records are missing: Determine whether the site paginates, loads more on scroll, or updates a list in stages. Inspect subsequent data requests or wait for a specific completion condition before enumerating results.
- Selectors stop working: The site’s markup may have changed, or you may be reading a different page state. Inspect the current HTML and revise the selectors; keep extraction code small enough to validate against sample records.
- Requests are slow or unreliable: Set timeouts, check response status, and handle network failures explicitly. Avoid unnecessary browser startup when an HTTP request can retrieve the same data. Keep collection volume appropriate for the site.
- The page blocks or challenges the request: Do not treat a challenge as a selector or timing bug. Review access rules and whether your collection is permitted; do not assume automation is authorization.
Validate results, crawl carefully, and review permission
Before scaling up, check that the output has the fields and shape you intended. Compare a few extracted records with the page or the observed data response, account for missing values, and confirm the record count makes sense for the page and pagination state. A scraper that runs without exceptions can still silently collect incomplete or misaligned data.
Best Value
Review the site’s terms and its robots.txt before collecting. The Robots Exclusion Protocol is standardized in IETF RFC 9309. Python’s urllib.robotparser can parse a robots file and answer whether a user agent may fetch a URL; see the Python documentation. Robots guidance is not a substitute for site-specific permission, terms review, or legal analysis. Rules and data availability vary by site, and a data endpoint may change over time.
Or skip the browser setup
If your goal is a visual capture rather than structured records, ScreenshotNeo returns a website screenshot or PDF from one API request. A screenshot is an image of the page, not a replacement for parsing JSON or extracting structured fields. Before a capture, ScreenshotNeo can accept cookie or consent banners and remove more than 60 known consent platforms, newsletter popups, and chat widgets; each of those steps can be turned off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and responses identify the page verdict and billing status in headers. Its MCP server provides take_screenshot, get_page_info, and capture_pdf tools for AI agents and MCP clients. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 shots. Every feature is on every plan.
Here is a one-request capture using Python. Replace the example URL and provide your API key:
Quick wins for a faster PC:
Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Repair Windows errors before they cause bigger problemsFix Now →import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
open("shot.webp", "wb").write(r.content)
For other request options and parameter compatibility, see the ScreenshotNeo documentation. cURL equivalent:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Sign up for 1,000 free screenshots a month with no card.
Frequently Asked Questions
Can I scrape a website just because its data endpoint is visible in developer tools?
No. A visible request helps explain how the page obtains data; it does not establish permission to collect or reuse that data. Check the site’s terms and applicable rules for your use case.
Does a screenshot API extract a page’s product names and prices as structured data?
No. A screenshot is a visual capture. For structured records, parse the permitted HTML or data response, or use browser automation to inspect the rendered page.
Free tools Windows power users keep installed
One-click scans. No signup required.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

