Free tools Windows power users keep installed
One-click scans. No signup required.
Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
When a page fills in data after it loads, its first HTML response may not contain what you want. The practical approach is to inspect the page’s network activity, then either request the data endpoint directly with Python or use Playwright to run the page’s JavaScript. Wait for the specific response or rendered content you need, and validate it before extracting data.
What makes an AJAX page different?
A traditional page can return its content in the initial HTML. An AJAX-style page may instead load a shell first, then use JavaScript to fetch data and update the page. That means a successful navigation—or a call to requests.get()—does not necessarily include the content visible in a browser.
Playwright’s navigation guide explains why a generic page-ready signal is unreliable: “There is no way to tell that the page is loaded, it depends on the page, framework, etc.” The page may do further work after the browser’s load event. See Navigations | Playwright Python.
Scraping such a page is therefore a choice between two routes: request suitable data directly, or automate a browser when rendering or interaction is necessary. Neither route makes a site’s endpoint stable or establishes that a particular use is permitted; check the target’s rules and applicable requirements.
#1 Best Overall
Inspect the page before writing the scraper
- Open the page in a browser. Open Developer Tools and select the Network panel.
- Reload and reproduce the relevant action. If results appear only after clicking a button, changing a filter, or scrolling, perform that action while observing requests.
- Find the response that contains the data. Look for XHR or fetch requests and inspect their response. Determine whether the useful content is JSON, HTML, or another format.
- Record request details. Note the method, URL and query parameters, and whether the request appears to depend on a session or other state.
- Choose the least complex route that fits. If an appropriate HTTP request returns the data and you can reproduce it, use a direct request. If the page must execute JavaScript or interact with controls to reveal the data, use a browser.
Playwright can monitor browser network activity, including XHR and fetch, and its network guide shows how to wait for a response after an action: Network | Playwright Python. Treat a discovered endpoint as a clue to how the page works, not as proof that it is intended for unrestricted or high-volume access.
Choose direct HTTP or browser automation
| Question | Direct HTTP request | Playwright browser |
|---|---|---|
| Does the data arrive without page JavaScript? | Usually the simpler option when a suitable request returns the needed data. | Useful when the page itself must run JavaScript. |
| Where is the data? | Often in a structured endpoint response, such as JSON. | May be read from a network response or from the rendered DOM. |
| What do you wait for? | The HTTP response, followed by status and body validation. | The particular response or page-content condition associated with the task. |
| What does setup involve? | Python HTTP client and response parsing. | Browser installation, page lifecycle, interaction, and cleanup. |
| What operational caveat matters? | Endpoint behavior and access requirements are specific to the target. | Playwright’s Python API is not thread-safe; account for this if using multiple threads. |
Playwright’s Python library also includes APIRequestContext for sending HTTP requests, alongside page APIs for observing browser traffic. Using Playwright does not mean every request has to be made by a rendered page; choose the method that suits the data and workflow. See the network guide and Getting started – Library.
Request an appropriate endpoint directly with Python
If inspection shows that a suitable endpoint returns the needed data without browser-side execution, use an HTTP client and parse its response. The following is a template, not a working endpoint for a particular site. Replace the URL and parameters with values observed for your target, and adapt parsing to its response format.
Rank #2
import requests
url = "https://example.com/api/data" # Replace with the observed endpoint.
params = {"page": 1} # Replace with the target's required parameters.
response = requests.get(url, params=params, timeout=30)
response.raise_for_status()
payload = response.json()
print(payload)
raise_for_status() makes an unsuccessful HTTP status an explicit failure rather than allowing the script to treat an error page as the intended data. If the response is HTML rather than JSON, parse it as HTML instead of calling response.json(). Add only request details that the target actually requires; do not assume an endpoint will behave the same way without inspecting it.
Use Playwright when the page must run JavaScript
Install Playwright for Python and its browser with the documented commands:
python -m pip install playwright
python -m playwright install chromium
The example below waits for a response triggered by a click, checks its HTTP status, and parses JSON. The domain, endpoint pattern, and control label are illustrative placeholders: replace them with the page, response pattern, and locator you observed. The response matcher should be narrow enough to identify the request you need.
from playwright.sync_api import sync_playwright
with sync_playwright() as p:
browser = p.chromium.launch()
page = browser.new_page()
try:
page.goto("https://example.com") # Replace with the target page.
with page.expect_response("**/api/data") as response_info:
page.get_by_text("Load data").click() # Replace with the target control.
response = response_info.value
if not response.ok:
raise RuntimeError(f"Unexpected status: {response.status}")
payload = response.json()
print(payload)
finally:
browser.close()
The event-wait pattern is documented in Playwright’s network guide. If the data is rendered into the DOM and there is no useful response to capture, wait for a locator or a content condition that represents the result you need, then read that element. Avoid treating navigation completion as proof that later content is ready.
Wait for the right condition, not an arbitrary pause
When an action triggers a known request
Wrap the action in page.expect_response(), as in the example above. Match the response by a URL pattern or a predicate that identifies the right request. This avoids racing ahead of the request triggered by a click.
When the result is visible in the page
Wait for the relevant locator or an explicit content condition. A fixed delay may be too short on a slow response and needlessly long on a fast one; use it only when a delay itself is genuinely part of the task, not as the default definition of readiness.
When the page appears loaded but data is missing
Check whether the page makes additional requests after the initial navigation, or only after a user action or scroll. Playwright’s navigation guidance describes why pages may continue working after load: Navigations | Playwright Python.
Validate responses before extracting data
A completed request is not necessarily a successful one. Playwright documents that HTTP error responses—including statuses such as 404 and 503—still complete as HTTP responses. Check the status and inspect the body before assuming it contains the expected data: Page | Playwright Python.
- Confirm the expected response arrived. A response to a different request may match a broad URL pattern.
- Check the HTTP status. Decide which statuses your task accepts; raise or report an actionable error otherwise.
- Check the body’s shape. Verify expected keys or elements exist before indexing or transforming them.
- Handle missing or changed data explicitly. A valid response can still have a different structure from the one your parser expects.
For a browser workflow, raise an error or log useful context when a wait times out; do not silently return an empty result. For direct requests, apply the same principle: validate both the response and the parsed data.
Best Value
Troubleshoot common failures
| Symptom | Likely cause | What to check or change |
|---|---|---|
| The initial HTML has no desired data. | The page fetches it after navigation. | Inspect Network activity while reloading and triggering the relevant action. Use the data request directly if suitable, or automate the page. |
| The script continues before results appear. | Navigation completed, but asynchronous work has not. | Wait for the matching response or a locator/content condition instead of relying on load or a routine fixed sleep. |
| The response wait times out. | The action did not trigger the request, or the matcher does not match its URL. | Reproduce the action in Network tools; check the URL pattern, method, and whether a different interaction is required. |
| A response arrived, but extraction fails. | It may be an HTTP error or have a different body format or structure. | Inspect status and body, then validate the expected keys or elements before parsing. |
| Request interception misses some traffic. | A service worker may handle requests outside the route interception you expect. | Check Playwright’s service-worker caveat; its Page reference recommends blocking service workers when request routing must observe those requests. See Page | Playwright Python. |
| A multithreaded scraper behaves unpredictably. | Playwright’s Python API is not thread-safe. | If threads are necessary, create a Playwright instance per thread. See Getting started – Library. |
Keep the scraper practical and responsible
Direct HTTP calls avoid browser rendering when the endpoint fits, while browser automation adds browser setup and lifecycle work. Choose based on the target’s actual behavior rather than assuming one method is universally faster or more reliable; no general performance figure follows from the documentation cited here.
- Use a specific response matcher and stable page-content conditions so failures point to a concrete missing event or element.
- Set timeouts and surface failures with context such as the target action or unexpected status.
- Check the site’s terms and applicable rules for your use case. Technical documentation does not determine whether collecting a particular site’s data is permitted.
Or skip the browser setup
If your goal is a screenshot of a JavaScript-rendered page rather than extracting its underlying data, ScreenshotNeo provides a website screenshot API and MCP server. It can return PNG, JPEG, WebP, or PDF; it is a screenshot service, not a general-purpose data-extraction API.
One GET request can capture a page. This cURL example saves a WebP screenshot of the target URL; replace the URL with the page you want to capture. See the ScreenshotNeo documentation for API details.
The Tool Desk
Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp
- Cookie/consent banners are accepted before capture, and 60+ known consent platforms, newsletter popups, and chat widgets are removed; each step can be turned off.
- Bot checks/CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing; response headers identify the page verdict and whether it was billed.
- An MCP server provides
take_screenshot,get_page_info, andcapture_pdftools for Claude, Cursor, and other MCP clients. - The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 screenshots. Every feature is on every plan.
Sign up for 1,000 free screenshots a month—no card required.
Frequently Asked Questions
Does every AJAX page need Playwright?
No. If an appropriate endpoint returns the needed data and your Python request can reproduce it, direct HTTP may be enough; use a browser when rendering or interaction is required.
Can I use ScreenshotNeo to extract the data behind a page?
ScreenshotNeo returns screenshots or PDFs; it is not a general-purpose data-extraction API. Use the endpoint or browser workflow described above when you need the page’s underlying data.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.
Do these 3 things before closing this tab:
1Fix the driver behind crashes, sound loss and screen glitches2Clear out junk files and repair common Windows errors3Scan for outdated or missing drivers - takes under a minute

