Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
If requests returns HTML without the content you see in a browser, the cause is usually that Requests does not execute JavaScript. Use a browser controlled by Pyppeteer when the page needs JavaScript, wait for the specific content or API response you need, and diagnose launch, navigation, readiness, network, and evaluation errors separately. If the site exposes a documented data endpoint, a direct HTTP request may be simpler.
Why Requests does not return JavaScript-rendered content
Python’s requests library fetches an HTTP response; it does not run the page’s JavaScript. A response can therefore contain the initial HTML shell while a browser later loads data from an API and inserts it into the DOM. If the text you need is absent from r.text but appears in a normal browser, first determine whether the page makes a documented JSON request. If it does, and you are authorized to use it, requesting that endpoint directly is often simpler than starting a browser.
When content is created by client-side JavaScript and there is no suitable direct endpoint, use a browser runtime such as Chromium with Pyppeteer. The key is not to wait an arbitrary amount of time: navigate, wait for the relevant response or page state, then extract the content.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Check the raw response first
import requests
url = "https://example.com/page"
r = requests.get(url, timeout=30)
r.raise_for_status()
print("URL:", r.url, "status:", r.status_code)
print("target present in raw HTML:", "target-text" in r.text)
This check distinguishes a missing JavaScript step from basic HTTP problems such as an unsuccessful status or unexpected redirect. It does not prove that a page is accessible to automation or that its data endpoint is public.
#1 Best Overall
Run Pyppeteer and wait for the content you need
Pyppeteer controls Chromium, so it can execute page JavaScript. Chromium must be installed or downloadable and runnable in the environment. Pyppeteer can download Chromium on first use; installation guidance documents the pyppeteer-install command and using a suitable existing Chrome binary. In containers and CI, verify the browser executable, permissions, and required Linux shared libraries before debugging page selectors. See the Pyppeteer repository and its installation notes.
Minimal asynchronous capture
import asyncio
from pyppeteer import launch
async def main():
browser = await launch(headless=True)
try:
page = await browser.newPage()
await page.goto(
"https://example.com/page",
{"waitUntil": "domcontentloaded", "timeout": 30_000},
)
await page.waitForSelector("#results", {"timeout": 30_000})
html = await page.content()
print(html)
finally:
await browser.close()
asyncio.run(main())
Replace the URL and selector with values for the actual page. domcontentloaded means the initial document was parsed; it does not mean an application has fetched and displayed its data. The selector wait expresses the condition that matters. Set bounded timeouts and close the browser in a finally block so it is shut down even if navigation or extraction fails.
Wait for an API response and then for the rendered state
If the page obtains results through a known request, wait for that response and for the corresponding DOM state. This helps distinguish an API that never succeeded from a page that received data but has not yet rendered it.
await page.waitForResponse(
lambda response: "/api/results" in response.url and response.status == 200,
{"timeout": 30_000},
)
await page.waitForFunction(
"() => document.querySelectorAll('#results li').length > 0",
{"timeout": 30_000},
)
Use a condition that reflects the target site’s real response URL and populated state. A successful HTTP status alone does not guarantee the expected data is present, and a longer timeout cannot fix an unauthorized, blocked, or failed API call. Pyppeteer documents waitForSelector, waitForFunction, and request and response waits in its API reference.
Rank #2
Fix common Pyppeteer errors by failure layer
Chromium fails to launch
- Confirm that Pyppeteer’s Chromium download completed, or provide the path to an installed, compatible browser using
executablePath. - Check file permissions and, on Linux, whether the shared libraries required by the browser are present.
- In a container or CI environment, check how the runtime handles Chromium’s sandbox. Do not add
--no-sandboxas a routine fix: understand the security implications and the environment first. - If the browser download is blocked, arrange an approved browser installation for the environment rather than repeatedly changing page-navigation code.
For example, when the environment has a known Chrome executable, set its real path rather than copying a placeholder:
browser = await launch(
headless=True,
executablePath="/usr/bin/google-chrome",
)
The path above is an example only; use the actual executable installed in your environment.
goto() times out or fails
Record the exception, requested URL, final URL if available, and response status. Check for an invalid URL, SSL error, failed main-document request, redirect, or a page that genuinely takes longer to respond. Increase the timeout only after checking those conditions. Pyppeteer’s navigation options and timeout behavior are documented in its API reference.
waitForSelector() times out
- Verify the selector in the live page’s DOM, including whether the content is inside an iframe or appears only after interaction.
- Check that the expected API request completed and returned usable data.
- Wait for a meaningful selector or predicate tied to the content, not merely the initial document load.
- Inspect the final URL and cookies if the page redirected to a sign-in, consent, or access-denied state.
A timeout identifies a condition that was not met before the deadline; it does not identify why. A selector typo, absent data, blocked request, and missing authentication need different fixes.
The page loads, but its data request fails
Inspect the relevant request and response, including status and URL. The page may require a session cookie or headers, or the request may be denied. Transfer cookies or headers only when the site requires them and you are entitled to access the data. If a documented endpoint is available, consider making a direct HTTP request instead of rendering the full page.
evaluate() says an expression is not a function
Pyppeteer tries to determine whether a string passed to evaluate() is a function or an expression. If it classifies an expression incorrectly, set force_expr=True. For a callback, pass an explicit function string and its argument:
text = await page.evaluate("document.body.textContent", force_expr=True)
heading_element = await page.querySelector("h1")
if heading_element is None:
raise RuntimeError("The page has no h1 element")
heading = await page.evaluate(
"element => element.textContent",
heading_element,
)
See Pyppeteer’s repository for its evaluation guidance and current project status.
Quick wins for a faster PC:
Clear out junk files and repair common Windows errorsFree Scan →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Repair Windows errors before they cause bigger problemsFix Now →Avoid races when an action triggers navigation
If a click triggers a full navigation, start waiting for navigation before clicking. Otherwise the navigation may begin before the wait is registered:
navigation = asyncio.ensure_future(
page.waitForNavigation({"waitUntil": "networkidle2"})
)
await page.click("a.next")
await navigation
await page.waitForSelector("#results")
Follow navigation with a content-specific wait. A site may change its URL using the History API without loading a new main document; in that case, a navigation wait is not a substitute for waiting on the new state or its API response. Pyppeteer documents navigation and waiting behavior in its API reference.
Use requests-html when you want its Requests-style parsing interface
requests-html combines a Requests-style session and HTML parsing with an optional browser-backed render step. Its documentation explains that rendering uses Pyppeteer and that the first call to render() downloads Chromium into the user’s home directory, such as ~/.pyppeteer/. This means it does not remove the browser installation and deployment considerations.
from requests_html import HTMLSession
session = HTMLSession()
r = session.get("https://example.com/page")
r.html.render(timeout=30, retries=2, wait=0.2)
items = r.html.find("#results li", first=False)
for item in items:
print(item.text)
The documented render options include retries, wait, sleep, reload, cookies, send_cookies_session, and keep_page. Use an option to address a known behavior rather than enabling all of them as a general cure. For asynchronous programs, use AsyncHTMLSession, await the response, and call await r.html.arender(...). Details are in the requests-html documentation.
Choose direct HTTP, Pyppeteer, or a maintained alternative
| Approach | Best fit | Trade-off |
|---|---|---|
| Direct Requests call | A stable, documented endpoint provides the data you need. | Does not execute page JavaScript or create browser-rendered content. |
| Pyppeteer | The required data appears only after a browser runs the page’s JavaScript, or you need browser controls and inspection. | Requires a runnable Chromium installation and careful waits; the project says it is unmaintained. |
| requests-html rendering | You want its Requests-style interface and HTML parsing around a browser render. | Rendering still depends on Pyppeteer and Chromium, including the documented first-use download. |
| Playwright for Python | You are starting new browser-automation work and want to assess a maintained alternative. | It is a separate automation library; migrating requires adapting code and deployment to its API and runtime. |
The Pyppeteer repository explicitly says it is unmaintained and recommends considering playwright-python as an alternative. That maintenance warning matters for new projects and long-lived deployments; it does not mean an existing Pyppeteer script cannot run in a controlled environment. Evaluate the alternative against your page requirements, async design, container or CI setup, and need for network inspection.
Best Value
Instrument the page before changing timeouts
Log the failure at the layer where it occurs. For a difficult page, collect the navigation exception, page errors, console output, failed requests, response status, final URL, relevant cookies, and the exact selector or JavaScript predicate being awaited. A compact set of listeners can reveal whether the page itself failed or merely failed to reach the expected state:
page.on("console", lambda message: print("console:", message.type, message.text))
page.on("pageerror", lambda error: print("page error:", error))
page.on("requestfailed", lambda request: print(
"request failed:", request.url, request.failure
))
page.on("response", lambda response: print(
"response:", response.status, response.url
))
Use this output carefully: requests and console messages may contain URLs or data you should not expose in logs. Avoid recording secrets such as authorization values or session cookies.
Or skip the browser setup
If your goal is a screenshot rather than extracting DOM data, ScreenshotNeo can return a screenshot or PDF from one API request. The example below captures a page to a WebP file; create an API key and replace the sample URL with your target. See the ScreenshotNeo API documentation for request options.
The Tool Desk
Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →curl -G "https://api.screenshotneo.com/v1/shot"
-d access_key=YOUR_API_KEY
--data-urlencode url=https://example.com
-o shot.webp
Before capture, it accepts the cookie or consent banner like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets; those steps can be turned off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing status. It also has an MCP server for AI agents such as Claude, Cursor, and other MCP clients. The free plan includes 1,000 screenshots per month without a card; paid plans start at $5 for 3,000 shots. These are plan allowances, not a substitute for Pyppeteer when you need to inspect or extract page data.
Sign up for ScreenshotNeo’s free plan to try 1,000 screenshots a month without a card.
Frequently Asked Questions
Does raising the timeout fix every Pyppeteer loading error?
No. It can help when a valid page or expected state needs more time, but it will not repair a blocked API request, missing authentication, invalid selector, or Chromium launch failure.
Is Pyppeteer maintained?
The Pyppeteer repository describes the project as unmaintained and recommends considering Playwright for Python as an alternative.
Free tools Windows power users keep installed
One-click scans. No signup required.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

