Do these 3 things before closing this tab:
1Repair Windows errors before they cause bigger problems2Scan for outdated or missing drivers - takes under a minute3Clear out junk files and repair common Windows errorsSome links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
A Pyppeteer navigation failure is not automatically proof that a website has blocked your scraper. First record the HTTP response (if any), final URL, exception, and returned page content. That evidence separates a server refusal such as 403 or 429 from an SSL error, timeout, invalid URL, browser-launch problem, or a failure in the page’s main resource. Then check the site’s published rules and use an approved access route. If automation is explicitly refused, stop rather than trying to disguise the bot or solve a CAPTCHA.
Start by capturing the failure, not by changing identities
Pyppeteer’s Page.goto can return the main-resource response or raise an exception. SSL failures, malformed URLs, timeouts and a failed main resource can therefore look very different from a site-generated denial page. Log enough context to tell those cases apart.
import asyncio
from pyppeteer import launch
async def inspect(url: str):
browser = await launch(headless=True, args=["--no-sandbox"])
page = await browser.newPage()
try:
response = await page.goto(url, {
"waitUntil": "domcontentloaded",
"timeout": 30_000,
})
print("requested:", url)
print("final:", page.url)
print("status:", response.status if response else None)
print("title:", await page.title())
html = await page.content()
print("body preview:", html[:1_000])
await page.screenshot({"path": "failure.png", "fullPage": True})
except Exception as exc:
print("requested:", url)
print("final:", page.url)
print("exception:", repr(exc))
finally:
await browser.close()
asyncio.run(inspect("https://example.com"))
Save the requested URL, final URL, status, exception text, timestamp, and a screenshot or HTML snapshot when your terms and privacy obligations permit it. A response object means the browser reached a main resource; an exception means you must investigate navigation, transport, certificate, URL or browser state as well as the target site.
Read the status and page together
- 403: commonly means the server refused the request, but the reason is site-specific. Inspect the response body for an explanation or an instruction to contact the operator.
- 429: indicates that the request rate is too high under HTTP semantics. Look for a
Retry-Afterheader and follow it. - 3xx: compare the requested and final URLs. A redirect to a login, consent or challenge page changes what your script actually received.
- 200 with a challenge page: a successful transport status does not mean that the intended content was delivered. Check the title, visible text and key selectors.
- No response and an exception: investigate DNS, TLS, timeout, invalid URL, browser launch and main-resource errors before calling it a block.
Honor rate limits and Retry-After
When a server returns 429 or another explicit rate signal, reduce concurrency and request frequency. HTTP’s Retry-After value is either a delay in seconds or an HTTP date. Parse it and wait at least that long before a follow-up request; if it is absent, use a conservative, documented backoff rather than immediately retrying.
#1 Best Overall
import asyncio
import email.utils
import time
from datetime import datetime, timezone
async def wait_from_retry_after(value: str | None):
if not value:
await asyncio.sleep(60)
return
try:
delay = max(0, int(value))
except ValueError:
target = email.utils.parsedate_to_datetime(value)
if target.tzinfo is None:
target = target.replace(tzinfo=timezone.utc)
delay = max(0, int((target - datetime.now(timezone.utc)).total_seconds()))
await asyncio.sleep(delay)
Do not treat retries as a way to wear down a refusal. A rate limit is a signal to slow down, review your workload and confirm that your use is permitted.
Check the site’s rules before continuing
Review the applicable host’s robots.txt, terms of service, API documentation, and any data-access or support page. A robots file communicates crawler preferences and can help manage crawler traffic; it is not an access-control mechanism, and some crawlers may ignore it. Its rules are scoped to the protocol, host and port where that file is served, so do not assume a rule on one host controls another.
Look for an official API, export, feed, partnership or written permission. Confirm whether authentication is required, which endpoints are intended for automation, and any published quota. Keep a record of the permission or license that covers your use and the fields you actually collect. The target site and your jurisdiction determine the legal meaning of its terms; a generic diagnostic guide cannot make that determination for you.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
When the site explicitly says stop
Stop automated requests when the site presents a CAPTCHA, asks you to sign in, states that automated access is forbidden, or tells you to contact the operator. Do not recommend proxy rotation, user-agent disguise, fingerprint manipulation or CAPTCHA-solving as routine fixes for an explicit denial. Those techniques attempt to defeat the site’s decision rather than resolve the underlying permission question.
- Ask the operator for access or a documented API.
- Use a licensed dataset or an export supplied by the site.
- Reduce the project to pages and endpoints the site identifies as public and automatable.
- Preserve the denial response and your contact request so the team can demonstrate that it paused responsibly.
Separate access permission from Pyppeteer’s maintenance
The Pyppeteer repository describes the project as unmaintained and recommends Playwright for Python. Replacing the library can fix stale browser integration, dependency incompatibilities and missing features, but it does not grant permission to access a site and cannot guarantee that a target will allow automation.
What Playwright changes
Playwright for Python provides both synchronous and asynchronous APIs and supports Chromium, WebKit and Firefox. Choose the API style that matches your existing tests and application. Its browser-engine coverage can help when a defect is specific to Chromium, but test the target and your own workload rather than assuming a different engine will bypass a policy.
Migration checklist
- Pin a current Playwright version in a reproducible environment and install the browser binaries required by your deployment.
- Translate launch, page creation, navigation timeout and screenshot code one operation at a time.
- Replace Pyppeteer’s selectors and event handlers with Playwright locators and explicit waits; remove sleeps that were compensating for timing races.
- Retain the logging from your diagnostic harness: requested URL, final URL, status, exception and a permitted snapshot.
- Run against a staging page or an approved target before increasing concurrency.
Compare maintenance signals, browser-engine requirements, API style, migration effort and fit with your existing tests. The available documentation does not establish a benchmark or a site-specific success rate, so do not promise that Playwright will defeat a 403 or CAPTCHA.
PC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minuteBuild a safe decision tree
- Did navigation return a response? If no, classify the exception as URL, TLS, timeout, network, browser or main-resource failure and fix that class first.
- What did the response deliver? Record status, final URL, title, body markers and any challenge or login text.
- Is the signal a rate limit? Honor
Retry-After, reduce activity and review your quota. - Do the site’s rules permit this activity? If unclear, pause and ask. If explicitly prohibited, stop and use an approved route.
- Is the remaining problem library maintenance? Plan a Playwright migration without changing the permission decision.
Troubleshooting common Pyppeteer symptoms
“Navigation Timeout Exceeded”
Check whether the page is slow, waiting on a never-ending resource, or failing before the main resource completes. Capture the final URL and partial content, then use a realistic timeout and a deliberate waitUntil condition. Do not respond by firing more concurrent retries.
Rank #3
SSL, DNS or invalid-URL exceptions
Validate the URL scheme and hostname, resolve DNS from the deployment network, and inspect the certificate chain. A transport failure is not evidence of a site block. Correct the environment or ask the site operator about network restrictions.
A 403 page
Preserve the response body and headers, check the terms and API route, and contact the operator if access is legitimate. Changing the user agent or rotating proxies is not an approved remedy for an explicit refusal.
A 429 page or Retry-After header
Stop the queue, wait the requested interval, lower concurrency and verify that your request volume fits the documented quota. Add bounded retries only after the wait.
CAPTCHA, “verify you are human” or login wall
Treat it as an access-control signal. Do not automate the challenge or search for a bypass. Use an API, permission, export or a human-reviewed workflow that the site allows.
Blank or incomplete content
Determine whether the main resource failed, JavaScript has not rendered, content is behind authentication, or the site intentionally serves a different response to automation. Compare the final URL, status, visible text and a permitted screenshot. Ask the operator for the supported access method instead of escalating evasion.
Operational safeguards for a permitted crawler
- Use a queue with a global rate limit and per-host concurrency limit.
- Cache responses where the site’s rules allow it, and avoid repeatedly fetching unchanged pages.
- Set a clear user agent that identifies your project and a contact address when appropriate.
- Store only the data you need, protect credentials and redact sensitive response bodies in logs.
- Make retries bounded, observable and cancellable; never retry a permanent denial indefinitely.
- Test navigation and parsing separately so a selector bug is not mistaken for an access block.
Or skip the browser setup
For a permitted screenshot rather than a full browser automation workflow, ScreenshotNeo provides a single HTTP request that returns PNG, JPEG, WebP or PDF. Before capture it accepts cookie or consent banners and removes more than 60 known consent platforms, newsletter popups and chat widgets; each step can be disabled. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads and cache hits are not billed, and response headers identify the page verdict and billing status. It also offers an MCP server with take_screenshot, get_page_info and capture_pdf for Claude, Cursor and other MCP clients.
Use the API only where the target permits automated capture. The complete option set includes full-page shots with lazy images loaded, CSS-selector element capture, dark mode, 12 device presets or a custom viewport, retina scale, PDF paper size/margins/landscape/page ranges, HTML/CSS-to-image, custom CSS and JavaScript, pre-capture clicks, hidden selectors, selector/delay/network-idle waits, blocking ads/trackers/requests/resource types, custom headers/cookies/user agent/Authorization, timezone and geolocation, transparent backgrounds, resizing, chosen cache TTL, signed image links, asynchronous jobs with signed webhooks, 100-URL bulk calls, usage API and OpenAPI specification. Common parameter names used by other screenshot APIs also work.
Recommended Free Tools
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
See the ScreenshotNeo documentation for authentication and options. Python:
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`${res.status} ${res.statusText}`);
const fs = await import('node:fs/promises');
await fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer()));
Pricing and fit
The Free plan includes 1,000 shots per month with no card. Paid plans are Starter $5 for 3,000 shots, Growth $15 for 15,000, Pro $39 for 60,000, Scale $99 for 250,000, and Business $249 for 1,000,000; yearly billing gives two months free, and every feature is on every plan. Only clean shots are billed, so a failed load is not charged. Create a free ScreenshotNeo account to start with 1,000 shots and no card.
Best Value
FAQ
Does a 403 prove that Pyppeteer is blocked?
No. It is evidence of a server refusal, but only the response, page content, rules and site-specific explanation establish why.
Can robots.txt authorize scraping?
No. It communicates crawler preferences; it is not an access-control mechanism or a legal permission grant.
Quick wins for a faster PC:
Repair Windows errors before they cause bigger problemsFix Now →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Will Playwright bypass a CAPTCHA?
There is no such guarantee, and bypassing an explicit restriction is not an appropriate remedy. Playwright is a maintained automation alternative, not permission to access a site.
Should I keep retrying a timeout?
Only after classifying the cause and applying a bounded, rate-limited retry policy. Repeating an unexplained failure can increase load and obscure the real problem.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

