Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

You can scrape a JavaScript-rendered page with Pyppeteer by launching Chromium, opening a page, waiting for the content you need, extracting a narrow set of text or attributes, and closing the browser in a finally block. However, Pyppeteer’s own README currently says the repository is unmaintained and recommends Playwright for Python. Keep Pyppeteer for an existing script or a learning exercise only after checking that its browser and API compatibility meet your needs; for a new production system, evaluate the maintained alternative before committing.

What Pyppeteer does—and the maintenance warning

Pyppeteer is an unofficial Python port of Puppeteer, the browser-automation library for headless Chrome and Chromium. Unlike an HTTP client that receives only the initial response, it runs a browser, executes page JavaScript, and lets you inspect the rendered DOM.

The project README includes this notice: “Attention: This repo is unmaintained and has been outside of minor changes for a long time. Please consider playwright-python as an alternative.” That warning matters for security updates, browser-version compatibility, documentation freshness, and production support. Existing Pyppeteer code may remain useful, but a new project should compare migration cost, required APIs, the browser version you must automate, and the current documentation of both projects. No independent benchmark is cited in the official project material, so do not assume one is faster or more reliable.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Requirements and installation

Python and package

The current README requires Python 3.8 or newer and gives this installation command:

python -m pip install pyppeteer

The first browser launch can download the project’s Chromium build. The README describes the download as approximately 150 MB; treat that as the project’s estimate, not a current measurement. In a container or CI job, budget disk space, network access, and a writable cache directory.

Using a local Chrome binary

You can configure an executable path or use the documented pyppeteer-install helper. The API reference cautions that compatibility with a non-bundled browser is not guaranteed and says Pyppeteer works best with its bundled Chromium. Pin and test the browser image if reproducibility matters.

Minimal scraper for rendered text

This documentation-based example navigates to a page, reads the rendered body text, prints it, and always closes the browser. The force_expr=True argument tells Pyppeteer that the string is a JavaScript expression; the project notes that expression-versus-function detection can otherwise be ambiguous.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
import asyncio
from pyppeteer import launch

async def main():
    browser = await launch()
    try:
        page = await browser.newPage()
        await page.goto("https://example.com")
        text = await page.evaluate("document.body.innerText", force_expr=True)
        print(text)
    finally:
        await browser.close()

asyncio.run(main())

The README’s examples use an event-loop runner; asyncio.run() is a modern illustrative wrapper. Check it against the Python and Pyppeteer versions in your environment, especially if your application already owns an event loop.

Wait for JavaScript content before extracting it

page.goto() finishing does not prove that an application has finished rendering. A single fixed sleep is brittle: fast pages waste time, while slow pages still fail. Wait for a stable condition that represents the content you actually need.

Wait for a selector

await page.goto("https://example.com/products")
await page.waitForSelector("article.product", {"visible": True})
items = await page.JJ("article.product")

In Python, selector helpers include querySelector(), querySelectorAll(), and xpath(), with the shorthands J(), JJ(), and Jx(). The exact wait options and method signatures are version-sensitive; verify them against the API reference shipped with your installed version. The snippet above illustrates the condition concept; use the selector and option spelling accepted by your release.

Extract a small structured payload

Prefer a focused extraction over dumping the entire HTML document. Returning JSON-compatible values also reduces post-processing:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
data = await page.evaluate("""() => Array.from(document.querySelectorAll('article.product')).map(el => ({
  name: el.querySelector('h2')?.innerText.trim() || null,
  price: el.querySelector('.price')?.innerText.trim() || null,
  url: el.querySelector('a')?.href || null
}))""")

If Pyppeteer misclassifies an expression string, add force_expr=True where the API supports it. For a function-style evaluation, pass a callable or function expression in the form documented for your version.

A production-shaped scraper

This pattern adds a viewport, navigation timeout, an explicit readiness selector, and cleanup. Selectors and timeout values are site-specific; the example does not imply a universal delay or reliability rate.

import asyncio
from pyppeteer import launch

URL = "https://example.com/catalog"

async def scrape():
    browser = await launch({
        "headless": True,
        "args": ["--no-sandbox"]  # use only when your container policy requires it
    })
    try:
        page = await browser.newPage()
        await page.setViewport({"width": 1365, "height": 900})
        page.setDefaultNavigationTimeout(45_000)
        response = await page.goto(URL, {"waitUntil": "networkidle2"})
        if response is None:
            raise RuntimeError("Navigation returned no response")
        await page.waitForSelector("main article", {"visible": True})
        rows = await page.evaluate("""() => Array.from(document.querySelectorAll('main article')).map(el => ({
            title: el.querySelector('h2')?.innerText.trim() || '',
            text: el.innerText.trim()
        }))""", force_expr=True)
        return rows
    finally:
        await browser.close()

if __name__ == "__main__":
    print(asyncio.run(scrape()))

Use the smallest readiness condition that is meaningful for the target. A network-idle condition can remain open on pages with analytics, streaming, or long polling, while a selector can appear before secondary fields are populated. When necessary, wait for a selector and then validate that the extracted fields are non-empty.

Selectors, attributes, and page actions

Text and attributes

Use DOM APIs inside evaluate() for values such as innerText, textContent, href, and data-* attributes. Normalize whitespace and preserve the source URL with each record so later users can audit what was collected.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Clicks, scrolling, and lazy content

Some pages render more results only after a click or scroll. Locate the control, click it, then wait for a new selector or a measurable change in the result count. Avoid unbounded scrolling loops; set a maximum number of rounds and stop when no new records appear.

Screenshots and PDFs for debugging

The README demonstrates page.screenshot(). Capture a diagnostic image when a selector is missing or a login flow behaves differently in CI. Screenshots document what the browser saw; they do not authorize collection of the underlying data.

Failures and recovery

Symptom Likely cause Practical fix
Chromium is missing or launch fails First-run download was blocked, cache is not writable, or the executable path is wrong. Run pyppeteer-install in an environment with network access, persist the cache, or configure a tested executable path. Remember that non-bundled browser compatibility is not guaranteed.
Navigation timeout The page is slow, keeps connections open, or the URL redirects. Set a deliberate timeout, choose an appropriate waitUntil condition, and log the final URL. Do not solve every timeout by using an unlimited wait.
Selector timeout The selector changed, content is behind a consent dialog, or the page returned an error shell. Save the URL and a screenshot, inspect the rendered DOM, verify the selector, and handle an error state explicitly.
Empty text despite a visible page Extraction ran before hydration, selected a hidden template, or targeted an iframe. Wait for the content selector, select the visible element, and inspect frames when the site actually embeds the data there.
evaluate() raises a parsing error Pyppeteer interpreted an expression as a function, or vice versa. Use the documented function form or add force_expr=True for an expression string.
Works locally but not in CI Different Chromium builds, missing libraries, sandbox policy, fonts, or viewport settings. Use a pinned image, install required system dependencies, record browser and Python versions, and reproduce the same headless settings.

Responsible collection and legal boundaries

Browser automation shows what a page renders; it does not grant permission to collect, store, or republish the data. Check the site’s terms, robots and access instructions, applicable law, and any contract governing your account. Prefer an official API or export when one exists. Limit request frequency, cache results where appropriate, identify your application honestly, and avoid personal or restricted data without authorization. Do not treat bypassing CAPTCHAs, bot checks, paywalls, or access controls as a normal scraping step.

Pyppeteer or Playwright for Python?

Decision factor Pyppeteer Playwright for Python
Project status The Pyppeteer README calls the repository unmaintained. The Pyppeteer README explicitly suggests Playwright for Python; current comparative details are not established here.
Existing code Keeping it may avoid a rewrite for a stable internal job. Migration effort depends on your selectors, fixtures, browser management, and integrations.
Browser compatibility Best compatibility is expected with bundled Chromium; external executable support is not guaranteed. Assess the browser versions and features your project requires.
Documentation The API reference lists version 0.0.25 and is legacy documentation. Review the current project documentation before choosing.

For a new production scraper, make maintenance and required browser support explicit acceptance criteria rather than selecting solely because a short example works.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup

If your goal is a clean image or PDF rather than DOM data, ScreenshotNeo provides a website screenshot API and MCP server. It accepts consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be turned off. Only clean shots are billed: bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and the response identifies the result with X-Page-Verdict and X-Billed headers.

One GET request returns PNG, JPEG, WebP, or PDF:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

See the complete parameter list and response details in the ScreenshotNeo documentation. Options include full-page capture with lazy images loaded, CSS-selector element capture, dark mode, 12 device presets or custom viewports, retina scale, PDF paper settings and page ranges, custom CSS and JavaScript, clicks, selector or network-idle waits, request blocking, headers, cookies, user agents, authorization, timezone, geolocation, transparent backgrounds, resizing, chosen cache TTLs, signed image links, asynchronous jobs with signed webhooks, bulk capture of 100 URLs per call, usage reporting, and an OpenAPI specification. An MCP server exposes take_screenshot, get_page_info, and capture_pdf to Claude, Cursor, and other MCP clients.

The Free plan includes 1,000 shots per month with no card. Paid plans start at $5 for 3,000 shots; yearly billing gives two months free, and every feature is on every plan. Create a free ScreenshotNeo account to try it.

Operational checklist

  • Confirm Python 3.8 or newer and decide whether Pyppeteer’s maintenance status is acceptable.
  • Provision Chromium storage and network access, or test a specific executable path.
  • Choose a stable readiness selector and extract only the fields you need.
  • Set bounded navigation and selector timeouts and close the browser in finally.
  • Log the final URL, status, browser version, and extraction counts without storing unnecessary personal data.
  • Review site policies and use an official API when available.
  • For images or PDFs, consider ScreenshotNeo instead of maintaining a browser runtime.

Frequently Asked Questions

Can Pyppeteer connect to an already running browser?

The legacy API reference documents connecting through a WebSocket endpoint; verify the endpoint and launch options against the version installed in your environment.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Does Pyppeteer scrape data hidden behind a login?

Only if you are authorized and provide a permitted session, such as approved cookies or credentials. Authentication does not override the site’s rules or your legal obligations.

What should I store for reproducibility?

Record the target URL, extraction version, selector definitions, Python and browser versions, timestamp, and a bounded error log; avoid retaining sensitive page content unless required.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.