Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

For a webpage that depends on JavaScript, the most reliable general-purpose Python approach is to open it in Chromium with Playwright, wait for the content you need, then save it with page.pdf(). Playwright prints using print CSS by default; you can choose paper size, margins, orientation, and whether to include backgrounds. This guide covers a runnable starting point, how to handle pages that load slowly or require a login, when to consider other renderers, and how to protect a server that accepts URLs.

Convert a webpage URL to PDF with Playwright

Playwright controls a real browser engine, so it can run page JavaScript and supports browser-context behavior that a direct HTML-to-PDF renderer may not. Its Python API documents page.pdf() as generating a PDF with print CSS media. PDF generation is a Chromium-oriented workflow; do not assume the same method works identically with Playwright’s other browser engines. See the Playwright page API.

Install the Python package and its browser binaries. The browser download is a separate step from installing the package, as described in the Playwright browser documentation.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
python -m pip install playwright
playwright install chromium

Save this as webpage_to_pdf.py, replacing the example URL with the page to capture:

from playwright.sync_api import sync_playwright

url = "https://example.com"

with sync_playwright() as p:
    browser = p.chromium.launch()
    page = browser.new_page()
    response = page.goto(url, wait_until="networkidle", timeout=60_000)

    if response is None:
        raise RuntimeError("Navigation did not return an HTTP response")
    if not response.ok:
        raise RuntimeError(f"Page returned HTTP {response.status}: {url}")

    page.pdf(
        path="page.pdf",
        format="A4",
        print_background=True,
        margin={"top": "15mm", "right": "12mm", "bottom": "15mm", "left": "12mm"},
    )
    browser.close()

Run it with python webpage_to_pdf.py. The output is written to page.pdf in the current working directory. The response checks help distinguish an HTTP error from a successful navigation, but a page can still return a successful status while showing an application error or incomplete content.

networkidle is a navigation wait condition, not proof that every application has finished rendering. Some sites poll continuously, load images lazily, or fetch data after navigation. For those pages, wait for a meaningful selector or an application-specific readiness signal before calling page.pdf().

Wait for the content the PDF actually needs

If the page has a known element that appears when its main content is ready, navigate to the page and wait for that element:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
page.goto(url, wait_until="domcontentloaded", timeout=60_000)
page.locator("main article").wait_for(state="visible", timeout=30_000)
page.pdf(path="page.pdf", format="A4", print_background=True)

Change main article to a selector that fits the target page. If content is rendered only after a user action, perform that action first using Playwright locators, then wait for the resulting content. A fixed delay can help with a known short animation, but it is usually less dependable than waiting for a specific element:

page.wait_for_timeout(2_000)

Use a wait condition suited to the site rather than increasing timeouts blindly. If a page never reaches network idle because it keeps a connection open, waiting for networkidle may time out even though the visible content is ready.

Choose print layout, paper, and page breaks

page.pdf() uses print media by default. That means print-specific CSS such as @media print can change visibility, colors, and layout compared with the browser window. The available controls and their behavior are documented in the Playwright PDF API.

Setting What it controls Example
format Standard paper size format="A4" or format="Letter"
landscape Page orientation landscape=True
margin Space around each printed page margin={"top": "15mm", "bottom": "15mm"}
print_background Whether to include background graphics and colors print_background=True
prefer_css_page_size Whether CSS @page dimensions take precedence over the format option prefer_css_page_size=True
page_ranges Which PDF pages to include page_ranges="1-3"
scale Scale of the page rendering scale=0.9

For example, to make CSS page sizing control the output and print only the first two pages:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
page.pdf(
    path="page.pdf",
    print_background=True,
    prefer_css_page_size=True,
    page_ranges="1-2",
)

Background printing is opt-in. Without it, colored panels and background images may be missing from the PDF. If a document defines its own page dimensions with CSS, decide whether to honor them using prefer_css_page_size=True or impose a standard size with format. Margins and page size influence pagination, so inspect the resulting PDF when layout precision matters.

To render screen styling rather than print styling, emulate screen media before generating the PDF:

page.emulate_media(media="screen")
page.pdf(path="page.pdf", format="A4", print_background=True)

PDF colors are adjusted for printing by default. The Playwright documentation identifies the CSS property -webkit-print-color-adjust as a way for a page to request exact colors. Whether you can change the page’s styles depends on whether you control its CSS or can inject a stylesheet.

Headers and footers

Playwright supports header and footer templates through PDF options. They have important limits: scripts inside the templates are not evaluated, and the page’s styles are not visible inside them. Do not depend on page JavaScript or page CSS to populate or style template content. Check the current API documentation for the option names and supported template classes for your installed version.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use authentication, cookies, and browser state

When a page requires a login, use a browser context with the required state rather than expecting an unauthenticated request to produce the same content. Playwright lets you create a context, add cookies, or load saved authentication state; consult its documentation for the authentication flow that matches your application. Keep credentials and authentication-state files private, and do not commit them to source control.

For a simple cookie-based case, create a context and add a cookie before opening the page. The cookie’s domain, path, security, and expiration must match the target site’s requirements:

context = browser.new_context()
context.add_cookies([{
    "name": "session",
    "value": "YOUR_SESSION_VALUE",
    "domain": "example.com",
    "path": "/",
    "httpOnly": True,
    "secure": True,
}])
page = context.new_page()
page.goto(url, wait_until="domcontentloaded")

This is only an example of cookie injection; it is not a substitute for the site’s intended login flow. A session cookie may expire or be bound to additional browser state. If the PDF contains a login screen, first confirm that the browser context is authenticated and that the target content is accessible in that same context.

When to use WeasyPrint or Selenium instead

The right renderer depends on what the page needs. The official documentation describes capabilities, not comparative speed or quality benchmarks, so there is no evidence-based universal performance winner.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Option Good fit Important limitation
Playwright Pages that need browser-side JavaScript, browser interactions, or browser-context cookies and authentication Requires browser binaries; page.pdf() is a Chromium-oriented workflow
WeasyPrint HTML and CSS pages that fit its rendering and resource-fetching model Do not assume browser-equivalent JavaScript execution; its default URL fetcher does not provide advanced cookie or authentication support
Selenium Projects that already automate a browser with Selenium WebDriver Its documented PDF flow returns encoded PDF data that must be decoded and saved

WeasyPrint’s official documentation shows a direct URL-to-PDF pattern: HTML('https://weasyprint.org/').write_pdf('/tmp/weasyprint-website.pdf'). See WeasyPrint documentation. Selenium’s WebDriver documentation describes printing a page to PDF and returning encoded data; see Selenium print page. Choose by checking JavaScript needs, interaction, authentication, print CSS behavior, and deployment dependencies—not by assuming one tool is always faster.

Performance, reliability, and operating costs

A browser-based PDF capture performs navigation, page rendering, and PDF generation, so its duration and resource use depend on the page and browser environment. The cited documentation does not establish comparative performance figures. For batch work, reuse browser processes where appropriate instead of launching a new browser for every URL, but isolate browser contexts and handle each page’s failure so one bad URL does not silently corrupt the rest of a job.

  • Set navigation and selector timeouts appropriate to the pages you control; report timeout failures rather than saving an empty or partial file as if it succeeded.
  • Check HTTP response status and, when important, verify expected page content before printing.
  • Close pages, contexts, and browsers reliably, including on exceptions. A context manager or try/finally cleanup pattern helps keep long-running workers from leaking browser processes.
  • Pin your Playwright version in deployment and install the corresponding browser binaries. New PDF options may not exist in older versions; check the documentation matching the installed release.
  • Budget for browser binaries, memory, CPU, and timeouts in the environment where the script runs. The sources document how to install and use the browser, but do not provide a universal cost or capacity estimate.

Protect services that convert user-supplied URLs

If a server accepts a URL from a user and fetches it to make a PDF, it can become a server-side request forgery (SSRF) surface. An attacker may try to make the service contact internal or otherwise unintended network resources. OWASP warns that validating complete URLs is difficult and that parsers can disagree; consult the OWASP SSRF Prevention Cheat Sheet.

  • For constrained workflows, allowlist permitted destination hosts instead of accepting arbitrary URLs.
  • Apply network-level restrictions as a second line of defense; do not rely on URL string checks alone.
  • Account for redirects, which can bypass simplistic checks that validate only the initial URL. Disable redirects where appropriate or validate each destination.
  • Remember that a browser loads subresources as well as the initial page. Treat the renderer as a network client with an explicit policy, and restrict access to internal services and local files.

These protections belong in the application and its network environment. Using Playwright does not by itself make an untrusted URL safe to fetch.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Troubleshoot common failures

Playwright reports that no browser executable exists

The Python package and browser binaries are installed separately. Run playwright install chromium in the same environment as the script. In deployment, ensure the install step runs in the image or environment that executes the code.

Navigation times out at network idle

Some pages keep network activity open or continually poll. Use a less restrictive navigation condition such as domcontentloaded, then wait for a specific content selector or application readiness signal. Increasing the timeout alone will not help if the page never becomes idle.

The PDF is blank or misses content

The page may have rendered its shell before loading the content, require a user action, or load images lazily. Wait for a content selector, perform the required interaction, and verify the expected text or element before printing. For content that appears only after scrolling, scroll through the relevant page area before capture and confirm the site’s behavior.

Colors or background graphics are missing

Set print_background=True for background graphics. If the page’s print stylesheet changes its colors or layout, use page.emulate_media(media="screen") before printing when screen styling is the desired result.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Pages split awkwardly or use the wrong paper size

Review print CSS, paper format, margins, and prefer_css_page_size. The site’s @page rules can affect dimensions and pagination. Try a suitable standard format, or allow CSS page size to take precedence when the page defines the intended size.

The PDF contains a sign-in page instead of the requested content

The page was not authenticated in the browser context that performed the navigation, or the session expired. Establish the appropriate login or cookie state in that context and confirm the target content is visible before generating the PDF.

The page loads but returns an error response

Check the HTTP status before printing, as in the example script. A non-success response may indicate an invalid URL, access restriction, or server error. A successful status is not enough to prove the correct content loaded, so verify a known page element when the output matters.

Or skip the browser setup

If you want a hosted screenshot or PDF capture rather than installing and managing a browser, ScreenshotNeo offers a website screenshot API and MCP server. Its PDF endpoint can capture a URL in one request; see the ScreenshotNeo documentation.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o page.pdf -d format=pdf

Replace the URL with the page you need and use your API key. ScreenshotNeo can accept cookie and consent banners before capture and remove known consent platforms, newsletter popups, and chat widgets; each cleanup step can be turned off. Bot checks and CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing status. An MCP server provides take_screenshot, get_page_info, and capture_pdf tools for AI agents. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 shots.

Sign up free for 1,000 screenshots a month, with no card required.

Frequently Asked Questions

Can Playwright create a PDF with Firefox or WebKit?

The documented page.pdf() workflow is Chromium-oriented; do not assume equivalent support across the other engines.

Does networkidle guarantee the page is ready to print?

No. It is a navigation wait condition. For late-rendered content, wait for a selector or application-specific readiness condition.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Can WeasyPrint run JavaScript from the webpage?

The cited WeasyPrint documentation describes HTML/CSS rendering and does not establish browser-equivalent JavaScript execution.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.