Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

aiohttp fetches a URL; it does not render HTML into a PDF. A reliable Python pipeline uses aiohttp for asynchronous retrieval, then passes the response to a PDF engine. Use WeasyPrint when the page is ordinary, server-rendered HTML/CSS. Use Playwright when JavaScript, browser fonts, client-side data, or browser layout are required. The examples below include redirects, reusable sessions, streaming, size limits, authentication considerations, and troubleshooting.

What aiohttp does—and what it does not do

aiohttp is an asynchronous HTTP client. It can download HTML, follow redirects, send headers and cookies, and reuse connections. It cannot interpret CSS layout, execute JavaScript, or create a PDF by itself. PDF generation is a second stage performed by a renderer such as WeasyPrint or Playwright’s Chromium engine.

That separation matters. A page whose content is present in the initial HTML can usually be rendered with WeasyPrint. A page that fills its content after JavaScript runs needs a real browser. If the URL already returns a PDF, save those bytes directly rather than converting them again.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Install the components

Create an isolated environment and install the HTTP client plus the renderer you choose:

python -m venv .venv
source .venv/bin/activate       # Windows: .venvScriptsactivate
pip install aiohttp weasyprint
# For JavaScript-heavy pages:
pip install playwright
playwright install chromium

WeasyPrint also depends on native libraries on some operating systems; follow its platform installation instructions if the import fails. Playwright requires a browser binary in addition to its Python package.

Basic URL-to-PDF pipeline with aiohttp and WeasyPrint

This reusable coroutine downloads the page, checks the HTTP status, records the final URL after redirects, and renders the HTML. Passing that final URL as base_url lets relative stylesheets, images, fonts, and links resolve correctly.

import asyncio
from pathlib import Path

import aiohttp
from weasyprint import HTML


async def url_to_pdf(url: str, output: str = "out.pdf") -> None:
    timeout = aiohttp.ClientTimeout(total=60)
    async with aiohttp.ClientSession(timeout=timeout) as session:
        async with session.get(url, allow_redirects=True) as response:
            response.raise_for_status()
            html = await response.text()
            final_url = str(response.url)

    HTML(string=html, base_url=final_url).write_pdf(output)


if __name__ == "__main__":
    asyncio.run(url_to_pdf("https://example.com/", "example.pdf"))

ClientSession is the normal aiohttp interface and should generally be reused instead of creating one session per request. The async with blocks release the response and session cleanly, even when an exception occurs. raise_for_status() stops a 404 or 500 page from being mistaken for valid input.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Keep the HTML encoding explicit when needed

await response.text() decodes according to the response headers and detected encoding. If a site declares a broken charset, read bytes and decode with a known encoding, or let WeasyPrint consume a saved HTML file. Do not silently replace undecodable bytes if the document’s text must be preserved.

Stream large responses instead of loading them all

The convenient text(), read(), and json() methods load the complete response into memory. For large pages, impose a limit and write chunks to a temporary file:

import asyncio
import tempfile
from pathlib import Path

import aiohttp
from weasyprint import HTML


async def streamed_url_to_pdf(url: str, output: str = "out.pdf",
                              max_bytes: int = 25 * 1024 * 1024) -> None:
    timeout = aiohttp.ClientTimeout(total=90)
    async with aiohttp.ClientSession(timeout=timeout) as session:
        async with session.get(url, allow_redirects=True) as response:
            response.raise_for_status()
            length = response.headers.get("Content-Length")
            if length and int(length) > max_bytes:
                raise ValueError("response exceeds the configured size limit")

            with tempfile.NamedTemporaryFile(suffix=".html", delete=False) as tmp:
                html_path = Path(tmp.name)
                total = 0
                async for chunk in response.content.iter_chunked(64 * 1024):
                    total += len(chunk)
                    if total > max_bytes:
                        raise ValueError("response exceeds the configured size limit")
                    tmp.write(chunk)
            final_url = str(response.url)

    try:
        HTML(filename=str(html_path), base_url=final_url).write_pdf(output)
    finally:
        html_path.unlink(missing_ok=True)


asyncio.run(streamed_url_to_pdf("https://example.com/"))

The limit is an application safeguard, not a property of aiohttp. Choose it for your workload and reject unexpectedly huge documents before rendering.

Choose the right PDF renderer

WeasyPrint: server-rendered HTML and CSS

WeasyPrint exposes HTML(...).write_pdf(...) and is a good fit when the initial response already contains the page’s content. It supports print-oriented CSS, but it is not a JavaScript browser. Unsupported CSS, web fonts, cross-origin restrictions, or missing assets can change the appearance.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use the redirected URL as base_url. Without it, relative paths such as /styles/site.css and images/logo.svg have no dependable origin and may disappear from the PDF.

Playwright: JavaScript and browser layout

Playwright’s page.pdf() generates a PDF using print CSS media. Use it for single-page applications, client-side data loading, browser fonts, animations that must settle, or layout that depends on Chromium behavior.

import asyncio
from playwright.async_api import async_playwright


async def browser_url_to_pdf(url: str, output: str = "out.pdf") -> None:
    async with async_playwright() as playwright:
        browser = await playwright.chromium.launch()
        page = await browser.new_page()
        await page.goto(url, wait_until="networkidle")
        await page.pdf(path=output, print_background=True)
        await browser.close()


asyncio.run(browser_url_to_pdf("https://example.com/", "example.pdf"))

networkidle is useful for pages that load data, but some sites keep analytics or streaming connections open indefinitely. In that case, wait for a meaningful selector or use a bounded timeout rather than waiting forever.

Use both stages when it helps

An aiohttp request can cheaply check status, inspect headers, authenticate, or determine whether the response is already a PDF. For a JavaScript page, let Playwright perform the actual navigation and rendering; do not expect HTML downloaded by aiohttp to contain content that JavaScript would have created in the browser.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Redirects, existing PDFs, and content checks

Follow redirects deliberately. The examples allow them and preserve response.url. If your security policy permits only one host, validate every redirect destination before continuing. Check the Content-Type header:

content_type = response.headers.get("Content-Type", "").lower()
if "application/pdf" in content_type:
    data = await response.read()
    Path("out.pdf").write_bytes(data)
    return

Some servers omit or mislabel the header, so a production service may also inspect the first bytes for the PDF signature (%PDF-) while keeping its size limit.

Cookies, authentication, and protected pages

Send credentials in the request that fetches the HTML:

headers = {"Authorization": "Bearer YOUR_TOKEN"}
cookies = {"session": "YOUR_SESSION_COOKIE"}
async with session.get(url, headers=headers, cookies=cookies) as response:
    response.raise_for_status()
    html = await response.text()
    final_url = str(response.url)

Those credentials are not automatically available to WeasyPrint when it fetches the page’s CSS and images. Its default URL fetcher handles ordinary HTTP and file URLs, but advanced cookies, authentication, and custom headers require a custom URL fetcher—or you must supply authenticated content and assets yourself. A browser context is often simpler for sites whose access depends on JavaScript or session state.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Never log bearer tokens or session cookies. Treat the input URL as untrusted: restrict schemes to HTTP(S), apply an allowlist or egress policy where appropriate, and block access to internal metadata and private network ranges in production.

Page size, margins, and print appearance

For WeasyPrint, put print rules in the document or an additional stylesheet:

<style>
@page { size: A4; margin: 18mm; }
@media print { .screen-only { display: none; } }
</style>

For Playwright, choose options such as format="A4", landscape=True, margin, and print_background=True in page.pdf(). Remember that Playwright uses print media; if the site has only screen styles, add an explicit print stylesheet or call page.emulate_media(media="screen") when that is the intended output.

Batch conversion and performance

Reuse one aiohttp session

For many URLs, create one ClientSession and schedule bounded tasks. Connection pooling and keep-alives avoid repeating TCP and TLS setup. Use a semaphore so a batch does not overwhelm the remote site or your renderer:

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
sem = asyncio.Semaphore(8)

async def limited(url, session):
    async with sem:
        async with session.get(url, allow_redirects=True) as response:
            response.raise_for_status()
            return await response.text(), str(response.url)

Rendering is CPU- and memory-intensive, especially in Chromium. Limit concurrent browser pages, close contexts and browsers in finally blocks, and use per-request timeouts. There is no universal throughput number: page size, fonts, JavaScript, network latency, and renderer version dominate.

Cache only when the content may be reused

HTTP caching, an application cache, or a saved HTML snapshot can reduce repeated downloads. Include authentication state, query parameters, and relevant headers in the cache key. Do not cache private pages where another user could retrieve the result.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Troubleshooting common failures

“The PDF is blank or missing content”

  • The page may be JavaScript-dependent. Use Playwright and wait for a content selector.
  • Images or styles may be relative. Pass the final redirected URL as base_url.
  • The server may have returned a login, bot-check, or error page. Inspect status, final URL, and a short HTML sample before rendering.

“CSS, fonts, or images do not load”

  • Check that resource URLs are reachable from the renderer and use HTTPS where required.
  • Confirm the document's base URL and examine browser or WeasyPrint resource logs.
  • For authenticated assets, provide a custom fetcher or use a browser context carrying the session.

“TimeoutError” or a request that never finishes

  • Set a total aiohttp timeout and a navigation timeout in Playwright.
  • Replace an unbounded networkidle wait with wait_for_selector() for a stable element.
  • Retry transient network failures with capped exponential backoff, but do not blindly retry authentication failures or non-idempotent actions.

“ClientResponseError: 403 or 429”

Respect the site's access policy and rate limits. Use the required authentication, identify your client honestly, reduce concurrency, and honor Retry-After where supplied. A PDF renderer cannot bypass authorization or anti-bot controls.

“WeasyPrint import or library error”

Install the native dependencies for your operating system, then verify the package inside the same virtual environment that runs your script. Pin compatible versions in deployment and test fonts in the target image.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

“Playwright browser executable is missing”

Run playwright install chromium during image or environment setup. In containers, install the documented system dependencies as well.

Or skip the browser setup

ScreenshotNeo provides a hosted screenshot and PDF API when you do not want to operate Chromium or WeasyPrint. One GET request can return a PDF; cookie and consent banners are accepted and 60+ known consent platforms, newsletter popups, and chat widgets are removed before capture. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and response headers identify the page verdict and billing result. Its MCP server exposes take_screenshot, get_page_info, and capture_pdf to Claude, Cursor, and other MCP clients.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

For PDF output, add the API's PDF option to the request; the complete parameter reference is in the ScreenshotNeo documentation.

Python

import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)

Node.js

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`HTTP ${res.status}`);
require('fs').writeFileSync('shot.webp', Buffer.from(await res.arrayBuffer()));

ScreenshotNeo includes full-page capture, lazy-image loading, custom CSS and JavaScript, waits, headers and cookies, device and viewport controls, PDF page settings, signed links, asynchronous webhooks, bulk capture, caching, and a usage API. Every feature is on every plan: 1,000 shots per month are free with no card; paid plans start at $5 for 3,000 shots. Start with the free ScreenshotNeo account.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Production checklist

  • Reuse a session for batches and close it deterministically.
  • Check status, content type, redirects, and response size before rendering.
  • Use WeasyPrint for static HTML/CSS and Playwright for JavaScript or browser fidelity.
  • Set network, navigation, and application-level timeouts.
  • Carry cookies, authorization, and custom headers into every asset fetch.
  • Restrict untrusted URLs and prevent internal-network access.
  • Keep concurrency bounded and clean up temporary files and browser processes.
  • Validate PDFs with representative pages, fonts, images, and authenticated states.

Frequently Asked Questions

Can aiohttp convert HTML directly to PDF?

No. aiohttp downloads the response; a renderer such as WeasyPrint or Playwright must create the PDF.

Which renderer should I use for a React or Vue page?

Use Playwright when the page's content or layout appears only after JavaScript runs. Use WeasyPrint for HTML and CSS already present in the response.

Why is my PDF different from the browser view?

WeasyPrint is not Chromium and may differ in CSS, fonts, and asset handling. Compare with Playwright when browser layout fidelity is required.

How do I convert many URLs safely?

Reuse one ClientSession, bound concurrent tasks with a semaphore, set size and timeout limits, and close renderer resources after each job.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.