October DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsPC HealthRecommendedCrashes, freezes, slowdowns? Check your PC nowSpot repairable issues before they interrupt work.Check PCOctober DealsAmazon USDeal season is back - check today's better picksAmazon US: current deals, useful picks and tech finds.See Picks×
Skip to content
aiohttp

Convert Raw HTML to PDF in Python with aiohttp: WeasyPrint and Playwright

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use aiohttp to fetch the HTML asynchronously, then render it with a PDF engine: WeasyPrint for static, print-oriented HTML and CSS, or Playwright when the page depends on JavaScript or browser layout. The HTTP fetch and PDF rendering are separate jobs; aiohttp retrieves the document but does not turn it into a PDF.

How the conversion works

The pipeline has two stages: request a page with an aiohttp.ClientSession, then hand its HTML to a renderer. For a normal-sized response, check the HTTP status and read the body as text. Pass the original page URL as the renderer’s base URL so relative stylesheets, images, and fonts can be resolved.

  1. Fetch the URL with an explicit timeout and an HTTP status check.
  2. Read and decode the HTML, validating response type and size where appropriate.
  3. Render it to PDF with WeasyPrint or, if browser execution is needed, with Playwright.
  4. Run remote or user-supplied input under suitable network, time, and resource limits.

Static HTML: aiohttp and WeasyPrint

This complete example fetches a page and writes a PDF. Install aiohttp and weasyprint in an environment supported by your platform, then save as html_to_pdf.py and run it with Python.

import asyncio
import aiohttp
from weasyprint import HTML

async def html_to_pdf(url: str, output_path: str) -> None:
    timeout = aiohttp.ClientTimeout(total=30)
    async with aiohttp.ClientSession(timeout=timeout) as session:
        async with session.get(url) as response:
            response.raise_for_status()
            html = await response.text()

    HTML(string=html, base_url=url).write_pdf(output_path)

asyncio.run(html_to_pdf("https://example.com", "out.pdf"))

raise_for_status() prevents an error page such as an HTTP 404 response from silently being treated as the intended document. HTML(string=html, base_url=url) tells WeasyPrint both what markup to render and where relative resource URLs originate. WeasyPrint supports rendering an HTML string with HTML.write_pdf(); its default URL fetcher can retrieve HTTP and file resources. Authentication and other advanced resource-fetching requirements may call for a custom URL fetcher. See the WeasyPrint documentation.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use an existing HTML string

If your application already has the markup, skip the HTTP request and render the string directly. Still supply a meaningful base_url if the document contains relative URLs; without one, relative resource references may not resolve as intended.

from weasyprint import HTML

html = "<h1>Report</h1><p>Generated by Python.</p>"
HTML(string=html, base_url="https://example.com/").write_pdf("report.pdf")

Large responses: stream and cap the body

response.text(), response.read(), and response.json() load the complete response body into memory. For a potentially large document, read chunks and enforce a maximum size before decoding. The renderer will still need the HTML to render, so streaming the download does not make the entire PDF pipeline constant-memory; it lets you reject an oversized response before retaining an unlimited body.

async def read_limited(response: aiohttp.ClientResponse, max_bytes: int) -> bytes:
    chunks = []
    size = 0
    async for chunk in response.content.iter_chunked(64 * 1024):
        size += len(chunk)
        if size > max_bytes:
            raise ValueError(f"HTML response exceeds {max_bytes} bytes")
        chunks.append(chunk)
    return b"".join(chunks)

async def fetch_limited(session: aiohttp.ClientSession, url: str) -> tuple[str, str]:
    async with session.get(url) as response:
        response.raise_for_status()
        content_type = response.headers.get("Content-Type", "")
        if "text/html" not in content_type.lower():
            raise ValueError(f"Expected HTML; got {content_type or 'unknown content type'}")
        raw = await read_limited(response, max_bytes=5 * 1024 * 1024)
        encoding = response.charset or "utf-8"
        return raw.decode(encoding, errors="replace"), str(response.url)

This helper returns the final response URL after redirects, which can be a more accurate base URL than the original request if the document uses relative links. Choose a limit suitable for your application; the 5 MiB value here is an example policy, not an aiohttp or WeasyPrint requirement. For stricter decoding, replace errors="replace" with strict decoding and handle a decoding error explicitly.

JavaScript-driven pages: use Playwright

WeasyPrint is not a browser JavaScript runtime. If the HTML only becomes complete after scripts run, or the output must reflect browser layout and print behavior, load the page in Playwright and generate the PDF from the rendered page. The example below is asynchronous and sets a navigation timeout.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
import asyncio
from playwright.async_api import async_playwright

async def page_to_pdf(url: str, output_path: str) -> None:
    async with async_playwright() as p:
        browser = await p.chromium.launch()
        try:
            page = await browser.new_page()
            page.set_default_navigation_timeout(30_000)
            await page.goto(url, wait_until="networkidle")
            await page.pdf(path=output_path, format="A4", print_background=True)
        finally:
            await browser.close()

asyncio.run(page_to_pdf("https://example.com", "out.pdf"))

Install Playwright and its supported browser binaries following the Playwright Python documentation. page.pdf() uses print CSS media by default. If the desired PDF should use screen styles instead, call await page.emulate_media(media="screen") before page.pdf(). For content that appears after a particular application event, waiting for networkidle may not be the right readiness test; wait for a selector that represents the actual content instead.

WeasyPrint or Playwright?

Need WeasyPrint Playwright
Run page JavaScript Not a browser JavaScript runtime; use for already-rendered HTML and CSS. Uses a browser page, so suitable when scripts must execute.
Print-oriented HTML and CSS Often the simpler path for static, print-oriented documents. Can generate browser PDFs; uses print CSS media by default.
Browser layout or screen media Not the browser-rendering option. Use when browser layout matters; emulate screen media before PDF generation if needed.
Remote resources and authentication Default fetcher can retrieve HTTP and file resources; advanced cookies or authentication require a custom URL fetcher. Use browser-context request and page controls as appropriate; consult the Playwright documentation for the specific control.
Startup and memory considerations Choose based on your deployment and document workload; no comparative benchmark is established here. Choose based on your deployment and document workload; no comparative benchmark is established here.

Choose by required behavior rather than a presumed speed ranking: use WeasyPrint for a fetched, static document that needs PDF output, and Playwright if scripts or browser-specific rendering are part of the result. There is no performance benchmark here to establish that either option is universally faster.

Reliability, resource loading, and security

Reuse sessions and set time limits

A reusable ClientSession is the normal aiohttp pattern, especially when fetching multiple pages. Set both connect and total timeouts appropriate to the source and workload. Rendering can take additional time beyond the HTTP request, so production code should also run the renderer in a controlled worker with an overall job deadline.

Validate what you fetched

  • Check the status before rendering and reject unexpected content types.
  • Decide whether redirects are allowed and how many; validate the final URL as well as the requested URL when inputs are untrusted.
  • Enforce a response-size limit and handle character encoding deliberately.
  • Log enough context to diagnose failures without logging credentials or sensitive HTML.

Restrict untrusted input and outbound requests

HTML can reference stylesheets, images, and fonts, and those resources can trigger additional network or file access during rendering. If users can submit URLs or markup, apply outbound network restrictions, redirect controls, resource limits, and process isolation. WeasyPrint explicitly warns that untrusted HTML or CSS may create security problems; see its URL fetcher and security guidance. Do not assume that a successful initial aiohttp fetch makes every resource referenced by the document safe.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Troubleshooting common failures

  • HTTP error or a PDF containing an error page: The source may have returned a non-success status. Call raise_for_status() before reading and rendering, and inspect the response URL and status.
  • Relative images or styles are missing: The HTML string lacks the correct document base. Pass the final page URL as base_url to WeasyPrint, especially after redirects.
  • PDF lacks content created by JavaScript: A static HTML renderer cannot execute the page’s scripts. Use Playwright and wait for the relevant content or selector before printing.
  • Unexpected page colors or layout in Playwright: PDF generation uses print media by default. Emulate screen media before printing if screen CSS is required.
  • Login-only resources do not load in WeasyPrint: The default resource fetcher may not carry the authentication context you need. Implement a custom URL fetcher for the needed cookies or authentication, or choose a browser workflow suited to the page.
  • Memory use grows with document size: The response is being buffered, and rendering also consumes resources. Stream and cap the fetch as shown, then reject or isolate documents too large for the service’s limits.
  • Fetch or render hangs: Set network timeouts, constrain resource access, and impose a job deadline around the renderer. A navigation timeout and a PDF-job timeout address different stages.
  • Fonts or images fail inconsistently: Check resource URLs, accessibility from the rendering environment, and any required authentication. A correct base URL helps resolve relative paths but does not grant access to protected resources.

Or skip the browser setup

If you need a screenshot or PDF of a live website rather than a custom Python rendering pipeline, ScreenshotNeo is a website screenshot API and MCP server. One GET request can return a PNG, JPEG, WebP, or PDF. Its clean-shot options accept consent banners and remove more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks and CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers report the page verdict and billing status. AI agents can use its MCP server tools: take_screenshot, get_page_info, and capture_pdf.

For a quick PDF capture, request PDF output from the API using the documented PDF parameters for your needs. The following is the one-call screenshot form, which returns a WebP image as written; see the ScreenshotNeo API documentation for PDF options and other parameters.

curl -G "https://api.screenshotneo.com/v1/shot" 
  -d access_key=YOUR_API_KEY 
  --data-urlencode url=https://stripe.com 
  -o shot.webp

The free plan includes 1,000 shots per month with no card. Paid plans start at $5 for 3,000 shots; yearly billing gives two months free. Every feature is available on every plan. If that fits your workflow, sign up for ScreenshotNeo to get 1,000 free screenshots a month with no card.

Frequently Asked Questions

Does aiohttp convert HTML to PDF by itself?

No. aiohttp fetches HTTP content asynchronously; a separate renderer such as WeasyPrint or Playwright creates the PDF.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Which renderer should I use for a page that runs JavaScript?

Use Playwright when the rendered result depends on page scripts or browser layout. WeasyPrint is for static HTML and CSS rendering.

Can I use these examples for untrusted URLs?

Only with safeguards. Validate URLs and redirects, restrict outbound resources, cap response size, and isolate rendering; untrusted HTML and CSS can be unsafe.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Read next

Recommended PC Tool
Recommended PC Tool
Windows Errors? Fix Them Before They SpreadFree repair scan
Outdated Drivers Are Slowing You DownFree scan - exact matches

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.