Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use WeasyPrint to render HTML and CSS to PDF, pass that PDF to pdf2image when you need PNG or JPEG pages, and use python-docx to create a structured Word document from selected content. These are different jobs: python-docx is not a general, faithful HTML-to-DOCX renderer. If you need a hosted renderer instead of installing browser-like dependencies, an API such as HTML2Image is another option, subject to its current service terms.

Choose the conversion path first

The right pipeline depends on the output you actually need:

Goal Recommended Python approach What it preserves
Printable PDF WeasyPrint: HTML/CSS to PDF Document layout, styles, images, fonts and page breaks that the renderer supports
PNG or JPEG page images WeasyPrint, then pdf2image The rendered PDF pages as raster images; image quality depends on chosen resolution and format
Editable Word file Build a DOCX with python-docx Paragraphs, headings, tables and pictures you explicitly add
Hosted HTML rendering HTML2Image’s documented Python client or HTML-to-PDF API Depends on the vendor’s current renderer, limits and handling terms

There is no neutral benchmark in the available documentation proving that one route is universally fastest or most faithful. Test representative pages containing your real CSS, fonts, images and page-break rules before selecting a production workflow.

Prepare a reproducible Python environment

Create an isolated environment and install only the packages needed by your output:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
python -m venv .venv
# macOS/Linux
source .venv/bin/activate
# Windows PowerShell
.venvScriptsActivate.ps1

pip install weasyprint pdf2image python-docx pillow

WeasyPrint may require operating-system libraries in addition to the Python package. Check its first-steps installation instructions for the operating system and deployment image you use. pdf2image also relies on PDF conversion utilities supplied by the target system; verify those utilities and their paths during deployment rather than assuming a laptop setup will transfer unchanged.

Keep your input deterministic: store the HTML, CSS, images and fonts in a known directory, and pin package versions in your project after testing. A conversion can succeed while silently omitting an external asset, so inspect the output rather than checking only for an exit code.

Convert HTML and CSS to PDF with WeasyPrint

Render a local HTML file

WeasyPrint accepts a filename, URL, readable file object or in-memory string. The simplest local-file program is:

from pathlib import Path
from weasyprint import HTML

source = Path("invoice.html").resolve()
output = Path("build/invoice.pdf")
output.parent.mkdir(parents=True, exist_ok=True)

HTML(filename=str(source)).write_pdf(str(output))
print(f"Wrote {output}")

Use an absolute path (or a correctly configured base URL) so relative references such as css/site.css and images/logo.png can be resolved. For an HTML string, provide a base URL pointing to the directory containing relative assets:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
from weasyprint import HTML

html = """
<!doctype html>
<html>
  <head>
    <meta charset='utf-8'>
    <style>
      @page { size: A4; margin: 18mm; }
      h1 { color: #17324d; }
    </style>
  </head>
  <body><h1>Report</h1><p>Generated from a string.</p></body>
</html>
"""

HTML(string=html, base_url="/absolute/path/to/project").write_pdf("report.pdf")

When you need the bytes in memory—for example, to upload them—omit the destination:

from weasyprint import HTML

pdf_bytes = HTML(string=html, base_url="/absolute/path/to/project").write_pdf()
with open("report.pdf", "wb") as file:
    file.write(pdf_bytes)

Fonts, page rules and external resources

Use CSS @page for paper size, margins, headers or footers, and test page-break behavior with long tables and headings. For custom fonts, WeasyPrint’s documented pattern creates a FontConfiguration and passes it to the HTML/CSS objects:

from weasyprint import HTML, CSS
from weasyprint.text.fonts import FontConfiguration

font_config = FontConfiguration()
html = HTML(filename="invoice.html")
css = CSS(filename="print.css", font_config=font_config)
html.write_pdf("invoice.pdf", stylesheets=[css], font_config=font_config)

The ordinary URL fetcher can retrieve linked stylesheets and images, but cookies and authentication are not supported by default. Protected assets may therefore be missing even though public assets load. A custom URL fetcher can supply the access behavior your application needs; alternatively, download authorized assets yourself and render local copies. Test fonts, images, CSS, redirects and access restrictions with production-like data.

Turn the PDF into PNG or JPEG images

pdf2image converts PDF input to images; it is not an HTML renderer. The reliable multi-page pipeline is therefore HTML → PDF → raster images.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
from pathlib import Path
from weasyprint import HTML
from pdf2image import convert_from_path

pdf_path = Path("build/report.pdf")
images_dir = Path("build/pages")
images_dir.mkdir(parents=True, exist_ok=True)

HTML(filename="report.html").write_pdf(str(pdf_path))
pages = convert_from_path(
    str(pdf_path),
    dpi=150,
    fmt="png",
    output_folder=str(images_dir),
    paths_only=True,
)
for page_number, path in enumerate(pages, start=1):
    print(page_number, path)

Increase dpi for print-quality output and reduce it for thumbnails or large batches. Use fmt="jpeg" for smaller photographic files, or PNG for sharp text and diagrams. Convert only a page range when your pdf2image version and installed PDF utility support the corresponding options; confirm those options against the current project documentation.

For predictable filenames and post-processing, render pages in memory and save them yourself:

from pdf2image import convert_from_path

for number, image in enumerate(convert_from_path("report.pdf", dpi=144), start=1):
    image.save(f"page-{number:03d}.webp", "WEBP", quality=90)

Large PDFs can consume substantial memory because each raster page is a bitmap. Process pages in smaller ranges, write them to disk, and release image objects when handling long documents.

Create a Word document with python-docx

python-docx is designed to create and update .docx files. It can add paragraphs, headings, tables and pictures, which makes it useful when you control the document structure:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
from docx import Document
from docx.shared import Inches

source = Document()
source.add_heading("Quarterly report", level=1)
source.add_paragraph("This paragraph was selected from the HTML application's data model.")

source.add_heading("Results", level=2)
table = source.add_table(rows=1, cols=2)
table.style = "Table Grid"
table.rows[0].cells[0].text = "Metric"
table.rows[0].cells[1].text = "Value"
for metric, value in [("Orders", "128"), ("Refunds", "4")]:
    cells = table.add_row().cells
    cells[0].text = metric
    cells[1].text = value

source.add_picture("logo.png", width=Inches(1.5))
source.save("report.docx")

This is not evidence of faithful, general-purpose HTML-to-DOCX conversion. Arbitrary web layouts, CSS positioning, scripts and responsive behavior do not map automatically to Word’s document model. For an existing HTML page, parse the parts you need, normalize them into headings, paragraphs, lists and tables, then add those structures with python-docx. If pixel-level preservation is the requirement, keep the PDF or page images as the authoritative rendering and evaluate a dedicated HTML-to-DOCX converter separately.

Use a hosted renderer when local setup is a poor fit

HTML2Image documents a Python client for an HTML-to-image API and an HTML-to-PDF API. Its vendor page stated Python 3.9 or newer and 50 starting free credits when it was crawled; verify the current offer, pricing, privacy terms, limits and output behavior before adoption. A hosted service can reduce local system-library maintenance, but it moves your HTML and assets to a third party and does not automatically solve fidelity, authentication or data-residency requirements.

Handle assets, authentication and untrusted input

  • Relative URLs: use a correct base_url or absolute asset paths. Confirm that the process can read every local file.
  • Authenticated pages: WeasyPrint’s default fetcher does not send cookies or authentication. Supply a custom fetcher or stage authorized assets locally.
  • Remote failures: a timeout, redirect or blocked host can produce a PDF with missing images. Log fetch failures and inspect each page.
  • Fonts: package the exact font files and configure them explicitly when typography matters.
  • Security: never render untrusted HTML with unrestricted file or network access. Apply an allowlist, isolate conversion workers and limit runtime, memory and output size.

Troubleshoot common failures

“Library” or system dependency errors during installation

Install the platform-specific libraries described by WeasyPrint, and install the PDF utilities required by pdf2image. Rebuild the deployment image from a documented, repeatable setup rather than copying a virtual environment from another operating system.

PDF is blank or missing CSS

Check the HTML base URL, stylesheet paths and file permissions. For remote resources, verify connectivity and whether authentication is required. Render a tiny document with one local image to separate path problems from CSS problems.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Fonts or images differ from the browser

Browser rendering and WeasyPrint do not support identical feature sets. Package fonts, inspect network-dependent assets and simplify unsupported CSS. Test the exact page rather than assuming browser pixel parity.

pdf2image cannot open the PDF

Confirm that the PDF was fully written, that the PDF utility is installed and discoverable, and that the file is not zero bytes. Open the PDF independently before debugging rasterization.

DOCX layout is not preserved

That is an architectural limitation of using python-docx for document construction. Map the HTML into Word paragraphs, runs, tables and images, or choose a dedicated conversion product after evaluating it against your pages.

Process is slow or runs out of memory

Reuse a worker process, avoid unnecessarily high image DPI, render PDF pages in ranges, and stream results to disk. Measure your own pages; the available documentation provides no neutral speed comparison.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server. It accepts consent banners before capture and removes more than 60 known consent platforms, newsletter popups and chat widgets; each cleanup step can be disabled. Only clean shots are billed: bot checks or CAPTCHAs, blank pages, timeouts, failed loads and cache hits are free, and response headers identify the page verdict and billing status.

For a one-call image or PDF capture, see the ScreenshotNeo documentation and use:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`HTTP ${res.status}`);
const data = Buffer.from(await res.arrayBuffer());
require('fs').writeFileSync('shot.webp', data);

It also provides an MCP server with take_screenshot, get_page_info and capture_pdf for Claude, Cursor and other MCP clients. Every plan includes its capture options, including full-page and element screenshots, device and viewport settings, retina scale, PDF paper and margin controls, custom CSS or JavaScript, waits, request blocking, headers, cookies, user agent, timezone, geolocation, resizing, caching, signed links, webhooks, bulk capture and a usage API. The Free plan includes 1,000 shots per month without a card; paid plans start at $5 for 3,000 shots. Sign up for the free plan.

Operational checklist

  • Choose PDF, raster images or editable DOCX before writing conversion code.
  • Test real CSS, fonts, images, page breaks and authenticated resources.
  • Record package, operating-system and PDF-utility versions.
  • Inspect output files, not just successful process exit codes.
  • Set time, memory and output-size limits for untrusted or large input.
  • For hosted rendering, verify current service terms and data handling.

Frequently Asked Questions

Can I convert HTML directly to an editable DOCX with python-docx?

python-docx constructs DOCX content; it does not document a general HTML parser and layout renderer. Extract the HTML elements you need and create corresponding Word structures, or evaluate a dedicated HTML-to-DOCX converter.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Why use PDF as the intermediate format for images?

pdf2image accepts PDF input and rasterizes its pages. Rendering HTML with WeasyPrint first gives the PDF renderer responsibility for CSS layout and page breaks.

Will WeasyPrint reproduce every browser CSS feature?

Not necessarily. Its supported HTML/CSS feature set differs from a browser, so test representative pages, especially those using advanced layout, scripts, remote assets or custom fonts.

Is a hosted API better than local Python libraries?

Neither is universally better. Hosted rendering can reduce system-dependency work, while local rendering offers more control over data and deployment. Compare fidelity, asset access, privacy, limits and current terms for your workload.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.