Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

Use a two-stage pipeline: normalize and validate your data first, then render it with a layout engine that matches the document you need. In Python, ReportLab is suited to programmatic drawing and report layouts, while WeasyPrint turns HTML and CSS into PDF. Keep data transformation separate from presentation, render short and long representative datasets, and inspect the resulting PDF for pagination, wrapping, links and required features.

Start with the data-to-document boundary

A reliable PDF generator does not format raw database rows directly on a page. It moves through explicit stages:

  1. Source: query results, API responses, spreadsheets or event data.
  2. Normalization: convert values into predictable types and names.
  3. Validation: reject or quarantine missing, invalid or contradictory values.
  4. Presentation mapping: decide labels, number formats, dates, colors and table columns.
  5. Rendering: pass the prepared representation to ReportLab or WeasyPrint.
  6. Verification: inspect the PDF and run automated checks before delivery.

This separation lets you change typography or page layout without changing business data, and lets you test transformations without requiring a PDF renderer.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Define output assumptions before coding

  • Who will read the document and whether it is printed or viewed on screen.
  • Page size and orientation (for example, A4 portrait or US Letter landscape).
  • Locale, currency, decimal precision and time zone.
  • How missing values, long labels and zero values should appear.
  • Whether the PDF needs hyperlinks, bookmarks, forms or attachments.
  • Whether tables can span pages and whether headers must repeat.

Normalize and validate structured data

Use a small, testable transformation layer. The example below converts records into values ready for either renderer.

from dataclasses import dataclass
from datetime import date
from decimal import Decimal, InvalidOperation

@dataclass
class SalesRow:
    product: str
    quantity: int
    unit_price: Decimal
    sold_on: date


def prepare_rows(records):
    rows = []
    for index, record in enumerate(records, start=1):
        product = str(record.get("product", "")).strip()
        if not product:
            raise ValueError(f"row {index}: product is required")
        try:
            quantity = int(record["quantity"])
            if quantity < 0:
                raise ValueError
        except (KeyError, TypeError, ValueError):
            raise ValueError(f"row {index}: quantity must be a non-negative integer")
        try:
            unit_price = Decimal(str(record["unit_price"]))
        except (KeyError, TypeError, InvalidOperation):
            raise ValueError(f"row {index}: invalid unit_price")
        rows.append(SalesRow(product, quantity, unit_price,
                             date.fromisoformat(record["sold_on"])))
    return rows


def display_rows(rows):
    return [
        [r.product, f"{r.quantity:,}", f"${r.unit_price:,.2f}", r.sold_on.isoformat()]
        for r in rows
    ]

Keep the internal value types (such as Decimal for money) until the presentation step. Do not use a binary floating-point value for financial totals merely because it is convenient to format.

Decide how exceptional values appear

  • Use an em dash or an explicit “Not provided” label for absent values; do not silently convert missing data to zero.
  • Choose one date format and apply it to every row.
  • Wrap or abbreviate long labels deliberately, and retain the full value in an accessible tooltip or note when appropriate.
  • Record validation failures with row identifiers so a caller can correct the source.

Choose ReportLab for Python-native layouts

ReportLab provides a lower-level pdfgen canvas for drawing directly on pages and higher-level flowables for reports containing paragraphs, tables and other layout elements. Its canvas uses points and lets you set the page size explicitly; do not rely on an implicit default when the output must print predictably.

Minimal ReportLab report with a repeating table header

from reportlab.lib import colors
from reportlab.lib.pagesizes import A4
from reportlab.lib.styles import getSampleStyleSheet
from reportlab.lib.units import mm
from reportlab.platypus import SimpleDocTemplate, Paragraph, Spacer, Table, TableStyle


def write_report(path, rows):
    doc = SimpleDocTemplate(
        path,
        pagesize=A4,
        rightMargin=15 * mm,
        leftMargin=15 * mm,
        topMargin=15 * mm,
        bottomMargin=15 * mm,
    )
    styles = getSampleStyleSheet()
    story = [Paragraph("Sales report", styles["Title"]), Spacer(1, 6 * mm)]
    data = [["Product", "Quantity", "Unit price", "Sold on"]]
    data.extend(display_rows(rows))

    table = Table(data, colWidths=[75 * mm, 25 * mm, 32 * mm, 32 * mm],
                  repeatRows=1, hAlign="LEFT")
    table.setStyle(TableStyle([
        ("BACKGROUND", (0, 0), (-1, 0), colors.HexColor("#e8eef7")),
        ("TEXTCOLOR", (0, 0), (-1, 0), colors.black),
        ("GRID", (0, 0), (-1, -1), 0.25, colors.grey),
        ("VALIGN", (0, 0), (-1, -1), "TOP"),
        ("ALIGN", (1, 1), (2, -1), "RIGHT"),
        ("BOTTOMPADDING", (0, 0), (-1, 0), 6),
    ]))
    story.append(table)
    doc.build(story)

# write_report("sales.pdf", prepare_rows(records))

ReportLab tables calculate row heights, can split across pages and can repeat header rows at page breaks. Explicit column widths are still important: they make wrapping and the available text width predictable. Test especially long product names and rows containing multiline text.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

When direct drawing is the better fit

Use the canvas when you need exact coordinates, custom diagrams, labels, invoices with fixed zones or page-by-page drawing logic. You must then manage text wrapping, page breaks and repeated elements yourself. Flowables and tables are usually safer for narrative reports because they can reflow as content changes.

Choose WeasyPrint for HTML and CSS templates

WeasyPrint accepts HTML and CSS and can write a PDF to a path or return PDF bytes. This model is useful when designers already work in templates, when styles are shared with a web view, or when the report contains familiar HTML structures such as headings, lists and tables.

Render a template to a PDF file

from weasyprint import HTML

html = """
<!doctype html>
<html>
<head>
  <meta charset="utf-8">
  <style>
    @page { size: A4; margin: 18mm; }
    body { font-family: sans-serif; font-size: 10pt; }
    h1 { color: #17365d; }
    table { width: 100%; border-collapse: collapse; }
    th, td { border: 0.25mm solid #999; padding: 4pt; }
    th { background: #e8eef7; }
    thead { display: table-header-group; }
    tr { break-inside: avoid; }
  </style>
</head>
<body>
  <h1>Sales report</h1>
  <table>
    <thead><tr><th>Product</th><th>Quantity</th><th>Price</th></tr></thead>
    <tbody>
      <tr><td>Example product</td><td>12</td><td>$19.00</td></tr>
    </tbody>
  </table>
</body>
</html>
"""
HTML(string=html).write_pdf("sales.pdf")

For generated content, escape text before inserting it into HTML and use a real template engine rather than concatenating untrusted strings. WeasyPrint documents supported HTML, CSS and PDF features and warns when a CSS property is unsupported. A warning is not proof that the layout is acceptable: verify the actual pages produced by your template.

Return bytes from an application endpoint

from weasyprint import HTML

def make_pdf(html_string: str) -> bytes:
    return HTML(string=html_string, base_url="/app/templates/").write_pdf()

# In a web framework, return the bytes with Content-Type: application/pdf.

Set a useful base_url when the document references local images, stylesheets or fonts. Without it, relative resources may not resolve in a server process.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

ReportLab or WeasyPrint?

Question ReportLab WeasyPrint
Authoring model Python drawing and document-layout objects HTML structure with CSS presentation
Best starting point Coordinates, flowables, tables and Python-controlled layouts Templates, web-style typography and CSS rules
Long tables Table splitting and repeated rows are documented features Use HTML table headers and verify page-break behavior
Primary risk More manual work for wrapping and complex visual layout when using the canvas Unsupported CSS or HTML features can produce warnings or unexpected output

Neither source establishes a controlled speed, fidelity or operating-cost winner. Select the representation that matches the report, then measure and inspect your own templates.

Make tables survive pagination

  1. Render a short dataset that fits on one page.
  2. Render enough rows to force several page breaks.
  3. Include very long labels, empty values and multiline descriptions.
  4. Check that column widths do not clip text or create unreadably narrow columns.
  5. Confirm that headers repeat and that a row is not split in a way that hides its meaning.
  6. Check totals and footnotes at the intended location.

For ReportLab, use repeatRows=1 and consider explicit widths. For WeasyPrint, use a table header group and page-break rules, then inspect the output because support depends on the exact CSS and HTML used.

Verification and reliability checklist

  • Open the PDF with more than one viewer.
  • Confirm the page size and orientation in document properties.
  • Search for representative text and verify that characters, accents and symbols survive.
  • Check hyperlinks, bookmarks, images and any required form or attachment feature.
  • Compare row counts and totals in the PDF with the validated input.
  • Keep source data, templates and generated PDFs as separate artifacts.
  • Repeat the same checks after changing a library, font, stylesheet or template.

Common failures and fixes

Text or columns are clipped

The available width is smaller than the content. Reduce padding, set realistic column widths, allow wrapping, or change the page orientation. Do not merely shrink the entire document until it is unreadable.

A table header disappears after a page break

In ReportLab, set repeatRows. In HTML, place headings in <thead> and apply the documented table-header display rule. Test with enough rows to create multiple breaks.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Relative images or CSS do not load

Give WeasyPrint an appropriate base_url, use resolvable paths, and verify permissions in the service account’s runtime environment.

A CSS rule is ignored

Read the renderer warning, check whether that property is supported, and replace it with a supported rule or redesign the layout. Always perform a representative render rather than assuming browser behavior is identical.

Dates or money look inconsistent

Normalize types before rendering and centralize formatting functions. Apply locale and time-zone decisions once, not independently in each template expression.

The process runs out of memory on a large report

Reduce unnecessary high-resolution images, avoid duplicating large strings, and split work into bounded jobs where your application permits it. The available documentation does not establish a universal row or file-size limit, so determine safe limits with your own data.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup

If your source is already an HTML page or dashboard and you need a PDF or a clean visual capture, ScreenshotNeo provides a website screenshot API. It accepts consent banners before capture and removes more than 60 known consent platforms, newsletter popups and chat widgets; each cleanup step can be disabled. Only clean shots are billed: bot checks, CAPTCHAs, blank pages, timeouts, failed loads and cache hits cost nothing, and the response identifies the result with X-Page-Verdict and X-Billed headers.

One GET request returns PNG, JPEG, WebP or PDF. The API supports full-page captures with lazy images loaded, CSS-selector elements, device presets and custom viewports, retina scale, PDF paper size, margins, landscape mode and page ranges, custom CSS or JavaScript, clicks, selector waits, delays, network-idle waits, request blocking, headers, cookies, user agents, authorization, timezone, geolocation, transparent backgrounds, resizing, chosen cache TTLs, signed links, asynchronous jobs with signed webhooks, bulk capture of up to 100 URLs per call, usage data and an OpenAPI specification. An MCP server exposes take_screenshot, get_page_info and capture_pdf to Claude, Cursor and other MCP clients.

See the ScreenshotNeo documentation for parameters and response behavior.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

The Free plan includes 1,000 shots per month without a card; paid plans start at $5 for 3,000 shots, with every feature on every plan. Create a free ScreenshotNeo account.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Commercial, template-driven options

ReportLab identifies RML as its commercial markup-based PDF-generation product, using a templating system populated with data. Current pricing and terms are not established here, so treat it as a procurement option to evaluate separately rather than assuming a particular plan or license.

Frequently Asked Questions

Can I use both libraries in one application?

Yes. A service can use ReportLab for coordinate-sensitive documents and WeasyPrint for HTML-based templates, provided each output path has its own tests and deployment dependencies.

What should I test when the dataset changes every day?

Keep a fixed validation fixture containing short, long, empty and malformed cases, and add a representative production sample to each release check.

Is PDF generation deterministic across machines?

Not automatically. Fonts, library versions, resource paths and renderer support can change pagination, so pin dependencies and validate output after environment changes.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.