October DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsSlow PC?RecommendedPC slow today? Run a repair scan before it gets worseResolve common Windows issues and optimize system performance.Scan NowOctober DealsAmazon USDeal season is back - check today's better picksAmazon US: current deals, useful picks and tech finds.See Picks×
Skip to content
OCR

Build Your Own PDF Tools With Python: A Practical Guide to Generation, Editing, Extraction, and OCR

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

There is no single “best” Python PDF library. The right design separates four jobs: generating a new document, structurally editing existing pages, fast rendering or conversion, and layout-aware extraction. A small, deliberate stack is easier to test and deploy than one library forced to do everything.

Use ReportLab to create invoices, reports, and forms; pypdf to merge, split, crop, transform, encrypt, and add metadata; PyMuPDF for fast rendering, conversion, inspection, and broad manipulation; and pdfplumber when coordinates, lines, rectangles, tables, or visual debugging matter. For scanned pages, add OCR with a separately installed Tesseract program.

Choose the library for the job

Task First choice Why it fits Main caveat
Generate invoices, reports, forms, or other new PDFs ReportLab Generation-oriented APIs and an official Python PDF-generation guide Layout is programmatic; the commercial ReportLab PLUS edition has separate licensing
Merge, split, crop, transform, encrypt, or edit metadata pypdf Pure Python with explicit support for these page operations It is not a document-generation engine
Fast rendering, conversion, extraction, or document-wide inspection PyMuPDF High-performance and broad document manipulation Wheel/operating-system compatibility and MuPDF licensing require review; OCR needs Tesseract
Coordinates, lines, rectangles, tables, and visual debugging pdfplumber Detailed geometry access and table-extraction tools Works best with machine-generated PDFs; scans need OCR first

These choices are complementary. For example, ReportLab can generate a report, pypdf can append a cover page and password-protect the result, PyMuPDF can render thumbnails for a review screen, and pdfplumber can inspect table geometry.

Set up a reproducible Python project

Create an isolated environment and install only what the first workflow needs. Pin versions in your application once you have tested them; PDF output can change when a dependency changes.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
python -m venv .venv
# macOS/Linux
source .venv/bin/activate
# Windows PowerShell
.venvScriptsActivate.ps1

python -m pip install --upgrade pip
python -m pip install reportlab pypdf pymupdf pdfplumber

The package commonly imported as fitz is installed from the pymupdf package. PyMuPDF publishes wheels for Windows 32-bit and 64-bit Intel, Linux 64-bit Intel and ARM, and macOS 64-bit Intel and ARM. If your platform has no suitable wheel, pip may build from source and require C/C++ tools. Pillow is needed for PIL image methods, fontTools for font subsetting, and pymupdf-fonts for additional fonts. pdfplumber requires Python 3.8 or newer and is MIT licensed.

Before accepting uploads, set explicit limits for file size, page count, processing time, and output size. Reject malformed files before they reach a worker, write outputs to a separate directory, and never treat a user-supplied filename as a safe path.

Generate a PDF with ReportLab

ReportLab is the generation choice when your input is data rather than an existing PDF. The Platypus layer lets you compose paragraphs, tables, page breaks, and page templates while the lower-level canvas API gives precise drawing control.

from reportlab.lib import colors
from reportlab.lib.pagesizes import letter
from reportlab.lib.styles import getSampleStyleSheet
from reportlab.lib.units import inch
from reportlab.platypus import SimpleDocTemplate, Paragraph, Spacer, Table, TableStyle

out = "invoice.pdf"
doc = SimpleDocTemplate(out, pagesize=letter,
                        rightMargin=0.6*inch, leftMargin=0.6*inch,
                        topMargin=0.6*inch, bottomMargin=0.6*inch)
styles = getSampleStyleSheet()
story = [Paragraph("Invoice 1042", styles["Title"]),
         Paragraph("Acme Services", styles["Heading2"]),
         Spacer(1, 12)]
rows = [["Item", "Qty", "Price"],
        ["Design work", "2", "$300.00"],
        ["Support", "1", "$75.00"],
        ["Total", "", "$675.00"]]
table = Table(rows, colWidths=[3.8*inch, 0.7*inch, 1.2*inch])
table.setStyle(TableStyle([
    ("BACKGROUND", (0, 0), (-1, 0), colors.HexColor("#eeeeee")),
    ("GRID", (0, 0), (-1, -1), 0.5, colors.grey),
    ("ALIGN", (1, 1), (-1, -1), "RIGHT"),
    ("FONTNAME", (0, 0), (-1, 0), "Helvetica-Bold"),
]))
story.append(table)
doc.build(story)
print(f"wrote {out}")

For long documents, use Platypus flowables rather than manually calculating every y-coordinate. Define page headers and footers in an onFirstPage/onLaterPages callback, register fonts deliberately, and test text wrapping with the longest realistic values. ReportLab’s vendor distinguishes its open-source software from the separately licensed PLUS edition, so check the edition and license that match your distribution model.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Edit, merge, split, and protect PDFs with pypdf

pypdf is a pure-Python library for page-level work. It can split and merge documents, crop and transform pages, edit metadata, encrypt files, and perform basic text and metadata extraction.

Merge files and add metadata

from pypdf import PdfReader, PdfWriter

writer = PdfWriter()
for filename in ("cover.pdf", "report.pdf", "appendix.pdf"):
    reader = PdfReader(filename)
    for page in reader.pages:
        writer.add_page(page)
writer.add_metadata({
    "/Title": "Quarterly report",
    "/Author": "Acme Services",
})
with open("combined.pdf", "wb") as f:
    writer.write(f)

Split selected pages

from pypdf import PdfReader, PdfWriter

reader = PdfReader("combined.pdf")
for number, page_index in enumerate((0, 1, 4), start=1):
    writer = PdfWriter()
    writer.add_page(reader.pages[page_index])
    with open(f"page-{number}.pdf", "wb") as f:
        writer.write(f)

Crop, rotate, and encrypt

from pypdf import PdfReader, PdfWriter

reader = PdfReader("input.pdf")
writer = PdfWriter()
for page in reader.pages:
    # Coordinates use the PDF page's points; adjust for your source geometry.
    page.mediabox.left = 36
    page.mediabox.bottom = 36
    page.rotate(90)
    writer.add_page(page)
writer.encrypt("strong-password")
with open("secured.pdf", "wb") as f:
    writer.write(f)

Always inspect the resulting page boxes after cropping and rotation. A crop box changes what viewers display but does not necessarily remove hidden content; use a redaction workflow when permanent removal is required. Encryption also does not make an unsafe password safe, and permission flags are not a substitute for access control.

Use PyMuPDF for speed, rendering, and broad inspection

PyMuPDF is positioned as a high-performance library for extraction, analysis, conversion, and manipulation. It is useful when you need page images, fast text extraction, document-wide checks, or conversions in a service.

import pymupdf

src = pymupdf.open("input.pdf")
print("pages:", src.page_count)
for index, page in enumerate(src):
    print(index + 1, page.rect, len(page.get_text()))
    pix = page.get_pixmap(matrix=pymupdf.Matrix(2, 2), alpha=False)
    pix.save(f"preview-{index + 1}.png")
src.close()

Rendering at a larger matrix improves visual inspection but increases memory and output size. Close documents promptly, especially in workers processing many files. Keep representative PDFs in tests: a text-heavy file, one with rotated pages, one with embedded images, and one with unusual fonts.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Extract tables and coordinates with pdfplumber

pdfplumber exposes individual text characters, their positions, lines, rectangles, table extraction, and visual-debugging helpers. It is strongest when the PDF was generated from text and drawing commands rather than scanned as an image.

import pdfplumber

with pdfplumber.open("statement.pdf") as pdf:
    page = pdf.pages[0]
    words = page.extract_words()
    print(words[:5])
    table = page.extract_table()
    if table:
        for row in table:
            print(row)
    page.to_image(resolution=150).save("debug-page.png")

Table extraction depends on ruling lines, whitespace, and nearby text. Start by examining a debug image, then tune table settings for that document family. Do not assume one setting works for every statement or invoice template. If a page has no text layer, run OCR first and expect coordinate and table quality to be lower than on machine-generated source.

OCR scanned PDFs with Tesseract and PyMuPDF

A scan is a collection of page images, not searchable text. PyMuPDF’s OCR path depends on separately installed Tesseract-OCR software; installing the Python package alone is not enough. Install Tesseract using your operating system’s package manager or installer, verify that the executable is on the worker’s PATH, and check the language data you need.

import pymupdf

src = pymupdf.open("scan.pdf")
for page in src:
    # OCR requires a working Tesseract installation.
    text_page = page.get_textpage_ocr(language="eng", dpi=300)
    text = page.get_text("text", textpage=text_page)
    print(text[:500])

OCR is slower and can misread columns, punctuation, handwriting, and low-resolution pages. Preserve the original scan, record the OCR language and resolution, and route low-confidence results for review. After OCR, pdfplumber can work with the generated text layer, but table extraction still needs validation.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Design the processing pipeline

  1. Validate the boundary. Check MIME type, file signature, size, page count, and a processing deadline before opening the file.
  2. Choose the first operation. Generate with ReportLab, edit with pypdf, inspect or render with PyMuPDF, or extract geometry with pdfplumber.
  3. Keep passes explicit. Save intermediate files only when needed, and name each transformation so failures can be diagnosed.
  4. Preserve intentional metadata. Copy or replace title, author, creation date, page size, rotation, and bookmarks according to your retention policy.
  5. Validate output. Reopen the result with a second reader, check page count and text, render sample pages, and open it in a normal PDF viewer.
  6. Deploy with pinned dependencies. Test the exact Python version, operating system, native wheels, Tesseract installation, fonts, and locale used in production.

Common failures and fixes

“No module named fitz” or an import failure

Install pymupdf in the active virtual environment and confirm the interpreter with python -m pip show pymupdf. Avoid accidentally installing an unrelated package named fitz.

pip tries to compile PyMuPDF

Your platform may lack a matching wheel. Upgrade pip, verify that your Python architecture is supported, or provide the documented C/C++ build tools. Test the resulting image in the same operating system used by production.

Text extraction returns an empty string

The page may be scanned, have an unusable text layer, or use damaged encoding. Render it to an image to distinguish a visual page from a text page, then use Tesseract OCR if it is an image.

pdfplumber finds no table

Confirm the page is machine-generated, inspect character positions, save a visual debug image, and tune table settings for that template. OCR the page first if it is a scan.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Output is rotated, clipped, or unexpectedly large

Inspect media, crop, and rotation boxes after every pypdf transform. In PyMuPDF, lower the render matrix or image resolution when producing previews. Check font embedding and image compression in generated documents.

OCR works locally but not in production

Tesseract is external software. Install it in the worker image, include language data, expose its executable on PATH, and log the version and selected language.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Performance, reliability, and cost considerations

There is no universal speed ranking in the available documentation, so benchmark your own document families rather than relying on a generic claim. Measure parse time, OCR time, peak memory, output size, and error rate separately. Reuse workers carefully, close files, and place unusually large or high-page-count jobs on a bounded queue.

Cache deterministic intermediate results when input bytes and options are identical. For security, isolate untrusted processing, limit concurrency, scrub temporary files, and avoid logging extracted personal data. Keep a golden set of PDFs and compare page count, geometry, text, metadata, and rendered images after dependency upgrades.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Or skip the browser setup

If your PDF workflow starts by capturing a web page, ScreenshotNeo can return a clean PNG, JPEG, WebP, or PDF with one request. Its capture steps can accept cookie or consent banners and remove more than 60 known consent platforms, newsletter popups, and chat widgets before the shot; each step can be disabled. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing status.

Example using cURL (see the ScreenshotNeo API documentation):

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`HTTP ${res.status}`);
const fs = await import('node:fs/promises');
await fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer()));

For PDF capture, specify the PDF options in the request. The service also supports full-page capture with lazy images loaded, CSS-selector element capture, dark mode, device presets, retina scale, custom CSS and JavaScript, clicks, waits, blocking rules, headers, cookies, user agents, authorization, timezone, geolocation, transparent backgrounds, resizing, chosen cache TTLs, signed image links, asynchronous jobs with signed webhooks, bulk capture of up to 100 URLs per call, a usage API, and an OpenAPI specification. Its parameter names are compatible with those used by other screenshot APIs, which can simplify migration.

An MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients. The Free plan includes 1,000 shots per month with no card; paid plans start at $5 for 3,000 shots. Sign up for the free ScreenshotNeo plan.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Frequently Asked Questions

Can one library generate, edit, extract, and OCR every PDF well?

Not reliably. Separate generation, page editing, high-performance inspection, geometry-aware extraction, and OCR so each stage uses the tool designed for it.

Does OCR make a scanned PDF identical to a native text PDF?

No. OCR adds an inferred text layer and can misread characters, columns, or handwriting. Keep the original scan and validate important results.

Should I use pypdf or PyMuPDF for merging files?

pypdf is the straightforward pure-Python choice for page-level merging. Use PyMuPDF when the same pipeline also needs fast rendering, conversion, or broad inspection.

Why does a PDF look cropped after editing?

Page boxes and rotation are separate properties. Inspect media and crop boxes after transformations, then render a representative page in a viewer.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Read next

Recommended PC Tool
Recommended PC Tool
PC Slower Than It Used to Be?Free scan - under a minute
Outdated Drivers Are Slowing You DownFree scan - exact matches

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.