Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

First check whether the PDF page contains embedded text or only a scanned image. For embedded text, start with pypdf for straightforward extraction, PyMuPDF when you need layout or page geometry, and pdfplumber when you need detailed character and table inspection. Image-only pages require OCR; a text-extraction library cannot recover words from pixels by itself.

Choose the right approach for the PDF

PDFs preserve visual presentation, not necessarily semantic structure. Text may be positioned as individual fragments rather than stored as paragraphs in reading order, and a page can contain a mix of embedded text and images. As a result, extraction often involves inferring reading order, paragraph boundaries, headers, footers, and table structure. There may be no single uniquely correct text representation for every use.

Task Starting point Check before relying on output
Extract embedded text with a pure-Python library pypdf Reading order, unusual fonts, and whether the page is image-only. pypdf does not OCR images.
Extract text with positions or layout options PyMuPDF Whether the selected output mode reconstructs the order and layout your application needs.
Inspect characters, lines, rectangles, or tables pdfplumber Table settings, visible borders, and whether the PDF is machine-generated. The project says it works best on machine-generated rather than scanned PDFs.
Recognize text on scanned pages OCR workflow, such as PyMuPDF OCR Language support, recognition errors, and whether output has been checked against the page image.

These are capability-based choices, not a universal accuracy or speed ranking. Results depend on how the particular PDF was created and what output your application needs.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Install a library and extract text page by page

Install the package that matches your chosen approach. The examples below use a local file named input.pdf and write a text file with explicit page boundaries. Keeping those boundaries makes it easier to trace a suspicious result back to its source page.

Option A: pypdf for basic embedded text

Install with python -m pip install pypdf, then save and run:

from pathlib import Path
from pypdf import PdfReader

pdf_path = Path("input.pdf")
reader = PdfReader(pdf_path)

with Path("extracted.txt").open("w", encoding="utf-8") as output:
    for page_number, page in enumerate(reader.pages, start=1):
        text = page.extract_text() or ""
        output.write(f"n--- Page {page_number} ---n")
        output.write(text)
        output.write("n")

print(f"Extracted {len(reader.pages)} pages to extracted.txt")

An empty or very short result does not prove the PDF page is blank. It may be an image scan, have unusual text encoding, or use a layout that ordinary extraction cannot reconstruct. pypdf explicitly says, “pypdf is no OCR software.” Use it for text already represented in the PDF, not for recognition of text in page images.

Option B: PyMuPDF for text and layout-oriented work

Install with python -m pip install pymupdf. This example extracts page text and also shows how to request word coordinates for downstream layout handling:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
from pathlib import Path
import pymupdf

pdf_path = Path("input.pdf")

with pymupdf.open(pdf_path) as document:
    with Path("extracted.txt").open("w", encoding="utf-8") as output:
        for page_number, page in enumerate(document, start=1):
            output.write(f"n--- Page {page_number} ---n")
            output.write(page.get_text("text"))

    first_page = document[0] if len(document) else None
    if first_page is not None:
        words = first_page.get_text("words")
        print("First-page word records:", words[:10])

Word records include coordinates as well as text. Those positions can help you devise a reading order or inspect columns, but they do not automatically guarantee the order your application expects. PyMuPDF also documents OCR workflows for pages that need recognition.

Option C: pdfplumber when page geometry or tables matter

Install with python -m pip install pdfplumber. The basic text and table example keeps its output grouped by page:

from pathlib import Path
import pdfplumber

with pdfplumber.open("input.pdf") as pdf:
    with Path("extracted.txt").open("w", encoding="utf-8") as output:
        for page_number, page in enumerate(pdf.pages, start=1):
            output.write(f"n--- Page {page_number} ---n")
            output.write(page.extract_text() or "")

            tables = page.extract_tables()
            for table_number, table in enumerate(tables, start=1):
                print(f"Page {page_number}, table {table_number}:")
                for row in table:
                    print(row)

Table output is a starting point for inspection, not a guarantee that every cell has been identified correctly. pdfplumber exposes detailed page objects and table controls; whether those controls work depends on the PDF’s lines, alignment, and visual design.

Recognize text on scanned pages with OCR

A scan may look like a page of text while containing only a raster image. Ordinary text extraction reads a text layer; it does not recognize the letters drawn in pixels. When a page produces little or no text, inspect the page visually and route scan-like pages to OCR.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

PyMuPDF documents an OCR approach using its page OCR method. OCR support requires the separate Tesseract OCR engine to be installed and available in the environment, as well as the relevant language data. Check the PyMuPDF OCR documentation for setup and supported options for your platform.

import pymupdf

with pymupdf.open("input.pdf") as document:
    for page_number, page in enumerate(document, start=1):
        # OCR this page and obtain a text page for extraction.
        text_page = page.get_textpage_ocr()
        text = page.get_text(textpage=text_page)
        print(f"--- Page {page_number} ---")
        print(text)

OCR output is a recognition result, not ground truth. Verify names, numbers, punctuation, columns, and other high-impact fields against the page image. If a document contains both selectable text and scans, a per-page workflow lets you use ordinary extraction where it works and OCR where it is needed.

Extract tables without assuming every grid is a real table

A table’s visual appearance does not guarantee that the PDF encodes it as rows and cells. Some tables have drawn borders or vector lines; others align text without borders, or distinguish cells with background color. A line-based detector may find the former more readily than the latter. PyMuPDF’s FAQ notes that missing borders and background-color-only designs are harder to detect.

  1. Inspect the page and determine whether borders, lines, or aligned text mark its rows and columns.
  2. Try the selected library’s table feature on representative pages and inspect the returned rows and cells.
  3. If cells are merged, borderless, or misaligned, adjust the library’s table settings or use character and page coordinates to build document-specific spatial logic.
  4. Compare the resulting data with the visible page, including row order, empty cells, and values spanning multiple columns.

For detailed analysis of lines, rectangles, and characters, pdfplumber is a reasonable choice, particularly for machine-generated PDFs. Scanned tables still need OCR, and recognition of words alone does not resolve the table’s row and column relationships.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Validate extraction before using the data

Test with representative documents from the application, not only a clean one-page example. Decide what the result must preserve: page boundaries, line breaks, headers, footers, columns, or table cells. PDF extraction has to infer structure that the file may not encode semantically, so validation is part of the implementation.

  • Check multi-column pages for interleaved lines or incorrect reading order.
  • Look for missing glyphs, broken ligatures, unexpected spacing, and characters substituted by unusual fonts.
  • Decide whether repeated headers, footers, and page numbers belong in your output; remove them only when the downstream task calls for it.
  • Compare table cells with the page, especially merged or borderless cells and values aligned only by position.
  • For OCR, verify likely recognition errors against the page image before acting on extracted names, dates, amounts, or identifiers.
  • Retain page numbers or other source references so incorrect output can be investigated without searching the entire document.

There is no single accuracy or speed figure that applies across these packages and document types. Choose by the required capability, then evaluate on the files and output format your application actually uses.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Troubleshooting common extraction failures

Symptom Likely cause What to do
Text output is empty or nearly empty The page may be image-only, or its text encoding may be difficult to interpret. Check whether text can be selected in a PDF viewer. If the page is a scan, use OCR; if it has embedded text, try another extraction mode or inspect its structure.
Text appears but is out of order Text fragments may be stored by position or drawing order, not reading order. Use positional output such as PyMuPDF’s word data, inspect coordinates, and build or select an ordering approach appropriate to that layout.
Table rows or columns are missing The table may lack detectable borders, use background fills, or have merged cells. Inspect page geometry, adjust table settings, and consider custom spatial logic for the document’s layout.
OCR text has mistakes OCR can misread characters, particularly where print quality, layout, or language recognition is challenging. Check language and OCR setup, then compare important values against the image. Do not treat OCR output as verified data.
Scanned table text is readable but the cells are scrambled Recognizing words and recovering table structure are separate problems. Combine OCR with layout or table analysis and verify the cell assignments visually.
Results contain repeated titles or page numbers Headers and footers are page content and may be extracted like body text. Define whether those elements belong in your target data, then filter them with rules suited to the document set.

Or skip the browser setup

PDF parsing in Python is for files you already have; ScreenshotNeo is a separate website screenshot API for capturing web pages as images or PDFs. A single request can return a screenshot or PDF, which can be useful when the source document is a live web page rather than a local PDF. See the ScreenshotNeo API documentation.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

ScreenshotNeo accepts cookie or consent banners as a visitor and removes 60+ known consent platforms, newsletter popups, and chat widgets before capture; each of those steps can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing, and responses include X-Page-Verdict and X-Billed headers. Its MCP server provides take_screenshot, get_page_info, and capture_pdf tools for AI agents using Claude, Cursor, or another MCP client.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The Free plan includes 1,000 screenshots per month with no card required; paid plans start at $5 for 3,000 shots. Sign up for ScreenshotNeo’s free plan.

Frequently Asked Questions

Can pypdf extract text from a scanned PDF?

No. pypdf extracts embedded text; scanned page images need OCR.

Which Python library should I use for PDF tables?

Use a table-capable approach such as pdfplumber or PyMuPDF, then inspect the extracted cells against the page because table structure depends on how the PDF was authored.

Are PDF text extraction results guaranteed to follow reading order?

No. The PDF may store text fragments in a different order from the intended reading sequence, so check representative pages and use position data or document-specific logic where needed.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.