Free tools Windows power users keep installed
One-click scans. No signup required.
Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
There is no single best open-source PDF parser. Choose pypdf for straightforward text, metadata, and page operations; pdfplumber when you need to inspect coordinates and tune table extraction; PyMuPDF for an all-purpose extraction, rendering, manipulation, and OCR workflow (after reviewing its AGPL/commercial licensing); and Apache PDFBox when your application is Java-based or needs forms, PDF/A validation, creation, printing, or signing. Scanned pages need a separate OCR path, and complex scientific papers, patents, and tables should be tested on your own representative files before you choose.
Quick comparison
The right choice depends on the PDF’s internal structure and what your application must do after extraction. The table below is a practical starting point, not a universal ranking.
| Library | Best fit | Important capabilities | Limits and obligations |
|---|---|---|---|
| pypdf 5.4.0 | Python text and document operations | Text and metadata retrieval; split, merge, crop, and transform pages; pure-Python implementation | Not a natural choice for rendering, OCR, or detailed table reconstruction; headers, footers, and page numbers may be indistinguishable from body text |
| pdfplumber | Layout inspection and tunable table extraction in Python | Low-level PDF objects, character coordinates, crop-box filtering, visual debugging, and tables exposing cells, rows, columns, and bounding boxes | Documentation places OCR, PDF generation, and PDF modification outside its scope; tables from OCRed documents have weak support |
| PyMuPDF | Broad extraction, rendering, manipulation, and OCR workflows | Text, images, vectors, rendering, tables, page manipulation, and on-demand Tesseract OCR; optional PyMuPDF4LLM outputs Markdown, JSON, or TXT for LLM workflows | AGPL or commercial licensing requires review for commercial deployment; published speed figures apply only to the vendor’s test corpus and method |
| Apache PDFBox | Java applications and full document lifecycle work | Unicode text extraction, forms, split/merge, PDF/A-1b preflight, printing, page images, PDF creation, and digital signing | Use the currently supported release and migration notes; the Apache project listed 3.0.8 (2026-07-11) and 2.0.37 (2026-07-15) at the time of the cited project information |
A 2024 comparative study found that results varied by document category: PyMuPDF and pypdfium generally did well on text extraction in that evaluation, while all tested parsers struggled with scientific and patent material and table leaders changed by category. Those findings are tied to that study’s datasets, metrics, and implementation versions, so they are useful for forming a test plan rather than declaring a winner.
Start with the PDF, not the library
Determine whether a text layer exists
Select a sentence in a desktop viewer and paste it into a plain-text editor. If nothing meaningful is pasted, the page is probably an image scan and a parser alone cannot recover its words. If text is present but scrambled, the problem is reading order or layout heuristics rather than the absence of a text layer.
#1 Best Overall
Identify the structures you must preserve
- Simple prose: paragraphs and headings can often be extracted with pypdf or PyMuPDF.
- Columns and positioned labels: preserve coordinates and inspect page regions; pdfplumber or PyMuPDF gives you more control than a plain text stream.
- Tables: verify cell boundaries, merged cells, repeated headers, and numeric alignment. A PDF stores drawing positions, not a semantic table model.
- Forms and signatures: use PDFBox or a library with explicit form and signing support rather than trying to infer fields from text.
- Images and scans: add OCR, then measure OCR errors separately from parser errors.
Account for deployment constraints
Python teams can evaluate all three Python libraries quickly. Java services may prefer PDFBox to avoid wrapping a different runtime. For a commercial product, review PyMuPDF’s AGPL and commercial options with your legal or compliance team before shipping; the technical feature list does not remove that obligation.
pypdf: the simple, pure-Python starting point
The pypdf 5.4.0 user guide describes a free, open-source, pure-Python library for retrieving text and metadata and for splitting, merging, cropping, and transforming pages. Its lack of a native C-library dependency can simplify installation in some environments.
Install and extract text
python -m pip install pypdf
from pathlib import Path
from pypdf import PdfReader
pdf_path = Path("input.pdf")
reader = PdfReader(str(pdf_path))
print("pages:", len(reader.pages))
print("metadata:", reader.metadata)
for number, page in enumerate(reader.pages, start=1):
text = page.extract_text() or ""
print(f"n--- page {number} ---n{text}")
Use the result as a first-pass text layer, not as proof that visual order is correct. Headers, footers, and page numbers are not always identifiable from the PDF alone.
Quick wins for a faster PC:
Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Clear out junk files and repair common Windows errorsFree Scan →Page operations
from pypdf import PdfReader, PdfWriter
reader = PdfReader("input.pdf")
writer = PdfWriter()
for page in reader.pages[:5]:
writer.add_page(page)
with open("first-five-pages.pdf", "wb") as out:
writer.write(out)
Choose pypdf when these operations and basic extraction are the core job. Move to a layout-aware tool when you need rendering, visual inspection, or reliable table geometry.
pdfplumber: inspect and tune layout
pdfplumber is built on pdfminer.six and exposes detailed PDF objects, character positions, crop boxes, visual debugging, and customizable text and table extraction. Its table API can return cells, rows, columns, and bounding boxes, which makes it useful when default extraction needs tuning.
Extract words and a table region
python -m pip install pdfplumber
import pdfplumber
with pdfplumber.open("input.pdf") as pdf:
page = pdf.pages[0]
words = page.extract_words()
print(words[:10])
# Adjust these coordinates to the table on your page.
table_area = page.crop((40, 120, 560, 500))
table = table_area.extract_table()
for row in table or []:
print(row)
Debug before changing settings
Render or display the page with detected lines and bounding boxes, then compare them with the visible table. Tune crop areas and table strategies only after checking whether the PDF contains ruling lines, whitespace-separated columns, or neither. pdfplumber does not provide OCR, PDF generation, or PDF modification, and its documentation warns that tables extracted from OCRed documents are not strongly supported.
PyMuPDF: a broad toolkit, with a license decision
PyMuPDF documents text extraction, rendering, image and vector handling, table extraction, page manipulation, and integration with Tesseract OCR. Its optional PyMuPDF4LLM tooling is aimed at layout analysis and semantic extraction for Markdown, JSON, TXT, and LLM workflows. Treat those as documented capabilities and validate the output on your corpus.
Extract text and render a page
python -m pip install PyMuPDF
import fitz # PyMuPDF
with fitz.open("input.pdf") as document:
for index, page in enumerate(document):
print(f"--- page {index + 1} ---")
print(page.get_text("text"))
pixmap = page.get_pixmap(matrix=fitz.Matrix(2, 2), alpha=False)
pixmap.save(f"page-{index + 1}.png")
OCR only when needed
PyMuPDF’s OCR API is on-demand and uses Tesseract. OCR is slower and can introduce recognition errors, so detect pages without usable text first and OCR those pages rather than every page. Install the Tesseract engine and the language data required by your documents separately; the Python package alone does not install them.
import fitz
with fitz.open("scan.pdf") as document:
page = document[0]
# Tesseract must be installed and available to PyMuPDF.
ocr_page = page.get_textpage_ocr(language="eng", dpi=300, full=True)
print(page.get_text("text", textpage=ocr_page))
Understand the license
PyMuPDF and MuPDF are available under AGPL and commercial license agreements, and the documentation identifies Artifex as MuPDF’s exclusive commercial licensing agent. Decide which terms apply to your distribution, hosting, and linking model before adopting it in a commercial deployment.
PyMuPDF’s documentation also reports timings on eight PDFs totaling 7,031 pages. Those are vendor measurements on a specified corpus and methodology, not a guarantee for your files or hardware.
Apache PDFBox: the Java choice for more than extraction
The Apache Software Foundation describes PDFBox as “an open source Java tool for working with PDF documents.” It is licensed under Apache License 2.0. In addition to Unicode text extraction, its feature list covers splitting and merging, form filling and extraction, PDF/A-1b preflight validation, printing, saving pages as images, PDF creation, and digital signing.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Minimal Maven setup and extraction
<dependency>
<groupId>org.apache.pdfbox</groupId>
<artifactId>pdfbox></artifactId>
<version>3.0.8</version>
</dependency>
Verify the current supported version and migration guidance before copying a version into production; the project information cited here listed 3.0.8 on July 11, 2026 and 2.0.37 on July 15, 2026.
import java.io.File;
import org.apache.pdfbox.Loader;
import org.apache.pdfbox.pdmodel.PDDocument;
import org.apache.pdfbox.text.PDFTextStripper;
public class ExtractPdf {
public static void main(String[] args) throws Exception {
try (PDDocument document = Loader.loadPDF(new File("input.pdf"))) {
PDFTextStripper stripper = new PDFTextStripper();
stripper.setSortByPosition(true);
System.out.println(stripper.getText(document));
}
}
}
PDFBox is particularly attractive when extraction is one step in a Java workflow that also validates PDF/A, handles AcroForms, renders pages, or signs output.
A reliable OCR and RAG ingestion pipeline
- Classify each file: record page count, whether selectable text exists, encryption status, and likely layout type.
- Extract native text first: retain page numbers and, where available, coordinates. Do not OCR a page that already has a good text layer without a reason.
- Detect suspicious output: flag pages with near-zero characters, repeated garbage glyphs, missing columns, or text in an implausible order.
- Render and OCR flagged pages: use a tool such as PyMuPDF with Tesseract, selecting language data and a resolution appropriate to the scan.
- Preserve structure for retrieval: chunk by page and detected heading or section, keep table rows together when possible, and store source page numbers for citations.
- Validate visually: inspect representative paragraphs, multi-column pages, footnotes, equations, and tables before indexing embeddings.
- Keep provenance: store the original file hash, parser and OCR versions, settings, and any crop or post-processing rules.
A parser can return text while still losing reading order or table relationships. That is why a RAG system should retain page-level evidence and make it possible to open the source page when an answer is challenged.
How to benchmark candidates fairly
Build a small corpus that mirrors production: born-digital prose, two-column reports, financial tables, scanned forms, scientific papers, and patents if those matter to you. For each file, compare:
- character and word recall against a manually checked reference;
- reading order across columns, headers, footers, and footnotes;
- table cell boundaries, row continuity, merged cells, and numeric values;
- OCR substitutions, missing symbols, and language-specific errors;
- rendering fidelity and page coordinate stability when downstream highlighting is required;
- runtime, memory, external dependencies, and failure behavior on encrypted or malformed files.
Run the same corpus and settings for every candidate, then manually inspect failures. A parser that wins on prose may lose on tables or scans. Academic and vendor benchmark figures should remain labeled with their dataset, metric, and methodology rather than being generalized into an accuracy promise.
Reliability, performance, and cost considerations
Performance
Native text extraction is usually cheaper than rendering plus OCR. Batch work by file, avoid rendering pages that do not need visual inspection, and cache intermediate text and OCR results using a content hash. Parallelism can improve throughput, but cap workers according to available memory because high-resolution page images are large.
Failure handling
Record a structured status for every page: native-text success, OCR used, empty output, password required, or parser exception. Keep the original PDF so a failed transformation can be retried with a different library. Never silently index an empty extraction.
Licensing and operations
pypdf and pdfplumber are Python dependencies with different scopes; PyMuPDF requires an AGPL-versus-commercial review; PDFBox uses Apache License 2.0. Include license notices in your release process and recheck them when upgrading. Operational cost also includes OCR binaries, language packs, temporary storage, and visual QA—not only the Python or Java package.
Troubleshooting common failures
“The extracted text is empty.”
The page is likely a scan, an image-only PDF, or protected content. Confirm by attempting selection in a viewer, then route the page through OCR. Check that the Tesseract executable and language data are installed.
“Text appears in the wrong order.”
PDF text is positioned rather than semantically structured. Try coordinate-aware extraction, crop columns separately, enable position sorting where supported, and compare the result with a rendered page. Do not assume a different chunk size will repair incorrect reading order.
Rank #4
“The table is shifted or merged.”
Inspect cell boundaries and ruling lines visually. In pdfplumber, crop to the table and tune extraction settings; in PyMuPDF, compare its table output with rendered coordinates. If the source is OCRed, expect weaker table support and consider a specialized post-processing step.
“The PDF opens but extraction crashes.”
Isolate the failing page, check encryption and malformed objects, and retry with a current library release. Keep a timeout and memory limit around batch jobs; route irreparably damaged files to a quarantine queue instead of blocking the entire batch.
Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minuteWindows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstall“Our commercial review rejects the dependency.”
For PyMuPDF, determine whether AGPL obligations fit your distribution or whether a commercial agreement is needed. If that does not fit, evaluate pypdf, pdfplumber, or PDFBox against the required features and document the trade-off.
Or skip the browser setup
If the source you need to process is a web-hosted document or documentation page, ScreenshotNeo can produce a deterministic image or PDF capture before your parser or OCR pipeline. It is separate from a PDF parsing library, but useful when the “PDF” starts as a page whose cookie banner, newsletter modal, or chat widget would otherwise contaminate the capture.
Use the API documentation for all options: ScreenshotNeo API docs.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
Before capture, ScreenshotNeo accepts the consent banner like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be disabled. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing status. Its MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients. Every plan includes the full feature set, including full-page and element capture, device and retina settings, custom CSS and JavaScript, waits, request blocking, headers and cookies, geolocation, PDF controls, caching, signed links, asynchronous jobs, bulk capture of up to 100 URLs per call, and a usage API.
The Tool Desk
Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →The Free plan includes 1,000 screenshots per month with no card. Paid plans start at $5 for 3,000 shots; yearly billing gives two months free. Create a free ScreenshotNeo account to try it.
Which parser should you choose?
- Choose pypdf for clean, selectable text and page-level document operations in Python.
- Choose pdfplumber when you need to see coordinates, crop regions, and tune table extraction interactively.
- Choose PyMuPDF when one toolkit must render, extract, manipulate, inspect tables, and invoke OCR—after resolving its license fit.
- Choose Apache PDFBox for Java services or workflows involving forms, PDF/A validation, creation, printing, and signatures.
- Add OCR for image-only pages and benchmark every candidate on the layouts your RAG system will actually receive.
Frequently Asked Questions
Can I use more than one parser in the same application?
Yes. A common design uses a fast native-text pass, sends only suspicious pages to OCR, and invokes a layout-focused library for documents that contain tables or columns. Keep the parser choice and settings in your provenance record so outputs remain reproducible.
How can I tell whether OCR or parsing caused an error?
Render the source page, compare it with the OCR text, and then compare that text with the parser’s structured output. If the OCR transcript is already wrong, changing PDF extraction settings will not fix the recognition error.
What should be retained for RAG citations?
Store the original file identifier, page number, extracted chunk, and enough coordinates or rendered-page references to let a reviewer verify the passage visually.
Do these 3 things before closing this tab:
1Fix the driver behind crashes, sound loss and screen glitches2Repair Windows errors before they cause bigger problems3Scan for outdated or missing drivers - takes under a minuteQuick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

