Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
Start by identifying what kind of PDF you have. If you can select text, use a PDF parser such as PyMuPDF. If pages are scans or photographs, run OCR with Tesseract. For tables, choose a layout-aware method and compare every result with the original page. A successful parser call proves only that software returned characters—not that reading order, columns, or table cells are correct.
1. Diagnose the PDF before extracting anything
A .pdf filename does not describe the file’s internal content. A document can contain a normal text layer, page images, or a mixture of both. Diagnosis determines the rest of your workflow.
Check for selectable text
- Open the file in a viewer and try selecting and copying a sentence.
- Run a small PyMuPDF test and count characters on each page.
- Inspect pages individually: a report may have text on most pages and scanned signatures or appendices on others.
import fitz # PyMuPDF
pdf = fitz.open("input.pdf")
for number, page in enumerate(pdf, start=1):
text = page.get_text("text")
print(f"page {number}: {len(text)} characters")
Pages returning little or no text may need OCR. A page with a text layer can still contain an image-only table, so treat this as a triage step rather than a quality guarantee.
Recommended Free Tools
Decide what “data” means for your project
- Plain text: fastest for search, indexing, and summarization.
- Reading-order-aware text: necessary for columns, sidebars, headings, and footnotes.
- Tables: requires cell detection and validation.
- Scanned content: requires OCR before normal text processing.
- Structured JSON: useful when a hosted extraction API should return text, images, and tables together.
2. Extract text with PyMuPDF
PyMuPDF opens a PDF, lets you iterate through pages, and exposes text with page.get_text(). Preserve page boundaries so each value can be traced back to its source.
#1 Best Overall
- Scanner type: Document
- Connectivity technology: USB
- With Auto Scan Mode, the scanner automatically detects what you're scanning
- Digitize documents and images
import fitz
from pathlib import Path
source = Path("input.pdf")
out = []
with fitz.open(source) as document:
for page_number, page in enumerate(document, start=1):
text = page.get_text("text")
out.append(f"n--- Page {page_number} ---n{text}")
Path("extracted.txt").write_text("n".join(out), encoding="utf-8")
The plain-text mode is a good first pass, but it does not promise the order a person sees. PDFs may store a right-column paragraph before a left-column paragraph, or place headers and footers between body lines.
Use structured and spatial output when order matters
PyMuPDF also exposes blocks, words, and coordinates. Use those modes when you need to group content by page region, remove repeated headers, or reconstruct columns. A practical approach is to inspect words with their bounding boxes, sort within a known column region, and retain the original page number for auditability. Do not silently merge columns into prose without checking representative pages.
3. Fix wrong reading order and layout
“Why is the extracted PDF text in the wrong order?” Usually, the parser is reporting the PDF’s content order, not its visual reading order. Multi-column articles, floating captions, text boxes, headers, footers, and tables are common causes.
A validation routine that catches errors
- Print one extracted page beside the rendered page image.
- Check the title, first paragraph, column transitions, footnotes, and page numbers.
- Compare several page types, not just the first page.
- If order is wrong, switch to block or word coordinates and process regions separately.
- Keep an exception list for pages requiring manual review.
For searchable archives, store both the raw extraction and any cleaned version. Cleaning should never destroy the evidence needed to explain a disputed value.
4. OCR a scanned or image-only PDF
Scanned pages contain pixels rather than characters. PyMuPDF’s documented OCR integration uses Tesseract, which must be installed separately. The OCR flow creates a text page that can then be searched or extracted like other page content.
Rank #2
- FAST SPEEDS - Scans color and black and white documents a blazing speed up to 16ppm (1). Color scanning won’t slow you down as the color scan speed is the same as the black and white scan speed.
- ULTRA COMPACT – At less than 1 foot in length and only about 1. 5lbs in weight you can fit this device virtually anywhere (a bag, a purse, even a pocket).
- READY WHENEVER YOU ARE – The DS-640 mobile scanner is powered via an included micro USB 3. 0 cable allowing you to use it even where there is no outlet available. Plug it into you PC or laptop and you are ready to scan.
- WORKS YOUR WAY – Use the Brother free iPrint&Scan desktop app for scanning to multiple “Scan-to” destinations like PC, Network, cloud services, Email and OCR. (2) Supports Windows, Mac and Linux and TWAIN/WIA for PC/ICA for Mac/SANE drivers. (3)
- OPTIMIZE IMAGES AND TEXT – Automatic color detection/adjustment, image rotation (PC only), bleed through prevention/background removal, text enhancement, color drop to enhance scans. Software suite includes document management and OCR software. (4)
import fitz
with fitz.open("scan.pdf") as document:
for page_number, page in enumerate(document, start=1):
# OCR requires Tesseract installed and available to PyMuPDF.
ocr_page = page.get_textpage_ocr()
text = page.get_text("text", textpage=ocr_page)
print(f"--- Page {page_number} ---")
print(text)
Consult the PyMuPDF OCR recipe for installation-specific details and language configuration. OCR recognizes text; it does not reconstruct every visual or semantic feature. The documentation notes that Tesseract does not recognize vector graphics and that OCR text has simplified font properties.
Make OCR affordable and repeatable
OCR is materially slower than ordinary extraction. PyMuPDF documentation states: “Because optical character recognition is about one thousand times slower than standard text extraction, we make sure to do OCR only once per page and store the result in a TextPage.” Detect pages needing OCR, create the OCR text page once, cache it, and reuse it for searches and extraction. Do not OCR every page by default when only a few are image-only.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
OCR quality checks
- Review names, decimal points, minus signs, dates, and currency symbols.
- Check columns and reading order; OCR can recognize words while still misplacing them.
- Compare low-confidence or business-critical fields against the scan.
- Record the page image and OCR settings used for each accepted value.
5. Extract tables without trusting the first result
Table extraction is layout-dependent. Lines, whitespace, merged cells, rotated labels, and background shading all affect detection. PyMuPDF provides Page.find_tables(); table objects can be exported, including to pandas DataFrames.
import fitz
with fitz.open("report.pdf") as document:
for page_number, page in enumerate(document, start=1):
finder = page.find_tables()
for table_number, table in enumerate(finder.tables, start=1):
dataframe = table.to_pandas()
dataframe.to_csv(
f"page-{page_number}-table-{table_number}.csv",
index=False
)
Line-based detection depends on vector graphics such as drawn borders. For borderless tables, the PyMuPDF FAQ describes a text-based strategy:
finder = page.find_tables(strategy="text")
Background-color-only tables and unusual structures can remain difficult. When automatic extraction fails, combine words and their coordinates with explicit row and column rules, then validate row assignments against the rendered page.
Rank #3
- Up to 255 customize favorite scan file setting with "Single Touch" , Support Windows 7/8/10
- Turn paper documents into searchable, editable files - save scans as searchable PDF files; OCR function included
- Info Barcode function - automatic categorization of complicate documentation and data with 1D or 2D Barcode page.
- Intelligent color and image adjustments — Auto Rotate, Crop, Deskew and blank page remove with Plustek Image Processing Technology
- Easy send scanned files to FTP server or personal NAS (FTP) with PDFs , Jpeg , TIFF or Png format. User can download scanner driver from Plustek website
What to verify in every extracted table
- Header labels and units survived.
- Numbers stayed in the correct row and column.
- Thousands separators and decimal marks were not altered.
- Blank cells, merged cells, subtotals, and footnotes have an explicit representation.
- The number of rows and columns matches the source page.
6. When Camelot is a better fit
Camelot is designed for table extraction from text-based PDFs. It can be useful when you want a quick CSV or DataFrame and the document’s tables have consistent geometry. Scanned pages need OCR first, or Camelot’s documented OCR-enabled setup.
Quick wins for a faster PC:
Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Repair Windows errors before they cause bigger problemsFix Now →There is no universal winner. Choose after checking three facts: whether text is selectable, whether tables have ruling lines, and whether you need a fast dataset or a carefully reconstructed table. Test a few representative pages before processing a large batch.
7. Use a hosted API for structured JSON
The Adobe PDF Services API documents extraction of text, images, tables, and other content from native and scanned PDFs into structured JSON. This can reduce local dependency management when your application already uses a hosted workflow.
Confirm current pricing, quotas, data-handling suitability, geographic availability, and retention terms directly with the provider before sending sensitive documents or designing around the service. Those details are not established by the extraction documentation itself.
8. A decision framework
| Situation | Start with | Important qualification |
|---|---|---|
| Selectable text, simple pages | PyMuPDF get_text() |
Validate reading order and retain page boundaries. |
| Selectable text, complex columns | PyMuPDF blocks/words and coordinates | Reconstruct regions and inspect representative pages. |
| Bordered tables | PyMuPDF find_tables() or Camelot |
Check merged cells, units, and row alignment. |
| Borderless tables | Text strategy plus coordinate rules | Whitespace and alignment can be ambiguous. |
| Scanned pages | Tesseract through PyMuPDF OCR | Install Tesseract; cache OCR because it is much slower. |
| Mixed pages | Page-by-page detection | OCR only image-only pages and retain provenance. |
| Hosted structured output | Adobe PDF Services API | Verify current commercial and privacy terms first. |
9. Troubleshooting common failures
The output is empty
Cause: the page is image-only, encrypted, or the text layer is damaged. Fix: test permissions, render the page, and use OCR for image content. Handle mixed PDFs page by page.
The Tool Desk
Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Rank #4
- Note: No software installation is required. You need 2 AA batteries ( not included) and a memory card ( included) to use it directly. Scan mode: Press and hold "Scan" for 2 seconds to turn on the device, and then press "Scan", the green light is on. The scanner moves to scan the file until the green light turns off automatically (or press the "Scan" key and the green light goes out). The number shown on the display increases by 1 to indicate that the scan is complete.
- Portable Scanner scans images or pictures quickly: Store JPEG/PDF files within seconds, scan images or pictures quickly, plug and play, no need any software preinstalled. Compatible with Windows XP/7/Vista/Mac OS 10.4 or above version.
- Lightweight and travel-friendly: Stored in Micro SD card directly, support read data on your computer or phone with USB connected. Powered by 2pcs AA batteries, Compact Design, it is convenient to carry outside.
- 3 Image Resolution: 3 modes of resolution for your options: 300dpi/600dpi/900dpi, you can save it at the clearest way, picture and document are showed clear as it is. Freely choose your favorite resolution.File Format: JPEG/PDF format is all available, Great storage capacity as it supports 32G Micro SD card(Included 16GB Card),total meet your need for business trip or daily use.
- Widely Used: It is applicable in bank, insurance business, real estate agency,home, office, library or outdoors. suitable for lawyer, businessmen, students, travelers and amateur archivists. Scan your important files and save them immediately, no struggling in finding a printing shop, keep it confidential.
Words appear in a nonsensical order
Cause: internal content order differs from visual layout. Fix: use blocks or word coordinates, process columns as regions, and compare with the original.
Tables are shifted or missing columns
Cause: borderless, merged, shaded, or unusual table geometry. Fix: try strategy="text", inspect coordinates, and manually validate critical rows.
OCR is too slow
Cause: OCR has a substantial processing cost compared with normal extraction. Fix: detect image-only pages, OCR once per page, cache the resulting TextPage, and avoid rerunning it for each search.
OCR text looks readable but values are wrong
Cause: character confusion, low resolution, skew, or complex layouts. Fix: render at a suitable resolution, configure the language, and verify names, numbers, and symbols against the image.
Free tools Windows power users keep installed
One-click scans. No signup required.
10. “Or skip the browser setup:” capture a clean PDF or page image
If the document is published behind a web page and you first need a stable visual copy, ScreenshotNeo can return a screenshot or PDF from one GET request. It accepts consent banners like a visitor and removes 60+ known consent platforms, newsletter popups, and chat widgets before capture; bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing status. Its MCP server lets Claude, Cursor, or another MCP client call take_screenshot, get_page_info, and capture_pdf.
Use the API documentation at https://screenshotneo.com/docs/ for all options, including full-page capture, CSS selection, custom waits, headers, cookies, blocking rules, PDF page ranges, caching, signed links, asynchronous jobs, and bulk capture.
Best Value
- PORTABLE SCANNER FOR USE ON-THE-GO — The fastest and lightest mobile single-sheet-fed compact document scanner in its class¹
- QUICK DOCUMENT SCANNING ― This Epson ultra-fast scanner scans a single page as quickly as 5.5 seconds²; Windows and Mac compatible
- VERSATILE PAPER HANDLING ― Portable scanner scans documents up to 8.5 x 72 in; Also easily digitizes receipts and ID cards to make accounting, bookkeeping, and organizing simpler
- INTUITIVE, HIGH-SPEED SOFTWARE — Epson ScanSmart Software³ is a smart tool allowing you to easily scan, review, and save; Stay organized easily with the help of this Epson scanner
- EASY SETUP — USB-powered connect to your computer for quick and simple scanning; No batteries or external power supply required to operate portable document scanner; Standard Connectivity: USB 2.0
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
The Free plan includes 1,000 screenshots per month with no card. Paid plans start at $5 for 3,000 shots; every feature is available on every plan. Create a free ScreenshotNeo account.
11. Build an auditable extraction pipeline
- Store the original PDF immutably.
- Record file hash, page count, extraction library version, and OCR language/settings.
- Save raw per-page output before cleaning.
- Keep coordinates or table metadata for values that may be challenged.
- Sample pages from each layout type for visual review.
- Send only validated fields to downstream systems.
This separation between extraction and verification prevents a plausible-looking parser result from becoming an untraceable fact.
Frequently Asked Questions
Can one PDF require both normal extraction and OCR?
Yes. PDFs commonly mix text pages with scanned inserts, signatures, or image-only tables, so choose the method per page.
Does OCR recover charts and diagrams as data?
Not reliably. OCR supplies recognized text; vector graphics and visual semantics need separate analysis or manual review.
Should I delete page boundaries from the final text?
Keep them in the source and audit copy. You may create a second reader-friendly version, but page provenance is essential for checking values.
Is a successful table export proof that the table is correct?
No. Compare headers, units, row counts, merged cells, and representative values with the rendered PDF.
Windows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallOutdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchQuick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

