What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
To extract text from a PDF in PHP, install smalot/pdfparser with Composer, parse the file with parseFile(), and call getText(). The example below also shows how to parse PDF bytes already in memory, read an individual page, and retrieve available document details. The library’s Packagist page lists PHP 7.1+ as its requirement; check that your installed release supports the PHP version used by your application.
Install the package and extract a PDF’s text
Run Composer in your project directory:
composer require smalot/pdfparser
This adds the package and its dependencies to the project. In a PHP script, load Composer’s autoloader, create a parser, and pass the PDF’s path to parseFile():
<?php
require __DIR__ . '/vendor/autoload.php';
$parser = new SmalotPdfParserParser();
$pdf = $parser->parseFile(__DIR__ . '/document.pdf');
$text = $pdf->getText();
echo $text;
Save this as a PHP file in the project, put document.pdf beside it, and run the script from the command line with PHP. The output is the text returned by the parser; this minimal example sends it directly to standard output. In a web application, decide how the extracted text should be returned or stored rather than echoing untrusted document content into an HTML page without appropriate output handling.
Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minuteWindows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallThe package publisher describes smalot/pdfparser as a standalone PHP package for extracting data from PDF files. Its documented basic flow is the one shown here: instantiate Parser, parse a file, and request its text. See the Packagist package page for installation and package information and the official usage documentation for the API examples.
#1 Best Overall
- PORTABLE SCANNER FOR USE ON-THE-GO — The fastest and lightest mobile single-sheet-fed compact document scanner in its class¹
- QUICK DOCUMENT SCANNING ― This Epson ultra-fast scanner scans a single page as quickly as 5.5 seconds²; Windows and Mac compatible
- VERSATILE PAPER HANDLING ― Portable scanner scans documents up to 8.5 x 72 in; Also easily digitizes receipts and ID cards to make accounting, bookkeeping, and organizing simpler
- INTUITIVE, HIGH-SPEED SOFTWARE — Epson ScanSmart Software³ is a smart tool allowing you to easily scan, review, and save; Stay organized easily with the help of this Epson scanner
- EASY SETUP — USB-powered connect to your computer for quick and simple scanning; No batteries or external power supply required to operate portable document scanner; Standard Connectivity: USB 2.0
Choose the input method that matches your PDF
Parse a file by path
Use parseFile($path) when the PDF is already available as a file that PHP can read. The basic example uses __DIR__ so the path is anchored to the script’s directory rather than depending on the process’s current working directory. Replace document.pdf with the actual path in your application.
If the path comes from a request, do not treat arbitrary user input as a safe filesystem path. Resolve user-supplied identifiers to files your application is permitted to access, and apply your own upload validation, access controls, and resource limits. The package’s usage documentation is an API guide, not a complete recipe for safely accepting uploads.
Parse PDF bytes already in memory
If another part of your program has already read the PDF into a string of bytes, use parseContent() instead of writing it to a temporary file solely for parsing:
<?php
require __DIR__ . '/vendor/autoload.php';
$content = file_get_contents(__DIR__ . '/document.pdf');
if ($content === false) {
throw new RuntimeException('Could not read the PDF.');
}
$parser = new SmalotPdfParserParser();
$pdf = $parser->parseContent($content);
echo $pdf->getText();
The explicit read check distinguishes a failed file read from a string containing the file’s bytes. For an uploaded or remotely supplied PDF, obtain the bytes through your application’s validated input flow; this example only demonstrates the parser call.
Rank #2
- FAST SPEEDS - Scans color and black and white documents a blazing speed up to 16ppm (1). Color scanning won’t slow you down as the color scan speed is the same as the black and white scan speed.
- ULTRA COMPACT – At less than 1 foot in length and only about 1. 5lbs in weight you can fit this device virtually anywhere (a bag, a purse, even a pocket).
- READY WHENEVER YOU ARE – The DS-640 mobile scanner is powered via an included micro USB 3. 0 cable allowing you to use it even where there is no outlet available. Plug it into you PC or laptop and you are ready to scan.
- WORKS YOUR WAY – Use the Brother free iPrint&Scan desktop app for scanning to multiple “Scan-to” destinations like PC, Network, cloud services, Email and OCR. (2) Supports Windows, Mac and Linux and TWAIN/WIA for PC/ICA for Mac/SANE drivers. (3)
- OPTIMIZE IMAGES AND TEXT – Automatic color detection/adjustment, image rotation (PC only), bleed through prevention/background removal, text enhancement, color drop to enhance scans. Software suite includes document management and OCR software. (4)
Decode Base64 before parsing
Base64 is an encoding of the PDF bytes, not a PDF parsing mode. Decode the string first, check that decoding succeeded, and then pass the resulting bytes to parseContent():
<?php
require __DIR__ . '/vendor/autoload.php';
$base64 = $encodedPdf;
$content = base64_decode($base64, true);
if ($content === false) {
throw new InvalidArgumentException('The input is not valid Base64.');
}
$parser = new SmalotPdfParserParser();
$pdf = $parser->parseContent($content);
echo $pdf->getText();
$encodedPdf must be a variable containing the Base64 string supplied to your application. If the incoming value includes a data-URL prefix, handle that format in your input layer before decoding; the parser expects PDF content bytes, not the Base64 representation.
Read a page or document details
Extract text from one page
After parsing, the usage documentation shows accessing the parsed pages through getPages(). For the first page, request its text like this:
$pages = $pdf->getPages();
if (isset($pages[0])) {
echo $pages[0]->getText();
}
PHP arrays use a zero-based index here, so $pages[0] refers to the first element returned. Check that the element exists before using it; a document with no accessible page should not be treated as though a first page were available. To inspect additional pages, iterate over the returned pages and call getText() on each one.
Rank #3
- STAY ORGANIZED – Easily convert your paper documents into digital formats like searchable PDF files, JPEGs, and more.Power Consumption : 2.5W or less (Energy Saving Mode: 0.7W). Suggested Daily Volume : 500 scans..Does it contain liquid: no
- CONVENIENT AND PORTABLE –lightweight and small in size, you can take the scanner anywhere from home offices, classrooms, remote offices, and anywhere in between
- HANDLES VARIOUS MEDIA TYPES – Digitize receipts, business cards, plastic or embossed cards, reports, legal documents, and more
- FAST AND EFFICIENT – No technical hurdles or complicated setups here; easily scan both sides of a document at the same time, in color or black-and-white, at up to 12 pages-per-minute, and with a 20 sheet automatic feeder
- BROAD COMPATIBILITY – Works with both Windows and Mac devices, be it laptop or computer
Retrieve available metadata
Call getDetails() on the parsed document to obtain the details exposed by the library:
$details = $pdf->getDetails();
foreach ($details as $name => $value) {
if (is_scalar($value) || $value === null) {
echo $name . ': ' . (string) $value . PHP_EOL;
} else {
echo $name . ': ' . json_encode($value) . PHP_EOL;
}
}
The documentation establishes that getDetails() returns document details; it does not guarantee that every PDF contains every metadata field. Treat the returned data as variable rather than assuming a particular title, author, or other property will always be present. The sample formats scalar values directly and encodes other values for readable command-line output.
What this example does—and does not—establish
This is text extraction from a PDF, not a guarantee that every visual element or every file will yield usable text. A PDF may contain text as page content or present information in ways that do not produce the expected extracted result. The cited package documentation does not establish OCR support for scanned or image-only pages, so do not rely on this example to recognize text inside page images.
Free tools Windows power users keep installed
One-click scans. No signup required.
Likewise, a successful call does not promise that extracted text will match the visual layout exactly. If your use case depends on reading order, tables, coordinates, or a particular document structure, validate that result against representative files and the relevant API documented by the project. The sources for this example do not provide a broad accuracy or performance comparison.
Rank #4
- IRIScan Express, portable scanner : scans color and black and white documents a blazing speed up to 8ppm simplex. Color scanning won’t slow you down as the color scan speed is the same as the black and white scan speed.
- IRIScan Express mobile scanner is powered via an included micro USB 2. 0 cable allowing you to use it even where there is no outlet available. Plug it into you PC or laptop and you are ready to scan. USB cable provided. AC Adapter not provided and not needed.
- IRIScan flatbed scanner uses a simplex scanning mode allows for quick and straightforward scanning of single-sided documents. IRIScan with its full portable features is the ideal document scanners for computers.
- IRIScan document scanner : Versatile scanning capabilities, including scanning to Word, PDF, and Excel formats with companion software provided Readiris OCR
- Receipt scanner and card scanner with Additional features include scanning business cards directly to Outlook, photo scanning, and receipt scanning for efficient document management
The package page says secured documents and form-data extraction are not supported. The usage documentation describes encrypted PDFs as unsupported by default and mentions a setIgnoreEncryption configuration option. An option to ignore encryption is not evidence that every encrypted file can be parsed or that its contents become accessible. Do not treat it as a substitute for valid access to a document or as a guarantee of successful extraction.
Release and maintenance considerations
At the time reflected by the package listing accessed in 2026, Packagist listed smalot/pdfparser version 2.13.0-beta1, published on 2026-09-25. That is a beta release, not a stable release label; check the package page for the version Composer will resolve and for updates before choosing a dependency for production. The same page states a PHP requirement of 7.1 or later and describes the project as under limited maintenance. The maintenance statement is the publisher’s stated status, not an independent assessment of support quality.
For production use, assess whether the documented file types, encryption behavior, PHP requirement, and maintenance status fit your application. Pin and update dependencies through your normal Composer workflow, and test your own representative PDFs when changing package versions. Neither the package page nor the usage examples cited here establish a performance benchmark, accuracy rate, security guarantee, or service-level commitment.
Troubleshooting common problems
- Composer cannot resolve or install the package: Check that Composer is running in the intended project directory and that the project’s PHP version satisfies the package requirement shown on Packagist. Review the exact Composer error and the version constraints in your project before changing dependencies.
Class 'Smalot\PdfParser\Parser' not found:Confirm thatvendor/autoload.phpis required from the correct path, that Composer installation completed successfully, and that the script is running inside the project containing thatvendordirectory.- The PDF cannot be read: Verify the path is correct relative to the script, that the file exists, and that the PHP process has permission to read it. If using
file_get_contents(), check forfalsebefore passing content toparseContent(). - The input is Base64 but parsing fails: Decode the Base64 string before calling
parseContent(). Use strict decoding and handle failure; do not pass the encoded characters as though they were the PDF bytes. - Extracted text is empty or incomplete: Confirm the intended file and page were parsed. If the page is scanned or image-only, this text-extraction example does not establish OCR capability. Check the PDF and the library’s documented supported behavior rather than assuming that an empty result means the PHP call was wired incorrectly.
- An encrypted or secured PDF does not parse: The package documentation identifies encrypted files as unsupported by default and secured documents as unsupported. Although an ignore-encryption configuration option is documented, it does not ensure that a particular protected file will work.
- Metadata output is missing a field:
getDetails()exposes available details; a particular field is not guaranteed by the cited documentation. Inspect what the document actually returns instead of assuming every PDF carries identical metadata.
Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server, not a PHP PDF parser; it does not replace the extraction examples above. If your separate task is to capture a website rather than parse a PDF, one GET request can return an image or PDF. The cURL example below uses the documented API endpoint and a sample target URL:
Best Value
- Scanner type: Document
- Connectivity technology: USB
- With Auto Scan Mode, the scanner automatically detects what you're scanning
- Digitize documents and images
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
See the ScreenshotNeo API documentation for the request options. ScreenshotNeo removes cookie/consent banners, newsletter popups, and chat widgets before capture; bot checks, blank pages, failed loads, and cache hits are not billed. Its MCP server offers screenshot tools for AI agents, and the free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000. Those features apply to website captures, not PDF text extraction.
Learn about ScreenshotNeo, or sign up free for 1,000 screenshots a month with no card.
Frequently Asked Questions
Does this package example convert an entire PDF into a searchable document?
No. The code demonstrates extracting text from a PDF. The cited documentation does not establish OCR for scanned or image-only pages.
Can I use the ScreenshotNeo request above to extract text from a PDF?
No. ScreenshotNeo captures websites as images or PDFs; it is separate from the PHP PDF parsing example and does not perform the text extraction shown here.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

