What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
For a PHP application, install smalot/pdfparser with Composer, load Composer’s autoloader, then call parseFile() and getText(). The complete workflow is below, including deployment, compatibility checks, limitations, and recovery steps.
What you need before installing
The parser is a standalone PHP implementation for extracting data from PDF files. Its package manifest declares these requirements:
- PHP 7.1 or newer.
- The
iconvextension. - The
zlibextension. symfony/polyfill-mbstringversion^1.18, which Composer resolves as a package dependency.
Check the PHP binary used by your application, not only the one in your interactive shell. For example, php -v shows the CLI version and php -m lists enabled extensions. A web server can use a different PHP installation or configuration, so verify the same requirements in the deployment runtime.
Install smalot/pdfparser with Composer
Open a terminal in the root directory of the PHP project (the directory that contains, or will contain, composer.json) and run:
#1 Best Overall
composer require smalot/pdfparser
Composer adds the package and its dependencies to composer.json, downloads them into vendor/, updates composer.lock, and generates an autoloader. If Composer reports a missing PHP extension or an incompatible PHP version, fix that platform requirement before retrying; do not bypass it with a platform override unless you have verified the extension is available in production.
Use the lockfile correctly
For an application, commit both composer.json and composer.lock. Use composer install during deployment when the lockfile is present; it installs the exact versions recorded there. Use composer update only when you intentionally want Composer to resolve newer versions allowed by your constraints and rewrite the lockfile. This keeps development, staging, and production on the same dependency versions.
Parse a local PDF and extract its text
Create a PHP script such as extract.php in the project directory. The following is the documented usage pattern, with basic file checks added so a bad path fails clearly:
<?php
require __DIR__ . '/vendor/autoload.php';
use SmalotPdfParserParser;
$path = __DIR__ . '/document.pdf';
if (!is_file($path) || !is_readable($path)) {
throw new RuntimeException('PDF file is missing or unreadable: ' . $path);
}
$parser = new Parser();
$pdf = $parser->parseFile($path);
$text = $pdf->getText();
echo $text;
Run it from the project root with php extract.php. The script loads the generated Composer autoloader, creates SmalotPdfParserParser, parses the file, and writes the extracted text to standard output. Keep the PDF path outside user-controlled input unless you have separately implemented path validation and access controls.
Rank #2
What the parser can extract
The project documentation describes parsing PDF objects and headers, extracting metadata, and extracting text in page order. It also lists support for compressed PDFs, Mac OS Roman text, and hexadecimal or octal encoded text. Custom parser configuration is available, but the exact settings should be selected from the version installed in your project rather than copied blindly from an unrelated example.
Preserve page boundaries when your application needs them
getText() gives you the document’s extracted text as one result. If your application needs page-level indexing, rendering, or search results, use the package’s page-oriented API documented for the installed release and store the page number alongside each extracted segment. Do not assume that visual columns, headers, footers, or table cells will remain in their original layout after text extraction; PDF text positioning is not the same thing as a semantic document structure.
Metadata is separate from readable text
PDF metadata such as title, author, or creation information is represented separately from the visible text stream. If metadata is part of your workflow, read it through the metadata methods documented by the package version you have installed and treat missing or malformed values as optional. A metadata field is not proof that the corresponding text appears on a page.
Important limitations
| Requirement or document type | What the documentation establishes | Practical consequence |
|---|---|---|
| Encrypted or secured PDFs | The README says secured documents are unsupported. | Do not design a password-protected-document workflow around this parser. Obtain an authorized, usable copy or choose a library whose current documentation explicitly supports the security mode you need. |
| PDF form data | The README says form-data extraction is unsupported. | Visible text and interactive field values are different data. Use a tool that supports AcroForm or XFA data when field extraction is required. |
| Scanned, image-only pages | The reviewed documentation does not claim OCR capability. | An image-only scan may produce little or no text. Add a separate OCR step if you have authority to process the document. |
| Malformed or unusual PDFs | No accuracy or coverage benchmark is established here. | Test representative files from your own sources and keep a failure path instead of assuming every PDF will parse identically. |
The package is licensed under LGPL-3.0. Check that license against your distribution model and your organization’s legal policy. The README characterizes the project as being in limited maintenance: it remains compatible with supported PHP versions, but there is no active feature development and pull requests may not be reviewed promptly. That maintenance posture matters if your application needs rapid fixes or new PDF features.
Make the extraction workflow safer in production
Validate files before parsing
- Check that the path resolves inside an approved directory rather than accepting arbitrary filesystem paths.
- Confirm the file is readable and apply an application-level size limit before handing it to the parser.
- Keep the original file and the extracted result under separate access controls if the PDF contains personal or confidential information.
- Catch failures at your job or request boundary and record a document identifier, not sensitive PDF contents, in logs.
Design for variable resource use
PDFs can differ greatly in page count, embedded images, compression, and object complexity. The available material does not establish a universal memory, time, or accuracy limit, so measure your own representative files. For large batches, process one document per queue job or another bounded unit, set limits appropriate to your PHP runtime, and retain the original file when a retry or alternate parser may be needed. Avoid treating an empty result as proof that the PDF was empty; it can also indicate an image-only page, unsupported encoding, or a parsing failure.
Separate extraction from indexing
Write the extracted text to a controlled intermediate format before sending it to search indexing or downstream language processing. Record the parser package version and the input file’s checksum with that record. This makes it possible to reproduce a result after a dependency update without storing duplicate copies of the document in every downstream system.
Troubleshoot common failures
Composer says the PHP version is unsupported
Cause: The runtime used by Composer is older than PHP 7.1, or a project constraint prevents a compatible dependency set.
Fix: Run php -v, compare it with the PHP runtime used by the application, and upgrade or select a compatible runtime. Re-run Composer after the runtime is corrected; do not claim compatibility merely because another machine has a newer PHP binary.
Rank #4
Composer reports missing iconv or zlib
Cause: The required extension is disabled or absent in the PHP configuration Composer is evaluating.
Fix: Enable the extension in that PHP installation, restart the relevant PHP service if required by the platform, and verify with php -m (or the equivalent check in the web runtime). Then run the install again.
Class "SmalotPdfParserParser" not found
Cause: The script did not load Composer’s autoloader, or it is pointing at the wrong project directory.
Do these 3 things before closing this tab:
1Scan for outdated or missing drivers - takes under a minute2Clear out junk files and repair common Windows errors3Fix the driver behind crashes, sound loss and screen glitchesFix: Confirm that vendor/autoload.php exists beside the project and that the require path is correct. Run the script from the same deployment that contains the installed vendor directory.
The PDF path works locally but not on the server
Cause: Relative paths, case-sensitive filenames, permissions, or a different working directory.
Fix: Build the path from a known application directory (as in __DIR__), check is_file() and is_readable(), and verify the service account can read the file. Do not fix this by granting broad filesystem permissions.
Output is empty or unreadable
Cause: The document may be image-only, secured, form-based, malformed, or encoded in a way that needs review.
Free tools Windows power users keep installed
One-click scans. No signup required.
Fix: Open the PDF in an authorized viewer to determine whether selectable text exists, then compare the file with a known-good sample. For scans, add OCR; for secured documents or form fields, use a component that explicitly supports those features. Preserve the failing sample for regression tests.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Version selection and upgrades
Release information visible at different times is inconsistent: one Packagist view displayed v2.12.5 dated 2026-04-17, while another search result displayed v2.13.0-beta1 dated 2026-09-25. Because those snapshots conflict about the latest stable release, this article does not declare either number to be the definitive current stable version. The unpinned composer require smalot/pdfparser command lets Composer select a compatible release; after reviewing the resolved dependency set and testing your files, keep the resulting lockfile. If you need a deliberate version policy, choose and document a constraint after checking the package’s current release metadata and your PHP support range.
Or skip the browser setup
If the file you need is really a webpage capture rather than an existing local PDF, ScreenshotNeo can return a clean PNG, JPEG, WebP, or PDF from one request. It accepts cookie or consent banners like a visitor, removes more than 60 known consent platforms plus newsletter popups and chat widgets before capture, and reports whether a response was a clean page, a bot check, a blank page, a timeout, a failed load, or a cache hit. Only clean shots are billed; the other outcomes are not.
For the API parameters and the other capture options, see the ScreenshotNeo documentation. A minimal cURL request is:
Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minutePC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
The same request in Python:
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
And in Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
ScreenshotNeo also provides an MCP server with take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients. Every feature is included on every plan; 1,000 shots per month are free without a card, and paid plans start at $5 for 3,000 shots. If you want to capture a source page before processing related files, sign up for the free ScreenshotNeo plan.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

