Do these 3 things before closing this tab:
1Repair Windows errors before they cause bigger problems2Scan for outdated or missing drivers - takes under a minute3Clear out junk files and repair common Windows errorsPowerShell can run a PDF text extractor and then read its output; it cannot extract PDF text with Get-Content alone. One documented route uses Apache PDFBox: for version 3.x, run its export:text command, then load the resulting text file in PowerShell. Scanned, image-only PDFs are a separate case: ordinary text extraction does not establish that their text has been recognized.
What you need to extract PDF text
Use a PDF parser such as Apache PDFBox to interpret the PDF and write extracted text to a file. PowerShell can launch the extractor and work with the resulting file. Microsoft documents Get-Content as a way to read file contents, not as a PDF parser; it does not convert a PDF by itself.
- A Java runtime accessible on your system, so the
javacommand can run. - The PDFBox application JAR for the release whose command syntax you intend to use.
- The PDF file you want to extract, with a path you can supply to the command.
The commands below are based on the official documentation, not a reported test. Replace example filenames and paths with real ones, and check the installed release’s help or documentation before relying on a particular option. See Apache PDFBox 3.0 Command-Line Tools and Apache PDFBox 2.0 Command-Line Tools.
Extract text with PDFBox 3.x
Open PowerShell in the folder containing the PDFBox JAR and input PDF, or use full paths. PDFBox 3.x documents the export:text command with explicit input and output arguments:
Free tools Windows power users keep installed
One-click scans. No signup required.
#1 Best Overall
- Book - powershell for sysadmins: workflow automation made easy
- Language: english
- Binding: paperback
java -jar .pdfbox-app-3.y.z.jar export:text -i=.input.pdf -o=.output.txt
3.y.z is not a literal version or filename. Substitute the actual PDFBox 3.x JAR filename you have, such as the complete filename downloaded for your release. Likewise, change input.pdf and output.txt to your files. The output path tells PDFBox where to write the extracted text.
- Check that Java is available by running
java -version. If PowerShell cannot findjava, install or configure Java so its executable is accessible. - Confirm the JAR and PDF paths. A simple check from the working folder is
Test-Path -LiteralPath .pdfbox-app-3.y.z.jarandTest-Path -LiteralPath .input.pdf, after replacing the example JAR name. - Run the PDFBox command with your actual paths.
- Check that the output file exists, then read it as described below. Review the text for missing content or unexpected reading order.
PDFBox 3.x documents UTF-8 as its default text encoding and provides page-range and sorting controls. It also documents Markdown output as available since version 3.0.4. Options vary by release, so consult the documentation or help for the exact JAR you run rather than assuming an option from another version applies.
Use a PowerShell variable for the JAR path
Variables can make it easier to change paths without editing the command in several places:
$jar = '.pdfbox-app-3.y.z.jar'
$inputPdf = '.input.pdf'
$outputTxt = '.output.txt'
java -jar $jar export:text "-i=$inputPdf" "-o=$outputTxt"
Replace the placeholder JAR filename and sample document names. This invokes Java directly from PowerShell; it is not a PDF conversion feature built into PowerShell.
Quick wins for a faster PC:
Clear out junk files and repair common Windows errorsFree Scan →Scan for outdated or missing drivers - takes under a minuteDriver Scan →Rank #2
Read or save the extracted text in PowerShell
Once PDFBox has created the text file, Get-Content can read it. Use -LiteralPath to treat the supplied path literally, and -Raw when you want one string instead of an array of lines.
$text = Get-Content -LiteralPath .output.txt -Raw
$text
To save a copy under another filename, use PowerShell file-writing commands after extraction. For example:
$text = Get-Content -LiteralPath .output.txt -Raw
Set-Content -LiteralPath .copy.txt -Value $text -Encoding utf8
The extraction step is still the PDFBox command; Get-Content and Set-Content operate on text files. Microsoft documents the difference between reading lines and returning the content as one string in its PowerShell 7.5 Get-Content reference.
Run extraction and check the output
You can put the steps in one script. This example checks the input path and confirms that an output file was created; it does not guarantee that the extracted text is complete or correctly ordered.
PC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchRank #3
$jar = '.pdfbox-app-3.y.z.jar'
$inputPdf = '.input.pdf'
$outputTxt = '.output.txt'
if (-not (Test-Path -LiteralPath $inputPdf)) {
throw "Input PDF not found: $inputPdf"
}
java -jar $jar export:text "-i=$inputPdf" "-o=$outputTxt"
if (-not (Test-Path -LiteralPath $outputTxt)) {
throw "PDFBox did not create the expected text file: $outputTxt"
}
$text = Get-Content -LiteralPath $outputTxt -Raw
$text
For production scripts, also inspect the Java process’s exit result and capture its diagnostic output as appropriate for your environment. If you launch the process with Start-Process, Microsoft documents that cmdlet for starting a specified executable and warns that untrusted data used with its FilePath parameter can create a security risk. Keep executable paths under your control and treat user-supplied paths as data, not as executable names. See the PowerShell 7.5 Start-Process reference.
PDFBox 2.x uses different syntax
Do not send the 3.x export:text form to a PDFBox 2.x JAR. The 2.x documentation uses the older ExtractText command form:
java -jar .pdfbox-app-2.y.z.jar ExtractText [OPTIONS] .input.pdf .output.txt
Here, too, replace 2.y.z with the actual JAR filename. The brackets around [OPTIONS] indicate where supported options go; they are not text to type literally. Consult the 2.x documentation for the options available in that release. Do not assume that 3.x flags, output controls, or behavior transfer unchanged to 2.x.
| Release documentation | Documented command form | Practical implication |
|---|---|---|
| PDFBox 3.0 | export:text -i=... -o=... |
Use the 3.x command form and check the installed release’s options. |
| PDFBox 2.0 | ExtractText [OPTIONS] <inputfile> [Text file] |
Use the older command form; do not mix it with 3.x syntax. |
What if the PDF is a scan?
A scanned PDF may contain page images instead of a text layer. The documented PDFBox text-extraction command does not establish that image text will be recognized. If the output is empty or lacks words visible on the page, first determine whether the PDF contains selectable text or only images. For an image-only scan, use an OCR process; the PDFBox command documented here should not be presented as OCR.
The Tool Desk
Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Rank #4
After OCR has created searchable text or a text layer, an extractor can read that text, but the OCR step itself is separate. Check extracted output against the source pages, especially for small print, columns, tables, handwriting, or skewed scans.
Options and output quality
Page ranges, ordering, and encoding
The PDFBox 3.0 command-line documentation describes page-range controls and sorting options, and identifies UTF-8 as the default encoding. These can matter when processing only selected pages, mixed layouts, or text with non-English characters. Verify the precise option names and accepted values for the release you installed; do not copy options from a different major version without checking.
Expect extracted text, not a perfect page replica
PDF text extraction yields text content, not a faithful reproduction of page layout. Multi-column pages, positioned labels, footnotes, tables, and unusual reading order can produce text that needs cleanup. Compare important results with the PDF itself instead of assuming the output’s line sequence matches visual reading order.
Troubleshooting common problems
javais not recognized: PowerShell cannot locate the Java executable. Install or configure Java and ensure its command is available to the shell, then verify withjava -version.- The JAR cannot be opened: Check the exact JAR filename and current directory, or pass a full path. Confirm the JAR exists with
Test-Path -LiteralPath. - PDFBox reports an unknown command or option: The command may belong to another major release. Use
export:textfor the documented 3.x form andExtractTextfor the 2.x form; check the installed JAR’s own documentation. - No output file appears: Confirm the input PDF path, output directory, and write permissions. Read Java/PDFBox error output rather than treating the absence of a file as successful extraction.
- Output exists but is empty or incomplete: The document may be image-only, protected, or structured in a way that extracts poorly. Check whether its text is selectable, verify any required password option against the release documentation, and inspect pages with complex layout.
- Text order looks wrong: PDF page placement does not always correspond to a simple reading sequence. Check the documented sorting controls for your PDFBox version and manually validate layout-sensitive content.
- Characters display incorrectly: Check the output encoding and how the text viewer reads it. PDFBox 3.x documents UTF-8 as the default, but verify version-specific behavior and the actual file.
Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server, not a PDF text extractor. If your real task is to capture a webpage rather than extract text from a PDF, one GET request can return a screenshot or PDF:
Best Value
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
See the ScreenshotNeo API documentation. It accepts cookie or consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing status. Its MCP server gives AI agents tools for screenshots, page information, and PDF capture. The free plan includes 1,000 shots per month with no card; paid plans start at $5 for 3,000 shots. For PDF text extraction, continue using an extractor such as PDFBox. Learn about ScreenshotNeo or sign up free for 1,000 screenshots a month with no card.
Frequently Asked Questions
Can PowerShell extract PDF text without installing another utility?
Not with Get-Content alone. It reads text files; a PDF parser or extractor must first interpret the PDF.
Can PDFBox extract text from a password-protected PDF?
The PDFBox 3.x command-line documentation lists a password option. Check the option syntax for the installed release and ensure you are authorized to access the document.
Can I extract only certain pages?
PDFBox 3.x documents page-range controls. Use the exact option syntax from the documentation matching your installed release.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




