Driver FixRecommendedSound, Wi-Fi or graphics acting up? Check drivers firstFind missing or outdated drivers fast.Check DriversOctober DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsWindows FixRecommendedWindows errors stealing your time? Find the fix fastScan stability, cleanup and performance issues.Fix Now×
Skip to content
Apache PDFBox

How to Convert PDF to Text with PowerShell

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

PowerShell can run a PDF text extractor and then read its output; it cannot extract PDF text with Get-Content alone. One documented route uses Apache PDFBox: for version 3.x, run its export:text command, then load the resulting text file in PowerShell. Scanned, image-only PDFs are a separate case: ordinary text extraction does not establish that their text has been recognized.

What you need to extract PDF text

Use a PDF parser such as Apache PDFBox to interpret the PDF and write extracted text to a file. PowerShell can launch the extractor and work with the resulting file. Microsoft documents Get-Content as a way to read file contents, not as a PDF parser; it does not convert a PDF by itself.

  • A Java runtime accessible on your system, so the java command can run.
  • The PDFBox application JAR for the release whose command syntax you intend to use.
  • The PDF file you want to extract, with a path you can supply to the command.

The commands below are based on the official documentation, not a reported test. Replace example filenames and paths with real ones, and check the installed release’s help or documentation before relying on a particular option. See Apache PDFBox 3.0 Command-Line Tools and Apache PDFBox 2.0 Command-Line Tools.

Extract text with PDFBox 3.x

Open PowerShell in the folder containing the PDFBox JAR and input PDF, or use full paths. PDFBox 3.x documents the export:text command with explicit input and output arguments:

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
#1 Best Overall
Sale
PowerShell for Sysadmins: Workflow Automation Made Easy
  • Book - powershell for sysadmins: workflow automation made easy
  • Language: english
  • Binding: paperback
java -jar .pdfbox-app-3.y.z.jar export:text -i=.input.pdf -o=.output.txt

3.y.z is not a literal version or filename. Substitute the actual PDFBox 3.x JAR filename you have, such as the complete filename downloaded for your release. Likewise, change input.pdf and output.txt to your files. The output path tells PDFBox where to write the extracted text.

  1. Check that Java is available by running java -version. If PowerShell cannot find java, install or configure Java so its executable is accessible.
  2. Confirm the JAR and PDF paths. A simple check from the working folder is Test-Path -LiteralPath .pdfbox-app-3.y.z.jar and Test-Path -LiteralPath .input.pdf, after replacing the example JAR name.
  3. Run the PDFBox command with your actual paths.
  4. Check that the output file exists, then read it as described below. Review the text for missing content or unexpected reading order.

PDFBox 3.x documents UTF-8 as its default text encoding and provides page-range and sorting controls. It also documents Markdown output as available since version 3.0.4. Options vary by release, so consult the documentation or help for the exact JAR you run rather than assuming an option from another version applies.

Use a PowerShell variable for the JAR path

Variables can make it easier to change paths without editing the command in several places:

$jar = '.pdfbox-app-3.y.z.jar'
$inputPdf = '.input.pdf'
$outputTxt = '.output.txt'

java -jar $jar export:text "-i=$inputPdf" "-o=$outputTxt"

Replace the placeholder JAR filename and sample document names. This invokes Java directly from PowerShell; it is not a PDF conversion feature built into PowerShell.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Read or save the extracted text in PowerShell

Once PDFBox has created the text file, Get-Content can read it. Use -LiteralPath to treat the supplied path literally, and -Raw when you want one string instead of an array of lines.

$text = Get-Content -LiteralPath .output.txt -Raw
$text

To save a copy under another filename, use PowerShell file-writing commands after extraction. For example:

$text = Get-Content -LiteralPath .output.txt -Raw
Set-Content -LiteralPath .copy.txt -Value $text -Encoding utf8

The extraction step is still the PDFBox command; Get-Content and Set-Content operate on text files. Microsoft documents the difference between reading lines and returning the content as one string in its PowerShell 7.5 Get-Content reference.

Run extraction and check the output

You can put the steps in one script. This example checks the input path and confirms that an output file was created; it does not guarantee that the extracted text is complete or correctly ordered.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
$jar = '.pdfbox-app-3.y.z.jar'
$inputPdf = '.input.pdf'
$outputTxt = '.output.txt'

if (-not (Test-Path -LiteralPath $inputPdf)) {
    throw "Input PDF not found: $inputPdf"
}

java -jar $jar export:text "-i=$inputPdf" "-o=$outputTxt"

if (-not (Test-Path -LiteralPath $outputTxt)) {
    throw "PDFBox did not create the expected text file: $outputTxt"
}

$text = Get-Content -LiteralPath $outputTxt -Raw
$text

For production scripts, also inspect the Java process’s exit result and capture its diagnostic output as appropriate for your environment. If you launch the process with Start-Process, Microsoft documents that cmdlet for starting a specified executable and warns that untrusted data used with its FilePath parameter can create a security risk. Keep executable paths under your control and treat user-supplied paths as data, not as executable names. See the PowerShell 7.5 Start-Process reference.

PDFBox 2.x uses different syntax

Do not send the 3.x export:text form to a PDFBox 2.x JAR. The 2.x documentation uses the older ExtractText command form:

java -jar .pdfbox-app-2.y.z.jar ExtractText [OPTIONS] .input.pdf .output.txt

Here, too, replace 2.y.z with the actual JAR filename. The brackets around [OPTIONS] indicate where supported options go; they are not text to type literally. Consult the 2.x documentation for the options available in that release. Do not assume that 3.x flags, output controls, or behavior transfer unchanged to 2.x.

Release documentation Documented command form Practical implication
PDFBox 3.0 export:text -i=... -o=... Use the 3.x command form and check the installed release’s options.
PDFBox 2.0 ExtractText [OPTIONS] <inputfile> [Text file] Use the older command form; do not mix it with 3.x syntax.

What if the PDF is a scan?

A scanned PDF may contain page images instead of a text layer. The documented PDFBox text-extraction command does not establish that image text will be recognized. If the output is empty or lacks words visible on the page, first determine whether the PDF contains selectable text or only images. For an image-only scan, use an OCR process; the PDFBox command documented here should not be presented as OCR.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

After OCR has created searchable text or a text layer, an extractor can read that text, but the OCR step itself is separate. Check extracted output against the source pages, especially for small print, columns, tables, handwriting, or skewed scans.

Options and output quality

Page ranges, ordering, and encoding

The PDFBox 3.0 command-line documentation describes page-range controls and sorting options, and identifies UTF-8 as the default encoding. These can matter when processing only selected pages, mixed layouts, or text with non-English characters. Verify the precise option names and accepted values for the release you installed; do not copy options from a different major version without checking.

Expect extracted text, not a perfect page replica

PDF text extraction yields text content, not a faithful reproduction of page layout. Multi-column pages, positioned labels, footnotes, tables, and unusual reading order can produce text that needs cleanup. Compare important results with the PDF itself instead of assuming the output’s line sequence matches visual reading order.

Troubleshooting common problems

  • java is not recognized: PowerShell cannot locate the Java executable. Install or configure Java and ensure its command is available to the shell, then verify with java -version.
  • The JAR cannot be opened: Check the exact JAR filename and current directory, or pass a full path. Confirm the JAR exists with Test-Path -LiteralPath.
  • PDFBox reports an unknown command or option: The command may belong to another major release. Use export:text for the documented 3.x form and ExtractText for the 2.x form; check the installed JAR’s own documentation.
  • No output file appears: Confirm the input PDF path, output directory, and write permissions. Read Java/PDFBox error output rather than treating the absence of a file as successful extraction.
  • Output exists but is empty or incomplete: The document may be image-only, protected, or structured in a way that extracts poorly. Check whether its text is selectable, verify any required password option against the release documentation, and inspect pages with complex layout.
  • Text order looks wrong: PDF page placement does not always correspond to a simple reading sequence. Check the documented sorting controls for your PDFBox version and manually validate layout-sensitive content.
  • Characters display incorrectly: Check the output encoding and how the text viewer reads it. PDFBox 3.x documents UTF-8 as the default, but verify version-specific behavior and the actual file.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server, not a PDF text extractor. If your real task is to capture a webpage rather than extract text from a PDF, one GET request can return a screenshot or PDF:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

See the ScreenshotNeo API documentation. It accepts cookie or consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing status. Its MCP server gives AI agents tools for screenshots, page information, and PDF capture. The free plan includes 1,000 shots per month with no card; paid plans start at $5 for 3,000 shots. For PDF text extraction, continue using an extractor such as PDFBox. Learn about ScreenshotNeo or sign up free for 1,000 screenshots a month with no card.

Frequently Asked Questions

Can PowerShell extract PDF text without installing another utility?

Not with Get-Content alone. It reads text files; a PDF parser or extractor must first interpret the PDF.

Can PDFBox extract text from a password-protected PDF?

The PDFBox 3.x command-line documentation lists a password option. Check the option syntax for the installed release and ensure you are authorized to access the document.

Can I extract only certain pages?

PDFBox 3.x documents page-range controls. Use the exact option syntax from the documentation matching your installed release.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Read next

Recommended PC Tool
Recommended PC Tool
Outdated Drivers Are Slowing You DownFree scan - exact matches
Windows Errors? Fix Them Before They SpreadFree repair scan

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.