Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
Yes—but PDFBox’s normal save is structural compression, not a complete “Optimize PDF” operation. In PDFBox 3.x, loading a document and saving it to a different file uses compressed saving by default. That may help a little, or may not reduce the file at all. Large reductions usually require inspecting and selectively recompressing or downsampling embedded images.
1. Add the current PDFBox dependency
This example targets Apache PDFBox 3.0.8, listed by the project as the current 3.x release on July 11, 2026. PDFBox 3.x requires Java 11 or newer; older 2.x examples use different loading APIs.
Use the dependency documented at pdfbox.apache.org/3.0/getting-started.html:
Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minuteWindows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstall<dependency>
<groupId>org.apache.pdfbox</groupId>
<artifactId>pdfbox</artifactId>
<version>3.0.8</version>
</dependency>
2. Try a compressed resave first
For a PDF that is only inefficiently structured, this is the safest first test:
#1 Best Overall
- EDIT text, images & designs in PDF documents. ORGANIZE PDFs. Convert PDFs to Word, Excel & ePub.
- READ and Comment PDFs – Intuitive reading modes & document commenting and mark up.
- CREATE, COMBINE, SCAN and COMPRESS PDFs
- FILL forms & Digitally Sign PDFs. PROTECT and Encrypt PDFs
- LIFETIME License for 1 Windows PC or Laptop. 5GB MobiDrive Cloud Storage Included.
import java.io.File;
import java.io.IOException;
import org.apache.pdfbox.Loader;
import org.apache.pdfbox.pdmodel.PDDocument;
public final class ResavePdf {
public static void main(String[] args) throws IOException {
File input = new File("input.pdf");
File output = new File("compressed.pdf");
try (PDDocument document = Loader.loadPDF(input)) {
document.save(output);
}
}
}
PDFBox 3.0 uses compressed saving by default. The migration notes at pdfbox.apache.org/3.0/migration.html also warn that the destination must be a different file. Never save over the source while it is being read; a failed write can destroy the only copy.
You can make the choice explicit:
import org.apache.pdfbox.pdfwriter.compress.CompressParameters;
try (PDDocument document = Loader.loadPDF(new File("input.pdf"))) {
document.save(
new File("compressed.pdf"),
CompressParameters.DEFAULT_COMPRESSION
);
}
DEFAULT_COMPRESSION controls PDFBox’s PDF-object and stream writing. It is not an image-quality setting. NO_COMPRESSION does the opposite and is mainly useful for special compatibility or PDF/A-1b workflows; it is not a size-reduction technique.
Measure instead of assuming
Compare the byte lengths of the original and output. A rewrite can be smaller, unchanged, or larger because of object layout, cross-reference data, metadata, fonts, or existing image filters. A resave does not automatically downsample images, convert photographic PNGs to JPEG, subset fonts, or remove every unused object.
3. Find out whether images are the problem
Image dimensions are a useful first signal, although they are not a complete byte-level profiler: color space, bit depth, masks, filters, duplication, and reuse also matter.
import java.io.File;
import java.io.IOException;
import org.apache.pdfbox.Loader;
import org.apache.pdfbox.cos.COSName;
import org.apache.pdfbox.pdmodel.PDDocument;
import org.apache.pdfbox.pdmodel.PDPage;
import org.apache.pdfbox.pdmodel.PDResources;
import org.apache.pdfbox.pdmodel.graphics.PDXObject;
import org.apache.pdfbox.pdmodel.graphics.image.PDImageXObject;
try (PDDocument document = Loader.loadPDF(new File("input.pdf"))) {
for (PDPage page : document.getPages()) {
PDResources resources = page.getResources();
if (resources == null) {
continue;
}
for (COSName name : resources.getXObjectNames()) {
PDXObject xObject = resources.getXObject(name);
if (xObject instanceof PDImageXObject image) {
System.out.printf(
"image=%s, width=%d, height=%d%n",
name.getName(), image.getWidth(), image.getHeight()
);
}
}
}
}
A small image drawn on a page can still contain millions of pixels. Conversely, the same image XObject may be referenced by many pages. A production optimizer should track object identity and avoid decoding and replacing one shared image repeatedly.
4. Recompress suitable images as JPEG
JPEG is generally appropriate for photographs and continuous-tone color scans. PDFBox documents JPEG creation from a BufferedImage, with a floating-point quality value and an optional DPI metadata value, in JPEGFactory.
import java.awt.image.BufferedImage;
import java.io.File;
import java.io.IOException;
import org.apache.pdfbox.Loader;
import org.apache.pdfbox.cos.COSName;
import org.apache.pdfbox.pdmodel.PDDocument;
import org.apache.pdfbox.pdmodel.PDPage;
import org.apache.pdfbox.pdmodel.PDResources;
import org.apache.pdfbox.pdmodel.graphics.PDXObject;
import org.apache.pdfbox.pdmodel.graphics.image.JPEGFactory;
import org.apache.pdfbox.pdmodel.graphics.image.PDImageXObject;
public final class RecompressImages {
public static void main(String[] args) throws IOException {
try (PDDocument document = Loader.loadPDF(new File("input.pdf"))) {
for (PDPage page : document.getPages()) {
PDResources resources = page.getResources();
if (resources == null) {
continue;
}
for (COSName name : resources.getXObjectNames()) {
PDXObject object = resources.getXObject(name);
if (!(object instanceof PDImageXObject oldImage)) {
continue;
}
BufferedImage image = oldImage.getImage();
PDImageXObject replacement =
JPEGFactory.createFromImage(document, image, 0.75f);
resources.put(name, replacement);
}
}
document.save(new File("compressed-images.pdf"));
}
}
}
This is deliberately a starting point, not a universal optimizer. It converts every encountered image, including images for which JPEG is a poor choice. Text scans, line art, screenshots, barcodes, transparency masks, indexed colors, and already-compressed JPEGs can become worse or even larger. Decoding and re-encoding an existing JPEG also adds another lossy generation. When the original JPEG bytes are acceptable, PDFBox’s JPEGFactory.createFromStream(...) can embed existing JPEG data without recompressing it.
Recommended Free Tools
The displayed placement is defined by the page content stream; replacing the XObject does not automatically change its on-page size. Replacing resources can also affect masks, color profiles, and transparency, so render and inspect the result.
Rank #2
- Edit PDFs with Ease. Modify text, images, and layouts directly within your PDF documents.
- Convert & Organize. Export PDFs to Word, Excel, or ePub, and organize files with ease.
- Read & Annotate. Enjoy intuitive reading modes and powerful tools to comment, highlight, and mark up PDFs.
- Create & Manage PDFs. Create new PDFs, combine multiple files, scan documents, and compress for easy sharing.
- Fill & Sign Forms. Complete forms and digitally sign documents with secure e-signature tools.
5. Downsample before encoding
Removing unnecessary pixels often saves more than changing JPEG quality alone. The following helper scales down only when the source exceeds the selected bounds:
import java.awt.Graphics2D;
import java.awt.RenderingHints;
import java.awt.image.BufferedImage;
static BufferedImage scaleToMaxDimension(
BufferedImage source, int maxWidth, int maxHeight) {
double scale = Math.min(1.0, Math.min(
(double) maxWidth / source.getWidth(),
(double) maxHeight / source.getHeight()));
if (scale >= 1.0) {
return source;
}
int width = Math.max(1, (int) Math.round(source.getWidth() * scale));
int height = Math.max(1, (int) Math.round(source.getHeight() * scale));
BufferedImage resized =
new BufferedImage(width, height, BufferedImage.TYPE_INT_RGB);
Graphics2D graphics = resized.createGraphics();
try {
graphics.setRenderingHint(RenderingHints.KEY_INTERPOLATION,
RenderingHints.VALUE_INTERPOLATION_BICUBIC);
graphics.setRenderingHint(RenderingHints.KEY_RENDERING,
RenderingHints.VALUE_RENDER_QUALITY);
graphics.drawImage(source, 0, 0, width, height, null);
} finally {
graphics.dispose();
}
return resized;
}
Use the resized image with JPEGFactory.createFromImage(document, resized, 0.75f). Values such as 2,000 pixels and quality 0.75 are examples, not universal settings. Choose them from the intended output:
| Content or use | Starting approach |
|---|---|
| Screen reading | Moderate dimensions and JPEG quality around 0.65–0.80, then check small text. |
| Office printing | Keep more pixels and use higher quality. |
| Photographs | JPEG is usually suitable. |
| Logos, diagrams, screenshots, fine line art | Prefer lossless encoding when artifacts would be visible. |
| Archival scans | Avoid uncontrolled lossy conversion; follow the required PDF/A and preservation workflow. |
| Text-only monochrome scans | Consider true 1-bit Group 4 compression after verifying readability. |
The DPI argument in the JPEG factory records metadata; it does not reduce pixels or file size by itself.
6. Pick an image format by content
PDFBox exposes JPEG, lossless, and CCITT factories, documented in the PDImageXObject factory documentation.
- JPEG: good for photographs and continuous-tone scans; introduces lossy artifacts.
- LosslessFactory: suitable for diagrams, screenshots, logos, text-like artwork, and transparency-sensitive images.
- CCITTFactory: suitable for genuinely black-and-white scans. Thresholding grayscale pages can erase faint characters, stamps, pencil marks, or signatures.
Do not blindly convert every image to RGB JPEG. JPEG does not preserve alpha transparency; use a lossless strategy or deliberately flatten against a known background. Barcodes and fine rules deserve special visual inspection.
7. Make newly generated PDFs small from the start
For new documents, resize source images before embedding, choose the factory that matches the content, and reuse one image XObject when the same image appears repeatedly. Avoid placing a 6,000-pixel source image when the page displays it at a small physical size.
The official image example uses PDImageXObject.createFromFile(...) and doc.save(...); see ImageToPDF.java. That convenience method does not guarantee the smallest encoding, so preprocess photographs and scans when size matters. Normal PDFBox saving still supplies structural compression.
8. Handle metadata and other content carefully
Removing document information, XMP metadata, attachments, annotations, form fields, thumbnails, duplicate resources, or unused pages can save space, but these objects may carry legal, archival, accessibility, or workflow value. XMP is a distinct metadata resource, as illustrated by the command-line documentation at pdfbox.apache.org/3.0/commandline.html. Do not delete objects merely because they look unused; indirect references and appearance streams can matter.
Rank #3
- EVERY PDF TOOL UNLOCKED - 30+ tools in one app: edit text and images, convert, merge, split, compress, sign, OCR, redact, watermark, batch process, and more. No feature gates, no upsells, nothing held back.
- PAY ONCE, OWN FOREVER — A one-time purchase, not a subscription. Other apps runs $240/year — Scrivar is yours for life, with free updates included.
- UNLIMITED eSIGN, BUILT IN — Send contracts and forms for signature and track every step. Recipients sign in their browser with no account or app needed. Replace DocuSign and save hundreds a year.
- PC, MAC, AND WEB — Install on any Win 10/11 PC or macOS 11+ Mac (Intel or Apple Silicon), or work in your browser at scrivar.com. Same tools, same account, everywhere you work.
- OCR + FULL OFFICE CONVERSION — Turn scanned documents into searchable, selectable text, and convert PDFs to and from Word, Excel, and PowerPoint with formatting kept intact.
9. Protect signatures, encryption, forms, and PDF/A requirements
Digital signatures
A normal full save rewrites the document and will generally invalidate existing signatures. Preserve the signed original and treat optimization as a pre-signing operation unless a signature-aware incremental or external-signing workflow has been designed and verified. PDFBox documents separate incremental-save and external-signing APIs in PDDocument.
Encryption
An encrypted input may require a password. Changing security settings changes the document’s usable state; follow the save and encryption constraints in the PDDocument API, and do not assume the same document object remains suitable for further work after activating encryption.
PDF/A and interactive content
Recompression choices can violate PDF/A requirements involving metadata, fonts, transparency, and image encoding. Forms, widgets, annotations, tagged accessibility content, hyperlinks, and appearance streams also need regression testing after resource replacement.
The Tool Desk
Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →10. Command-line limitations
The documented PDFBox command-line distribution provides rendering, image export, splitting, merging, and inspection utilities, but no general-purpose “optimize existing PDF” command. For example:
java -jar pdfbox-app-3.y.z.jar export:images -i=input.pdf
java -jar pdfbox-app-3.y.z.jar decode input.pdf output-decoded.pdf
decode decompresses PDF streams for inspection; it does not make a PDF smaller. A repeatable image optimization policy therefore requires Java code or another dedicated optimizer.
11. Validate every rewritten file
- Compare byte size and confirm the output opens in more than one PDF viewer.
- Search and select text; test extraction if downstream systems depend on it.
- Render representative pages and inspect photographs, small text, line art, transparency, and barcodes.
- Print pages and test hyperlinks, navigation, annotations, widgets, and form editing.
- Check embedded files, accessibility tags, and metadata that your workflow requires.
- Recheck digital signatures and PDF/A conformance where applicable.
12. Troubleshoot common results
The output is larger
The source may already use efficient filters, or the rewrite may have added less-compact structures. Decode and inspect image dimensions and filters, compare a resave-only copy, then test downsampling separately from quality changes. Avoid converting already-efficient JPEGs without a measurable reason.
The PDF is blurry
Increase pixel dimensions or JPEG quality, use lossless encoding for text and diagrams, and avoid repeated lossy recompression. Suitable monochrome scans may benefit from CCITT instead.
Memory use is excessive
PDFBox 3.x uses incremental parsing, but iterating through every page and decoding large images still consumes memory. Process documents individually, do not retain all BufferedImage objects, downsample in a controlled way, use temporary files, and set an appropriate JVM heap. Very large scans may need external or streaming image preprocessing.
The Bottom Line
Start with a separate-output Loader.loadPDF(...) and save(...) pass. If that is not enough, identify the largest image resources, downsample selectively, and choose JPEG, lossless, or CCITT according to the image content. Keep the original, and validate the rewritten PDF’s appearance, behavior, signatures, and conformance before replacing it.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

