DriversRecommendedOutdated drivers can make a good PC feel brokenScan driver issues before chasing fixes manually.Scan NowOctober DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsClean PCRecommendedOne scan can reveal what keeps slowing WindowsLook for cleanup and repair opportunities.Run Scan×
Skip to content
Apache PDFBox

How to Convert HTML to PDF with PDFBox (Java Guide)

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

PDFBox does not parse HTML or reproduce a browser page by itself. The reliable Java workflow is to use an HTML/CSS renderer such as OpenHTMLtoPDF to lay out the markup, configure its PDFBox integration, and then use Apache PDFBox APIs for any PDF-specific processing. This distinction determines which dependency you add, which CSS you can use, and how you troubleshoot missing content.

Can PDFBox convert HTML to PDF?

Not on its own. Apache PDFBox is an open-source Java library for creating and working with PDF documents. Its feature set covers PDF objects, pages, fonts, images, metadata, signing, extraction and rendering of existing PDF pages; it is not an HTML parser or browser layout engine.

For HTML input, add an HTML/CSS renderer. OpenHTMLtoPDF renders a supported subset of well-formed XML/XHTML (and some HTML5) with CSS into PDF or images. Its PDFBox integration writes the resulting PDF through PDFBox. In practical terms:

  • OpenHTMLtoPDF parses markup and applies supported CSS rules.
  • PDFBox is the PDF backend and the library you use for subsequent PDF operations.
  • Your application supplies valid markup, a base URL for relative resources, fonts, output streams and lifecycle management.

This is different from PDFBox’s PDFRenderer, which rasterizes an existing PDF page. It does not lay out HTML into a new PDF.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Choose the dependency that matches your PDFBox major version

First inspect the application’s dependency tree and identify whether it uses PDFBox 2 or PDFBox 3. OpenHTMLtoPDF publishes separate integration artifacts for those majors. Do not assume that an artifact name containing “pdfbox” is an Apache module; these are OpenHTMLtoPDF components.

Application stack PDFBox coordinate OpenHTMLtoPDF integration coordinate Notes
PDFBox 3 org.apache.pdfbox:pdfbox:3.0.8 in the current getting-started example io.github.openhtmltopdf:openhtmltopdf-pdfbox Use a renderer release declared compatible with PDFBox 3.
PDFBox 2 Your existing PDFBox 2.x version com.openhtmltopdf:openhtmltopdf-pdfbox Use the PDFBox 2 integration, not the PDFBox 3 artifact.

The project homepage reported PDFBox 2.0.37, released July 15, 2026, and PDFBox 3.0.8, released July 11, 2026. Release numbers change, so verify the current compatible versions immediately before updating a production build.

Maven coordinates for PDFBox 3

Add the OpenHTMLtoPDF PDFBox 3 integration and a compatible version of that library. The exact OpenHTMLtoPDF version is intentionally a build decision: select one whose dependency metadata matches the PDFBox major version already used by your application.

<dependency>
  <groupId>io.github.openhtmltopdf</groupId>
  <artifactId>openhtmltopdf-pdfbox</artifactId>
  <version>REPLACE_WITH_COMPATIBLE_VERSION</version>
</dependency>
<dependency>
  <groupId>org.apache.pdfbox</groupId>
  <artifactId>pdfbox</artifactId>
  <version>3.0.8</version>
</dependency>

Maven coordinates for PDFBox 2

<dependency>
  <groupId>com.openhtmltopdf</groupId>
  <artifactId>openhtmltopdf-pdfbox</artifactId>
  <version>REPLACE_WITH_COMPATIBLE_VERSION</version>
</dependency>
<dependency>
  <groupId>org.apache.pdfbox</groupId>
  <artifactId>pdfbox</artifactId>
  <version>YOUR_PDFBOX_2_VERSION</version>
</dependency>

Keep one PDFBox major version in the runtime class path. If another dependency pulls in a second major, resolve it in your build tool rather than relying on whichever jar happens to load first.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Minimal Java conversion example

The following program accepts an HTML file and writes a PDF. It uses the OpenHTMLtoPDF PDFBox integration; the same application pattern works with either supported PDFBox major when the matching artifact is on the class path.

import com.openhtmltopdf.pdfboxout.PdfRendererBuilder;

import java.io.OutputStream;
import java.nio.file.Files;
import java.nio.file.Path;
import java.nio.file.Paths;

public final class HtmlToPdf {
    public static void main(String[] args) throws Exception {
        if (args.length != 2) {
            System.err.println("Usage: java HtmlToPdf input.html output.pdf");
            System.exit(2);
        }

        Path htmlPath = Paths.get(args[0]).toAbsolutePath().normalize();
        Path pdfPath = Paths.get(args[1]).toAbsolutePath().normalize();
        Path parent = pdfPath.getParent();
        if (parent != null) {
            Files.createDirectories(parent);
        }

        String html = Files.readString(htmlPath);
        String baseUri = htmlPath.getParent().toUri().toString();

        try (OutputStream output = Files.newOutputStream(pdfPath)) {
            PdfRendererBuilder builder = new PdfRendererBuilder();
            builder.withHtmlContent(html, baseUri);
            builder.toStream(output);
            builder.run();
        }

        System.out.println("Wrote " + pdfPath);
    }
}

Compile this class with the renderer integration and its transitive dependencies, then run:

java HtmlToPdf invoice.html build/invoice.pdf

withHtmlContent is useful when the HTML is already in memory. The second argument is important: it gives relative links such as css/print.css, images/logo.png and web fonts a resolvable base. If you have a URL instead, use the renderer’s URI-based input method and make sure the process can access every referenced resource.

Registering a local font

PDF output is only as stable as the fonts available to the renderer. For a font file that is not installed on the server, register it explicitly with the builder and use the same family name in CSS. The exact font provider syntax can vary by OpenHTMLtoPDF release, so compile against the version selected in your build and verify that the font is embedded in representative output. Do not assume a developer workstation’s installed fonts exist in a container or CI runner.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Prepare HTML and CSS for a print renderer

OpenHTMLtoPDF describes its support as a reasonable subset of well-formed XML/XHTML, HTML5 features and CSS 2.1 (with some later standards). It explicitly is not a browser and does not run JavaScript. A page that looks correct in Chrome can therefore require a print-specific template.

Use deterministic, well-formed markup

  • Close elements and quote attributes. Treat malformed HTML as a conversion error, not something the renderer must repair like a browser.
  • Include a character encoding declaration and write the source as UTF-8.
  • Use a stable base URI for relative images, stylesheets and fonts.
  • Prefer simple block, inline, table and positioned layouts that the renderer supports.

Do not depend on browser-only behavior

  • JavaScript-generated text, charts or menus will not execute during conversion.
  • CSS flexbox and grid are not implemented as they are in modern browsers; redesign those sections with supported blocks or tables.
  • Responsive breakpoints aimed at a changing viewport may not produce the intended paper layout.
  • Browser APIs, canvas scripts, client-side data fetching and lazy-loading code cannot be treated as available content sources.

Design pagination deliberately

Use print-oriented CSS such as @page, explicit page size and margins, and page-break rules where supported. Test long tables, headings at the bottom of a page, images near page boundaries, nested lists and repeated headers. A renderer can produce a valid PDF while still placing a heading, row or image in an undesirable position.

Use PDFBox after conversion when you need PDF operations

Once the renderer has produced the file, load it with PDFBox for tasks such as merging documents, adding metadata, extracting text, applying a signature or rendering pages to images. Keep conversion and post-processing as separate stages so a layout failure is distinguishable from a PDF manipulation failure.

Close every PDDocument with try-with-resources. PDFBox documents that only one thread may access a single document at a time. You can process separate document instances independently, but do not share one mutable PDDocument among concurrent workers.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
import org.apache.pdfbox.pdmodel.PDDocument;

import java.nio.file.Path;

public class InspectPdf {
    public static void main(String[] args) throws Exception {
        Path pdf = Path.of(args[0]);
        try (PDDocument document = PDDocument.load(pdf.toFile())) {
            System.out.println("Pages: " + document.getNumberOfPages());
            document.getDocumentInformation().setTitle("Generated document");
            document.save(pdf.toFile());
        }
    }
}

For PDFBox 3, check any API changes when adapting older examples. In particular, historical PDFBox 2 migration notes explain that PDPage.convertToImage and PDFImageWriter were removed in 2.0.0; use PDFRenderer for rendering an existing PDF page to an image instead.

Validate output before shipping

No HTML-to-PDF renderer can guarantee browser-equivalent output for arbitrary sites. Build a fixture set that represents your actual documents and inspect both the PDF visually and its extracted content.

  • Fonts: Latin, non-Latin, bold/italic variants, fallback glyphs and embedded-font behavior.
  • Images: relative and absolute URLs, transparency, high-resolution images, missing resources and very large files.
  • Pagination: first and last page, long tables, forced breaks, headers/footers, widows and orphans.
  • Content: special characters, links, lists, nested tables and generated values supplied by the server.
  • Operations: output can be opened by your target PDF viewers, post-processing closes documents, and concurrent jobs use independent instances.

Keep representative PDFs as regression fixtures. Compare page count, extracted text and important visual regions after dependency upgrades; version changes can alter line wrapping or pagination without throwing an exception.

Troubleshooting common failures

Symptom Likely cause Fix
ClassNotFoundException: PdfRendererBuilder The OpenHTMLtoPDF integration is missing or the wrong artifact is selected. Add the matching io.github.openhtmltopdf artifact for PDFBox 3 or com.openhtmltopdf artifact for PDFBox 2, then refresh the dependency tree.
Linkage errors involving PDFBox classes PDFBox 2 and PDFBox 3 jars are mixed. Inspect transitive dependencies and align every renderer and PDFBox dependency to one major version.
Blank or incomplete dynamic content The source relies on JavaScript, client-side requests or browser APIs. Render the data server-side before conversion, provide a static print template, or use a browser automation workflow when full JavaScript execution is a requirement.
CSS layout collapses The template depends on flexbox, grid or another unsupported modern standard. Replace the layout with supported block/table/print CSS and test the exact document.
Images or styles are missing Relative URLs have no correct base URI, resources require authentication, or the process cannot reach them. Pass a base URI, use accessible resource URLs, provide required request configuration, and log resource loading during development.
Unexpected glyphs or substituted fonts The required font is absent or not registered. Install or register the font, verify the CSS family name, and inspect embedding in the resulting PDF.
Out-of-memory errors A large document, high-resolution images or retained PDF/rendering objects consume substantial memory. Reduce image resolution, avoid retaining unnecessary objects, process documents in bounded batches and use PDFBox scratch-file loading options where appropriate.
Intermittent corruption in a web service One PDDocument is being accessed by multiple threads or streams are closed inconsistently. Create separate document instances per job, serialize access to each instance and close every stream/document deterministically.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Performance, reliability and deployment notes

Measure your own workload instead of relying on a universal pages-per-second figure. HTML complexity, image dimensions, font files, page count, JVM memory and output resolution all affect conversion time and memory use.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Reuse immutable HTML templates, but create a fresh renderer and output stream for each conversion.
  • Cache stable assets such as stylesheets and fonts at the application layer when licensing and freshness rules allow.
  • Set job timeouts around remote resource access and fail clearly when a required asset cannot be loaded.
  • Limit input size and image dimensions for untrusted HTML. Do not allow arbitrary file or network access without a deliberate security policy.
  • Store failed input and diagnostic logs separately from successful PDFs so a malformed document can be reproduced.

Or skip the browser setup

If your goal is to capture a hosted web page rather than generate a PDF from an in-memory HTML string, ScreenshotNeo provides a website screenshot API that can return PNG, JPEG, WebP or PDF. It handles the browser session for you. Before capture, it accepts cookie/consent banners and removes more than 60 known consent platforms, newsletter popups and chat widgets; each cleanup step can be disabled. Bot checks/CAPTCHAs, blank pages, timeouts, failed loads and cache hits are not billed, and the response identifies the result with X-Page-Verdict and X-Billed headers. It also offers an MCP server for AI agents, with take_screenshot, get_page_info and capture_pdf tools.

For the full parameter list and PDF options, see the ScreenshotNeo documentation. A basic request looks like this:

curl -G "https://api.screenshotneo.com/v1/shot" 
  -d access_key=YOUR_API_KEY 
  --data-urlencode url=https://example.com 
  -o page.webp

The same endpoint can be called from Python:

import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://example.com"},
    timeout=90,
)
r.raise_for_status()
open("page.webp", "wb").write(r.content)

Or Node.js:

const q = new URLSearchParams({
  access_key: 'YOUR_API_KEY',
  url: 'https://example.com'
});
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`${res.status} ${res.statusText}`);
const fs = await import('node:fs/promises');
await fs.writeFile('page.webp', Buffer.from(await res.arrayBuffer()));

ScreenshotNeo is not a substitute when your Java service must lay out arbitrary private HTML in-process; OpenHTMLtoPDF remains the relevant renderer for that case. It is useful when the source is a reachable URL and you want a managed browser capture or PDF, without building and maintaining browser automation. One thousand screenshots per month are free with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account to get started.

Frequently Asked Questions

Can I feed a local HTML file directly to ScreenshotNeo?

ScreenshotNeo captures a URL. Host the page at a reachable address or keep the conversion local with the Java renderer when the HTML must remain on your machine.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Should I upgrade PDFBox before adding an HTML renderer?

First identify the major version already used by your application, then choose the corresponding OpenHTMLtoPDF integration. Upgrade only after testing your fixture documents and resolving any API or layout changes.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Read next

Recommended PC Tool
Recommended PC Tool
Windows Errors? Fix Them Before They SpreadFree repair scan
Outdated Drivers Are Slowing You DownFree scan - exact matches

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.