Do these 3 things before closing this tab:
1Fix the driver behind crashes, sound loss and screen glitches2Clear out junk files and repair common Windows errors3Scan for outdated or missing drivers - takes under a minuteSome links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
The most direct way to convert a web page to PDF in Java is to use an HTML renderer rather than PDFBox alone. For controlled HTML or XHTML, iText pdfHTML provides HtmlConverter.convertToPdf(...) and a base-URI setting for relative assets. OpenHTMLToPDF is a pure-Java, LGPL-compatible choice for a reasonable XHTML/CSS subset. If the source is a modern, JavaScript-heavy page, use a browser-backed renderer such as Flying Saucer’s flying-saucer-chrome-pdf path, or call a screenshot/PDF service.
Choose the renderer before writing code
“Java HTML to PDF” can mean two different jobs:
- Render content you control: a string or template containing stable HTML, CSS, images and fonts.
- Capture an arbitrary live web page: a URL that may depend on JavaScript, responsive layout, cookies, consent dialogs, lazy loading or browser APIs.
Those inputs need different tools. A Java library that parses XHTML is not automatically a browser. Decide whether you need browser fidelity, accessibility or PDF/A conformance, and whether a commercial license is acceptable.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
| Option | Best fit | Important limits or requirements |
|---|---|---|
| iText pdfHTML | Controlled HTML/CSS with a short, documented conversion API | Check the iText license for your version and deployment model; configure a base URI for relative resources. |
| OpenHTMLToPDF | Pure-Java XHTML/HTML and CSS 2.1-style documents | Not a web browser; it does not execute JavaScript and does not implement many modern standards, including flex and grid. |
| Flying Saucer | XHTML/CSS rendering, or browser-oriented output through its Chrome artifact | The flying-saucer-chrome-pdf artifact delegates to chrome-headless-shell. Verify the Java baseline for the version you select. |
| Apache PDFBox | Creating, editing, rendering or post-processing PDF files | It is PDF infrastructure, not a complete browser-grade HTML/CSS/JavaScript converter by itself. |
For an HTML string to PDF Java workflow, start with iText or OpenHTMLToPDF. For a JavaScript web page to PDF, use a Chromium-backed route or a service that runs a browser.
Convert HTML to PDF with iText pdfHTML
iText’s documented minimal path accepts HTML and writes a PDF to a stream or file. This example converts a string and creates the output file:
import com.itextpdf.html2pdf.HtmlConverter;
import java.io.FileOutputStream;
import java.io.IOException;
public class HtmlToPdf {
public static void createPdf(String html, String destination) throws IOException {
HtmlConverter.convertToPdf(
html,
new FileOutputStream(destination)
);
}
public static void main(String[] args) throws IOException {
String html = "<!doctype html>"
+ "<html><head><meta charset='UTF-8'>"
+ "<style>body{font-family:Arial,sans-serif}h1{color:#174ea6}</style>"
+ "</head><body><h1>Invoice</h1>"
+ "<p>Generated from HTML in Java.</p></body></html>";
createPdf(html, "output.pdf");
}
}
The API also accepts a File or InputStream as input and can write to an output stream, file, PdfWriter or PdfDocument. Keep the output stream in a try-with-resources block in production so it is closed even when conversion fails.
Resolve images, CSS and fonts with a base URI
Relative URLs such as images/logo.png have no meaning unless the renderer knows the document’s location. Set a base URI explicitly:
Recommended Free Tools
import com.itextpdf.html2pdf.ConverterProperties;
import com.itextpdf.html2pdf.HtmlConverter;
import java.io.FileOutputStream;
import java.io.IOException;
public class HtmlWithAssets {
public static void createPdf(String baseUri, String html, String destination)
throws IOException {
ConverterProperties properties = new ConverterProperties();
properties.setBaseUri(baseUri);
try (FileOutputStream output = new FileOutputStream(destination)) {
HtmlConverter.convertToPdf(html, output, properties);
}
}
}
Pass a directory or URL that makes every relative stylesheet, image and font resolvable. If an asset is remote, make sure the Java process can reach it and that the server permits the request. For repeatable builds, localize the assets and use a fixed base directory instead of relying on a changing website.
Use complete, valid markup
Prefer a declared character set, closed tags, absolute dimensions for critical images, and print-oriented CSS. A renderer may tolerate malformed HTML differently from a browser. Test page breaks, long tables, overflowing code blocks and missing fonts with the actual documents your application produces.
When OpenHTMLToPDF is the better Java choice
OpenHTMLToPDF is a pure-Java library that renders a reasonable subset of well-formed XML/XHTML (and some HTML5) with CSS 2.1 and later standards, producing PDF or images. It uses Apache PDFBox rather than iText and documents support for PDF/A, accessible PDF output, SVG, MathML and font fallback.
Rank #2
Its boundary is important: the project explicitly says it is not a web browser. It does not run JavaScript and does not implement many modern layout features such as flex and grid. Therefore it is a good match for server-rendered templates, email-like documents and controlled reports, but not for a site whose content appears only after client-side JavaScript executes.
Quick wins for a faster PC:
Repair Windows errors before they cause bigger problemsFix Now →Scan for outdated or missing drivers - takes under a minuteDriver Scan →Clear out junk files and repair common Windows errorsFree Scan →Choose it when a pure-Java, open-source-compatible deployment matters and your markup fits its supported subset. Before switching, create fixtures for your real CSS, SVGs, fonts and page-break rules; “HTML5” support does not mean browser-equivalent rendering.
Flying Saucer and browser-oriented rendering
Flying Saucer provides pure-Java XML/XHTML and CSS 2.1 rendering to PDF and images. Its project lists both org.xhtmlrenderer:flying-saucer-pdf and org.xhtmlrenderer:flying-saucer-chrome-pdf. The latter delegates PDF generation to chrome-headless-shell, making it the browser-oriented path when ordinary XHTML rendering cannot reproduce a modern page.
Check the selected artifact’s requirements before deployment. The documented baselines are Java 11 or later for versions from 9.5.0, Java 17 or later for 9.6.0, and Java 21 or later for 10.0.0. Verify the exact version and its transitive dependencies in your build, because the Java requirement is tied to the artifact version rather than to Flying Saucer as a whole.
A Chrome-backed process generally costs more operationally than an in-process parser: you must package or provision the browser, manage startup and concurrency, and isolate untrusted pages. In return, browser layout and script execution are closer to what a user sees.
Why PDFBox alone does not convert arbitrary pages
Apache describes PDFBox as an open-source Java tool for working with PDF documents. It is useful for creating and manipulating PDFs, rendering pages and post-processing a PDF generated by another renderer. It does not parse a modern web page, run JavaScript or implement browser layout by itself.
A practical pipeline is therefore HTML renderer first, PDFBox second: generate the document with iText, OpenHTMLToPDF or a browser-backed renderer, then use PDFBox for metadata, merging, stamping, splitting or validation.
Handling a live JavaScript web page
If the page is a URL rather than your own HTML, identify what happens before capture:
- Does JavaScript insert the main content?
- Are images lazy-loaded only after scrolling?
- Does the page require cookies, authentication, a user agent, timezone or geolocation?
- Do consent banners, chat widgets or bot checks obscure the page?
- Are the layout and fonts dependent on flex, grid or other browser features?
Non-browser renderers will not execute scripts or reproduce all browser CSS. A Chrome-backed renderer can, but it introduces browser-process management. For untrusted URLs, apply network restrictions, timeouts, output-size limits and a resource policy; never allow arbitrary page content to access internal services from a privileged network.
Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server for developers. One GET request can return a clean PNG, JPEG, WebP or PDF. Before capture it accepts the cookie or consent banner like a visitor and removes more than 60 known consent platforms, newsletter popups and chat widgets; each cleanup step can be turned off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads and cache hits are not billed, and the response identifies the result with X-Page-Verdict and X-Billed headers.
For a one-call capture, use the API base shown in the ScreenshotNeo documentation:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
The same request from Java can be made with java.net.http.HttpClient:
import java.net.URI;
import java.net.http.HttpClient;
import java.net.http.HttpRequest;
import java.net.http.HttpResponse;
import java.nio.file.Files;
import java.nio.file.Path;
public class ScreenshotNeoJava {
public static void main(String[] args) throws Exception {
String endpoint = "https://api.screenshotneo.com/v1/shot"
+ "?access_key=YOUR_API_KEY"
+ "&url=https%3A%2F%2Fstripe.com";
HttpRequest request = HttpRequest.newBuilder(URI.create(endpoint)).GET().build();
HttpResponse<byte[]> response = HttpClient.newHttpClient()
.send(request, HttpResponse.BodyHandlers.ofByteArray());
if (response.statusCode() / 100 != 2) {
throw new IllegalStateException("Screenshot request failed: " + response.statusCode());
}
Files.write(Path.of("shot.webp"), response.body());
}
}
ScreenshotNeo also provides take_screenshot, get_page_info and capture_pdf tools through an MCP server for Claude, Cursor and other MCP clients. Its options include full-page capture with lazy images loaded, CSS-selector element capture, dark mode, device presets and custom viewports, retina scale, PDF paper size/margins/orientation/page ranges, custom CSS and JavaScript, click-before-capture, selector hiding, selector or delay or network-idle waits, request and resource blocking, headers, cookies, user agent, authorization, timezone, geolocation, transparent backgrounds, resizing, configurable-TTL caching, signed links, asynchronous jobs with signed webhooks, bulk capture of up to 100 URLs per call, a usage API and an OpenAPI specification. Common parameter names used by other screenshot APIs also work, which eases migration.
Rank #4
Pricing is straightforward: the Free plan includes 1,000 shots per month with no card; Starter is $5 for 3,000, Growth $15 for 15,000, Pro $39 for 60,000, Scale $99 for 250,000 and Business $249 for 1,000,000. Yearly billing gives two months free, and every feature is available on every plan. Create a free ScreenshotNeo account to start without a card.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Production checklist
- Pin your renderer and Java runtime. Record the library version, Java baseline and license terms.
- Make resources deterministic. Bundle fonts, images and CSS where possible; set an explicit base URI.
- Set timeouts and limits. Bound conversion time, input size, page count and output size, especially for remote URLs.
- Separate browser work. Run browser-backed conversion in an isolated process or service with restricted network access.
- Test representative pages. Include long tables, missing assets, SVG, non-Latin text, print CSS, page breaks and malformed input.
- Validate the PDF. Check that text is searchable, fonts are embedded as required, links work and accessibility or PDF/A requirements are met.
Troubleshooting common failures
Images or styles are missing
The usual cause is an unresolved relative URL. Set ConverterProperties.setBaseUri(...), use correct file permissions, and verify that the process can reach remote assets. Check the generated HTML for typos and redirects.
The PDF is blank or nearly empty
If the page relies on JavaScript to insert content, OpenHTMLToPDF and ordinary Flying Saucer rendering will not see it. Render the final HTML after application-side data binding, or use the Chrome-backed artifact or a browser service.
Modern CSS looks wrong
Flexbox, grid and browser-specific behavior exceed OpenHTMLToPDF’s documented scope. Simplify the print stylesheet, provide a supported fallback layout, or move to a Chromium-backed renderer.
Free tools Windows power users keep installed
One-click scans. No signup required.
Fonts or non-Latin characters are incorrect
Install or bundle the required fonts, register them according to the renderer’s documentation and confirm that the PDF embeds the expected glyphs. Font fallback can hide missing-font problems until production text differs from test data.
Conversion hangs on remote pages
Use request and overall timeouts, limit redirects and external resources, and log the URL and stage at which conversion stopped. For browser processes, also cap concurrent jobs and clean up crashed or orphaned processes.
Best Value
License review blocks deployment
iText pdfHTML has licensing terms that depend on the version and deployment model. Have legal or procurement review the applicable terms. OpenHTMLToPDF, Flying Saucer and PDFBox are open-source projects, but you still need to comply with their licenses and notices.
Performance and reliability decisions
In-process renderers avoid browser startup overhead and are often easier to scale for controlled documents. Their output is predictable only when the HTML, CSS, fonts and assets are controlled. Browser-backed rendering handles more real-world web behavior but consumes more memory and requires process supervision. The OpenHTMLToPDF project describes its newer renderer as faster for very large documents, but no controlled benchmark figure or universal speed guarantee is established; measure your own templates and page sizes.
For repeat jobs, cache immutable assets, reuse renderer infrastructure where supported, and avoid fetching the same remote stylesheet for every document. Do not treat cache hits or failed remote loads as successful conversions: record status, output size and validation results so callers can retry safely.
FAQ
Can I convert a URL directly with iText pdfHTML?
iText’s conversion API accepts HTML content and common input streams or files. Fetch the URL yourself, resolve its assets with a base URI, and remember that fetching HTML is not the same as executing its browser-side JavaScript.
Which approach creates accessible PDFs?
iText pdfHTML describes standards-compliant, accessible, searchable output, while OpenHTMLToPDF documents accessible PDF and PDF/A support. Confirm the exact conformance level with your selected version and validate the generated files.
Do I need PDFBox if I use OpenHTMLToPDF?
OpenHTMLToPDF uses PDFBox internally. Add PDFBox directly only when your application needs explicit PDF manipulation or post-processing beyond the renderer’s output.
What is the safest way to process user-supplied URLs?
Use a sandboxed worker with strict timeouts, output limits and egress controls. Treat every remote page as untrusted content and never give the renderer access to internal administrative endpoints.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

