Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
For a static, well-formed XHTML page, Java can create a PDF locally with OpenHTMLtoPDF or Flying Saucer. Those libraries parse XML/XHTML and CSS and write a PDF without a browser. If the page needs JavaScript, flexbox, grid, login state, or other browser behavior, use a browser-backed renderer or a hosted HTML-to-PDF service instead. The right choice depends on the page you are converting, not just on the Java API.
Choose the renderer before writing code
A URL-to-PDF pipeline has two separate jobs: retrieve the URL and render its HTML, CSS, images, fonts, and scripts into PDF instructions. A Java PDF library that only creates PDF objects is not automatically an HTML renderer.
| Page or requirement | Best starting point | Reason |
|---|---|---|
| Controlled XHTML/XML and CSS 2.1-style layout | OpenHTMLtoPDF | Pure-Java URI and HTML-content APIs with PDF output. |
| Existing XHTML/XML application using Flying Saucer | Flying Saucer | Direct URL/file rendering methods and an XML/XHTML-focused model. |
| JavaScript-generated content, modern CSS, or authenticated browser state | Browser-backed renderer or hosted service | A real browser can execute scripts and apply browser layout rules. |
| Post-processing an already-created PDF | Apache PDFBox | PDFBox creates and manipulates PDFs; it is not an HTML/CSS URL renderer. |
Before conversion, confirm that the URL is reachable from the machine running Java, returns the expected HTML rather than a bot-check page, and does not require browser-only interaction. For private pages, plan how cookies, authorization headers, or a signed URL will be supplied. Also decide whether remote images, stylesheets, fonts, and redirects are allowed by your network policy.
Free tools Windows power users keep installed
One-click scans. No signup required.
Convert a URL with OpenHTMLtoPDF
OpenHTMLtoPDF is the most direct local implementation when the source is well-formed XHTML/XML. Its URI builder accepts a URL and writes the result to an output stream. Its HTML-content builder accepts a string plus a base document URI, which is important when the markup contains relative links.
Add the Maven artifact com.openhtmltopdf:openhtmltopdf-pdfbox to your build using the current version published by the project. The artifact uses PDFBox for PDF output. Keep the version in your dependency management so it can be updated independently of application code.
Java example: URL directly to a PDF file
import com.openhtmltopdf.pdfboxout.PdfRendererBuilder;
import java.io.BufferedOutputStream;
import java.io.IOException;
import java.io.OutputStream;
import java.net.URI;
import java.nio.file.Files;
import java.nio.file.Path;
public final class UrlToPdf {
private UrlToPdf() {}
public static void main(String[] args) throws Exception {
if (args.length != 2) {
throw new IllegalArgumentException("Usage: java UrlToPdf <url> <output.pdf>");
}
String source = args[0];
Path destination = Path.of(args[1]);
URI uri = URI.create(source);
String scheme = uri.getScheme();
if (!("https".equalsIgnoreCase(scheme) || "http".equalsIgnoreCase(scheme))) {
throw new IllegalArgumentException("Only http and https URLs are allowed");
}
Path parent = destination.toAbsolutePath().getParent();
if (parent != null) {
Files.createDirectories(parent);
}
try (OutputStream output = new BufferedOutputStream(Files.newOutputStream(destination))) {
PdfRendererBuilder builder = new PdfRendererBuilder();
builder.withUri(uri.toString());
builder.toStream(output);
builder.run();
}
System.out.println("Wrote " + destination.toAbsolutePath());
}
}
Run it with a reachable URL and a writable destination, for example java UrlToPdf https://example.com out/example.pdf. The renderer resolves relative resources against the document URI. If your application already fetched HTML, use withHtmlContent(html, baseDocumentUri) instead; pass the page URL as the base URI so relative CSS, images, and fonts still resolve.
HTML-content variant
String html = "<html><head><link rel="stylesheet" href="/css/print.css"></head>"
+ "<body><h1>Invoice</h1></body></html>";
try (OutputStream output = Files.newOutputStream(Path.of("invoice.pdf"))) {
new PdfRendererBuilder()
.withHtmlContent(html, "https://www.example.com/invoices/123")
.toStream(output)
.run();
}
The supplied markup must be sufficiently well formed for an XML/XHTML parser. Close elements, quote attributes, use valid nesting, and provide print-oriented CSS. A page that looks fine in a browser can still fail or render differently if its markup depends on browser error recovery.
Fonts, images, and resources
- Use absolute HTTPS URLs or a correct base URI for relative assets.
- Make sure the Java process can resolve DNS and establish outbound connections.
- Embed or make required fonts available to the renderer; a missing font can change line wrapping and page count.
- Use print CSS and explicit page-break rules for invoices, reports, and long documents.
- Do not assume browser extensions, session cookies, or JavaScript-created asset URLs exist in a local renderer.
Use Flying Saucer for XHTML/XML documents
Flying Saucer describes itself as a pure-Java XML/XHTML and CSS 2.1 renderer. Its PDF output module exposes direct URL and file methods, including PDFRenderer.renderToPDF(String url, String pdf) and file overloads.
Rank #2
Minimal Java example
import org.xhtmlrenderer.pdf.PDFRenderer;
public final class FlyingSaucerUrlToPdf {
private FlyingSaucerUrlToPdf() {}
public static void main(String[] args) throws Exception {
if (args.length != 2) {
throw new IllegalArgumentException("Usage: java FlyingSaucerUrlToPdf <url> <output.pdf>");
}
PDFRenderer.renderToPDF(args[0], args[1]);
}
}
Add the Flying Saucer PDF module appropriate for the release you select. Runtime requirements vary by release: the project documents Java 11 or later for 9.5.0, Java 17 or later for 9.6.0, and Java 21 or later for 10.0.0. Check the release documentation before choosing a Java runtime.
Flying Saucer is a good fit when you control the XHTML and the layout fits CSS 2.1. It is not a drop-in replacement for a current Chromium browser. Test tables, floats, page breaks, SVG, font embedding, and remote resources with representative documents.
Why JavaScript-heavy pages fail with local HTML renderers
OpenHTMLtoPDF’s own FAQ says it is not a web browser: it does not execute JavaScript and does not implement many modern standards such as flex and grid. Consequently, a single-page application may produce an empty shell, an old loading state, or a layout that differs substantially from Chrome or Firefox. Flying Saucer has the same fundamental limitation because it targets XML/XHTML and CSS 2.1 rather than browser scripting.
Do these 3 things before closing this tab:
1Clear out junk files and repair common Windows errors2Fix the driver behind crashes, sound loss and screen glitches3Repair Windows errors before they cause bigger problemsUse a browser-backed renderer when the final content appears only after client-side execution, when CSS relies on flexbox or grid, or when the page requires a user gesture, login session, or browser storage. A hosted option such as Adobe PDF Services documents HTML-to-PDF input from static HTML, dynamic HTML, ZIP files, and URLs, with Java integration guidance.
If you choose a browser service, define a readiness condition instead of capturing immediately. Wait for a selector containing the finished content, wait for network idle where appropriate, and set a hard timeout. Capture after authentication and consent handling, not before. For sensitive pages, review where the HTML and credentials are processed.
Keep PDFBox in the right role
Apache PDFBox is an open-source Java library for creating PDF documents, manipulating existing PDFs, and extracting content. It is useful after HTML rendering for merging files, adding metadata, stamping pages, encrypting output, or extracting text. It does not replace an HTML/CSS renderer for converting a URL. A typical pipeline is therefore: render HTML with OpenHTMLtoPDF, Flying Saucer, or a browser-backed service, then use PDFBox for PDF-specific post-processing.
Build a reliable URL-to-PDF service
Validate input and control network access
- Accept only
httpandhttpsunless local files are an explicit, protected feature. - Reject unsupported schemes and malformed URLs before invoking the renderer.
- Prevent server-side request forgery by restricting private IP ranges, metadata endpoints, and internal hostnames when users supply URLs.
- Set connection and read timeouts, cap response sizes, and limit redirect depth.
- Decide whether redirects may cross hosts and whether remote resources are permitted.
Make output deterministic
- Pin the Java runtime and renderer versions in deployment.
- Use the same fonts and font configuration in development and production.
- Set a consistent page size, margins, and print stylesheet.
- Record the source URL, renderer, runtime, elapsed time, output bytes, and failure category.
Handle failures without corrupt files
Render to a temporary file or memory stream, check that the operation completed, and atomically move the result into its final location. Do not publish a partially written PDF after a network or resource error. Return distinct errors for invalid input, unreachable URLs, malformed markup, missing resources, timeout, and renderer failure so callers can retry intelligently.
Troubleshooting common conversion errors
| Symptom | Likely cause | Fix |
|---|---|---|
| Blank PDF or missing body | Content is inserted by JavaScript. | Use a browser-backed renderer or fetch a server-rendered/export endpoint. |
| Flexbox or grid layout collapses | Local renderer lacks those modern layout features. | Rewrite the print stylesheet for supported CSS or use a browser engine. |
| Images or CSS are absent | Relative URLs have no correct base, or the process cannot reach the host. | Pass the page URI/base URI, use absolute resource URLs, and verify network access. |
| Fonts change line breaks | Font files are unavailable or not embedded. | Install/register the required fonts and test the deployed environment. |
| XML parsing exception | Malformed HTML, unclosed tags, or invalid nesting. | Serve XHTML-compatible markup or sanitize and normalize the HTML before rendering. |
| 401/403 response | The URL requires authentication or rejects non-browser requests. | Use an authenticated fetch, a signed export URL, or a browser session; do not silently publish an error page as a PDF. |
| Timeout or out-of-memory error | Large images, endless requests, or an unbounded page. | Set limits, optimize assets, impose page and resource budgets, and capture long documents in controlled chunks. |
| PDF file is zero bytes | The output stream was closed early or rendering failed before writing. | Use try-with-resources, write to a temporary destination, and inspect the original exception. |
Performance, reliability, and cost considerations
Local libraries avoid per-request service fees and keep source HTML inside your infrastructure, but they consume your CPU and memory and require you to operate fonts, networking, queues, and upgrades. Browser-backed services generally provide better fidelity for modern pages but add network latency, a provider dependency, and possible data-processing concerns. A hosted API can be simpler for bursty workloads; measure queue time, conversion time, output size, and retry behavior with your own pages because no authoritative cross-library performance figure is established here.
Rank #4
Cache only when the URL, authentication state, and page content are stable. A cache key should include relevant query parameters and rendering options. For dynamic pages, stale PDFs can be more damaging than a slower fresh conversion. Queue large jobs, enforce concurrency limits, and use exponential backoff only for transient network failures; malformed markup and authorization errors should not be retried indefinitely.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Or skip the browser setup
ScreenshotNeo is a hosted website screenshot API that can return PNG, JPEG, WebP, or PDF from one GET request. It accepts the cookie or consent banner before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each cleanup step can be disabled. Only clean shots are billed: bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers identify the page verdict and billing status.
For a URL-to-PDF call, use the API endpoint documented at https://screenshotneo.com/docs/:
Quick wins for a faster PC:
Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Clear out junk files and repair common Windows errorsFree Scan →curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
The same endpoint can be called from Python:
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
Or Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' }); const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
ScreenshotNeo also provides an MCP server with take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients. It supports full-page capture, selector targeting, device and viewport settings, dark mode, retina scale, PDF paper and margin options, custom CSS and JavaScript, waits, request blocking, headers, cookies, user agents, authorization, timezone, geolocation, resizing, selectable cache TTLs, signed links, asynchronous webhooks, bulk capture for up to 100 URLs per call, and a usage API. Every feature is included on every plan. The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.
Frequently Asked Questions
Can OpenHTMLtoPDF convert any public web page?
No. The page must be reachable and its markup and CSS must fit the renderer’s supported XHTML/XML and CSS subset. JavaScript-only content and many modern browser layouts require another approach.
Best Value
Should I use PDFBox instead of an HTML-to-PDF library?
Use PDFBox for creating or post-processing PDF files. Pair it with an HTML renderer when the input is a web page.
How do I preserve relative image and stylesheet URLs?
Use the URL form of the renderer or pass the original page URL as the base document URI when supplying HTML content.
Recommended Free Tools
What is the safest way to convert user-supplied URLs?
Validate the scheme, restrict internal destinations to prevent SSRF, enforce time and size limits, and isolate rendering from sensitive network resources.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

