Use iText pdfHTML when you need broad HTML/CSS support, tagged or PDF/A output, forms, and continued iText document processing. Use OpenHTMLtoPDF when an LGPL, PDFBox-based renderer is sufficient for controlled XHTML/CSS templates without JavaScript, flexbox, or grid. This guide shows runnable Java code for HTML strings and files, explains relative assets and fonts, and covers API choices, accessibility, troubleshooting, performance, and licensing.
Choose the renderer before writing code
HTML-to-PDF libraries are renderers, not full browsers. Your choice should follow the document’s layout and compliance requirements.
| Requirement | iText pdfHTML | OpenHTMLtoPDF |
|---|---|---|
| Best fit | Applications needing maintained HTML/CSS conversion, accessibility, tagging, PDF/A, forms, SVG, RTL content, or further iText manipulation. | Controlled, well-formed XHTML/CSS templates where an LGPL and PDFBox stack fit the project. |
| Browser features | Validate the exact feature set against the selected version; it is still a document renderer rather than Chrome. | Does not execute JavaScript and does not implement many modern standards, including flex and grid. |
| Layout guidance | Use the documented pdfHTML/CSS support and test your actual templates. | Prefer table layouts and avoid floats close to page breaks, as the project README advises. |
| Output and post-processing | Direct PDF conversion, tagged PDFs, PDF/A examples, and APIs that return a Document or parsed elements for additional content. | PDFBox-based PDF output with documented accessible and PDF/A capabilities. |
| Runtime and license | Check the iText version’s Java requirements and commercial/open-source terms for your deployment. | LGPL; the README states Java 8 is the minimum and records testing with OpenJDK 8 and 11 (plus 17 early access). Verify current releases before pinning; the changelog lists 1.0.10 (2021-09-13) and a later 1.0.11-SNAPSHOT heading. |
Do not start new code with iText’s removed HTMLWorker or the older XML Worker approach. The iText 7 renderer framework and pdfHTML add-on are the current API family described for HTML conversion.
Set up iText pdfHTML
Add the pdfHTML add-on and its compatible iText Core dependencies using the coordinates and version recommended in the official iText documentation for your project. Keep all iText modules on the same compatible release line; do not mix arbitrary versions. iText licensing and commercial-support requirements depend on how you distribute and use the library, so review the applicable terms before shipping.
Free tools Windows power users keep installed
One-click scans. No signup required.
Convert an HTML string or file
Smallest string example
import com.itextpdf.html2pdf.HtmlConverter;
import java.io.FileOutputStream;
import java.io.IOException;
public class HtmlStringToPdf {
public static void main(String[] args) throws IOException {
String html = "<html><body><h1>Hello, PDF</h1>"
+ "<p>Generated from a Java string.</p></body></html>";
try (FileOutputStream output = new FileOutputStream("out.pdf")) {
HtmlConverter.convertToPdf(html, output);
}
}
}
convertToPdf accepts an HTML string and writes to an output stream, file, PdfWriter, or PdfDocument. The destination stream is where the PDF bytes go; close it after conversion.
Convert an HTML file
import com.itextpdf.html2pdf.HtmlConverter;
import java.io.FileInputStream;
import java.io.FileOutputStream;
import java.io.IOException;
public class HtmlFileToPdf {
public static void main(String[] args) throws IOException {
try (FileInputStream html = new FileInputStream("invoice.html");
FileOutputStream pdf = new FileOutputStream("invoice.pdf")) {
HtmlConverter.convertToPdf(html, pdf);
}
}
}
For a File source, iText can use the file’s parent directory as the default base URI. Streams do not carry that location, so provide one explicitly whenever the HTML contains relative URLs.
Resolve relative CSS, images, and fonts with a base URI
An HTML reference such as img/logo.png is resolved relative to a base URI. A stream has no parent directory for the converter to infer, so configure ConverterProperties.
import com.itextpdf.html2pdf.ConverterProperties;
import com.itextpdf.html2pdf.HtmlConverter;
import java.io.FileInputStream;
import java.io.FileOutputStream;
import java.io.IOException;
public class AssetsToPdf {
public static void main(String[] args) throws IOException {
ConverterProperties properties = new ConverterProperties();
properties.setBaseUri("/opt/app/templates/invoice/");
try (FileInputStream html = new FileInputStream("/opt/app/templates/invoice/invoice.html");
FileOutputStream pdf = new FileOutputStream("invoice.pdf")) {
HtmlConverter.convertToPdf(html, pdf, properties);
}
}
}
Place the CSS, images, SVG files, and font resources beneath the configured directory, or use absolute URLs that your deployment can reach. In production, make resource access deterministic: package templates and assets together, restrict outbound fetching, and test behavior when an asset is missing or slow.
Recommended Free Tools
Rank #2
Choose the API for your document workflow
Write a finished PDF
Use convertToPdf(...) when the HTML is the complete document and you only need a PDF output stream.
Append content after HTML parsing
import com.itextpdf.html2pdf.HtmlConverter;
import com.itextpdf.layout.Document;
import com.itextpdf.kernel.pdf.PdfDocument;
import com.itextpdf.kernel.pdf.PdfWriter;
import java.io.IOException;
public class AppendAfterHtml {
public static void create(String html, String destination) throws IOException {
PdfDocument pdf = new PdfDocument(new PdfWriter(destination));
Document document = HtmlConverter.convertToDocument(html, pdf);
document.add(new com.itextpdf.layout.element.Paragraph("Appended by Java"));
document.close();
}
}
convertToDocument returns an iText Document so application code can add elements after conversion. Close the document to finish the file.
Insert parsed elements into a separately managed flow
convertToElements(...) returns parsed elements. This is useful when your application owns page setup, headers, footers, or a larger document lifecycle and wants to insert HTML-derived content into it.
Tagged, accessible, and archival PDFs
For a tagged PDF, enable tagging on the PDF document before conversion:
import com.itextpdf.html2pdf.HtmlConverter;
import com.itextpdf.kernel.pdf.PdfDocument;
import com.itextpdf.kernel.pdf.PdfWriter;
import java.io.IOException;
public class TaggedPdf {
public static void create(String html, String destination) throws IOException {
PdfDocument pdf = new PdfDocument(new PdfWriter(destination));
pdf.setTagged();
HtmlConverter.convertToPdf(html, pdf);
pdf.close();
}
}
Use semantic HTML—headings in order, lists for lists, table headers, meaningful link text, and alternative text for informative images. The iText examples also cover PDF/A-3B, custom fonts, HTML forms, Arabic and Hebrew, and SVG. Treat each as a documented capability to validate against the exact pdfHTML version, content, and conformance profile you deploy; conversion success alone does not prove accessibility or archival compliance.
OpenHTMLtoPDF implementation
OpenHTMLtoPDF is a pure-Java renderer for a reasonable subset of well-formed XML/XHTML and some HTML5. It uses CSS 2.1 and later standards and PDFBox. It is a practical alternative for controlled templates, but it will not execute JavaScript and does not implement many modern layout systems such as flex and grid.
Craft XHTML for the engine: close every element, use explicit character encoding, prefer tables for complex page layout, and avoid floats near page boundaries. Because dependency coordinates and releases can change, use the project’s current README and release metadata when adding it to Maven or Gradle, and verify Java compatibility before deployment.
CSS, images, fonts, and web-page limitations
- CSS: Test print-specific rules, page breaks, counters, and unsupported properties with representative documents. Browser CSS that relies on flex, grid, JavaScript-generated markup, or animation may not render as expected.
- Images: Confirm every relative path against the configured base URI. Check file permissions, URL reachability, MIME types, and image dimensions.
- Fonts: Register or package the fonts required for the target languages and verify that the resulting PDF embeds the intended glyphs. Missing fonts commonly appear as blank squares or fallback typography.
- Dynamic pages: Render the HTML first in a browser or server-side template engine if JavaScript is required, then pass the resulting static HTML to the PDF renderer.
- Security: Treat HTML and resource URLs as untrusted input. Apply size, time, and network policies appropriate to your service, and avoid allowing arbitrary local-file or internal-network reads.
Or skip the browser setup
If your real starting point is a public web page rather than a controlled HTML template, ScreenshotNeo can return a screenshot or PDF through one GET request. It accepts cookie and consent banners like a visitor, then removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each cleanup step can be disabled. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers report the page verdict and billing status.
Windows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallOutdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchcurl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
See the ScreenshotNeo API documentation for PDF parameters and the 63 capture options, including full-page lazy-image loading, CSS-selector elements, device and viewport settings, retina scale, custom CSS/JavaScript, clicks, waits, blocked resources, headers, cookies, user agents, authorization, timezone, geolocation, transparent backgrounds, resizing, caching, signed links, asynchronous webhooks, bulk capture, usage, and OpenAPI support. Its MCP server supplies take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients. The Free plan includes 1,000 screenshots per month without a card; paid plans start at $5 for 3,000, and every feature is available on every plan. Create a free ScreenshotNeo account.
Rank #4
Production performance and reliability checklist
- Reuse compatible renderer configuration and font resources where the library permits, rather than rebuilding them for every request.
- Set request-level timeouts and output-size limits around your conversion service.
- Measure conversion time and memory with your largest real templates; page count, high-resolution images, SVG complexity, and font embedding can dominate resource use.
- Write to a temporary file or controlled stream, then atomically publish the completed PDF so readers never receive partial output.
- Keep renderer versions pinned, test upgrades against visual fixtures, and inspect PDFs for missing assets, page-break changes, metadata, tagging, and font embedding.
- For parallel jobs, establish a safe concurrency limit from load tests instead of assuming unlimited thread safety or memory.
Troubleshooting common failures
Images or CSS are missing
Cause: A relative URL has no correct base URI, or the process cannot read the resource. Fix: call properties.setBaseUri(...), use paths relative to that directory, verify permissions, and log resource-resolution failures.
The output is blank or conversion throws a parsing error
Cause: Malformed markup, an unsupported construct, or an empty stream. Fix: save and inspect the exact HTML bytes, provide a declared UTF-8 encoding, close tags, and reduce the document to a minimal failing case.
Modern browser styling is rearranged
Cause: The renderer is not a browser; OpenHTMLtoPDF specifically lacks JavaScript, flex, and grid support. Fix: replace those layouts with supported print CSS or pre-render the page in a browser before conversion.
Non-Latin text shows boxes
Cause: The required font or glyph shaping is unavailable. Fix: package and configure a font covering the script, then verify embedding and shaping in the produced PDF.
Best Value
Pages break in the wrong places
Cause: floats, oversized blocks, or unsupported CSS page-break behavior. Fix: use print-oriented structure, table layouts where appropriate, explicit break rules, and content-sized images; compare output across representative page lengths.
Practical decision checklist
- Choose iText pdfHTML for the maintained iText ecosystem, accessibility/tagging, PDF/A, forms, or post-conversion document APIs.
- Choose OpenHTMLtoPDF when LGPL licensing and a PDFBox renderer meet the needs of stable XHTML/CSS templates.
- Inventory JavaScript, flex/grid, fonts, SVG, RTL text, forms, external assets, and compliance targets before selecting a library.
- Build a fixture set containing short and long pages, tables, images, missing assets, multiple languages, and intentional page breaks.
- Validate the generated PDFs for visual fidelity, text search, font embedding, accessibility, and archival requirements—not just whether a file was created.
Frequently Asked Questions
Can these libraries convert any URL directly?
Not reliably. They consume HTML and resources through their renderer rules; JavaScript-heavy pages should be rendered first in a browser or replaced with static HTML.
Which API should I use to add a footer after conversion?
Use iText’s convertToDocument(...) workflow, then add iText elements before closing the returned Document.
Quick wins for a faster PC:
Scan for outdated or missing drivers - takes under a minuteDriver Scan →Repair Windows errors before they cause bigger problemsFix Now →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Is a successful conversion proof that a PDF is accessible?
No. Use semantic source HTML, enable tagging where required, and test the resulting PDF against your accessibility and conformance criteria.
The Bottom Line
For a Java application that needs dependable HTML/CSS conversion and advanced PDF capabilities, start with iText pdfHTML and configure a base URI for every stream-based document. Select OpenHTMLtoPDF for controlled XHTML/CSS templates when its rendering limits and LGPL license fit. Whichever engine you choose, test real assets, fonts, page breaks, and compliance requirements before production.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

