Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

Give the PDF converter the same origin as the HTML, either with a base URI or with a custom resource resolver. In iText pdfHTML, set ConverterProperties.setBaseUri(...) and pass those properties to HtmlConverter. OpenHTMLtoPDF and Flying Saucer use the document or stylesheet URI as their base and expose resolver callbacks for authentication, filtering, URL rewriting, and custom schemes.

Why a base URI is required

An HTML string or input stream does not inherently tell a PDF renderer where css/site.css lives. A browser knows because it loaded the page from a URL; a renderer receiving only bytes does not. Without an origin, a relative stylesheet, font, image, or URL inside a stylesheet may be skipped or resolved incorrectly.

Use the URL of the HTML document as the base when the markup came from a web page. If the page uses an assets directory, use the directory that matches the relative links. For example, with href="css/site.css", a base of https://example.com/ produces https://example.com/css/site.css; a base of https://example.com/assets/ produces https://example.com/assets/css/site.css.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iText pdfHTML: set the base URI

In iText, ConverterProperties.setBaseUri supplies the parent location used to resolve linked resources. Pass the configured object to the overload of HtmlConverter.convertToPdf that accepts converter properties.

import com.itextpdf.html2pdf.ConverterProperties;
import com.itextpdf.html2pdf.HtmlConverter;

import java.io.InputStream;
import java.io.OutputStream;
import java.nio.file.Files;
import java.nio.file.Path;
import java.net.URI;
import java.net.http.HttpClient;
import java.net.http.HttpRequest;
import java.net.http.HttpResponse;

public class HtmlToPdf {
    public static void main(String[] args) throws Exception {
        URI page = URI.create("https://example.com/reports/invoice.html");

        HttpClient client = HttpClient.newHttpClient();
        HttpRequest request = HttpRequest.newBuilder(page).GET().build();
        HttpResponse<String> response = client.send(
                request, HttpResponse.BodyHandlers.ofString());
        if (response.statusCode() / 100 != 2) {
            throw new IllegalStateException("HTML request failed: " + response.statusCode());
        }

        ConverterProperties properties = new ConverterProperties()
                .setBaseUri(page.toString());

        try (OutputStream pdf = Files.newOutputStream(Path.of("invoice.pdf"))) {
            HtmlConverter.convertToPdf(response.body(), pdf, properties);
        }
    }
}

The HTML may contain an absolute link such as <link rel="stylesheet" href="https://example.com/assets/site.css"> or a relative link such as <link rel="stylesheet" href="../css/site.css">. The base URI is used for the latter and for other relative resources, including images and fonts.

When the HTML is already in a stream, the pattern is the same:

ConverterProperties properties = new ConverterProperties()
        .setBaseUri("https://example.com/assets/");

try (InputStream html = Files.newInputStream(Path.of("page.html"));
     OutputStream pdf = Files.newOutputStream(Path.of("page.pdf"))) {
    HtmlConverter.convertToPdf(html, pdf, properties);
}

Use a directory-style URI when links are relative to a directory. A base ending in / avoids accidentally treating the final path segment as a filename.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Fetch the page first without losing its origin

If you use Jsoup, retain the URL when parsing. Fetching with Jsoup.connect(...).get() retrieves an HTTP or HTTPS document and raises IOException for a failed request. If you parse an already fetched string, call the overload that accepts the document URL:

import org.jsoup.Jsoup;
import org.jsoup.nodes.Document;

Document document = Jsoup.connect("https://example.com/reports/invoice.html")
        .get();
String html = document.outerHtml();

// The second argument preserves the page origin for relative links.
Document reparsed = Jsoup.parse(
        html, "https://example.com/reports/invoice.html");

Pass the same origin, or an assets directory derived from it, to the renderer:

ConverterProperties properties = new ConverterProperties()
        .setBaseUri("https://example.com/reports/");
HtmlConverter.convertToPdf(html, outputStream, properties);

Do not replace the page URL with a local temporary-file path unless every resource has also been downloaded and rewritten to that path.

Authenticated, filtered, or rewritten resources in iText

A base URI solves address resolution, not access control. CSS may require an authorization header, cookies, a corporate proxy, or a trusted certificate. iText exposes a configurable resource retriever through its converter properties. Configure that retriever when the default HTTP retrieval cannot reach the stylesheet, or when you need to restrict which hosts can be contacted.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A safe retriever should:

  • Allow only the schemes and hostnames your document needs.
  • Apply authentication headers or cookies without logging their values.
  • Follow redirects according to your policy and validate TLS certificates.
  • Reject unexpected local-file or private-network targets to reduce server-side request risks.
  • Return the stylesheet bytes together with the stylesheet URL so relative fonts and images inside CSS resolve from the CSS file’s directory.

If your iText version provides a built-in HTTP client or resource retriever implementation, configure it through ConverterProperties rather than downloading CSS manually and changing every link. Verify the API names against the version used by your build, because resource-retriever classes vary between iText releases.

Rank #3
Sale
Play for Java: Covers Play 2
  • Used Book in Good Condition

OpenHTMLtoPDF: document URI and FSUriResolver

OpenHTMLtoPDF targets well-formed XML/XHTML and a CSS 2.1-oriented subset rather than full browser behavior. Relative URIs are resolved against the document URI or the stylesheet URI. Set the document’s base URL when creating the renderer, then install an FSUriResolver if resources need an allow-list, HTTPS-only enforcement, authentication, or URL rewriting.

// Representative configuration; adapt builder and resolver names to your OpenHTMLtoPDF version.
PdfRendererBuilder builder = new PdfRendererBuilder();
builder.withHtmlContent(xhtml, "https://example.com/reports/");
builder.toStream(outputStream);
// builder.useUriResolver(new MyAllowListedResolver(...));
builder.run();

The important value is the second argument to withHtmlContent: it is the origin used for links in the XHTML. If the CSS contains url("fonts/Inter.woff2"), the resolver must treat the stylesheet URL—not the original HTML URL—as that declaration’s base.

Flying Saucer: UserAgentCallback and setBaseURL

Flying Saucer provides the same extension point under different names. Its UserAgentCallback retrieves XML, CSS, and images and resolves URI and base-URI values. APIs include methods such as getCSSResource(String), resolveURI(String), and setBaseURL(String).

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use the callback when the CSS is behind authentication, when you must rewrite a legacy host, or when you need to deny requests outside an allow-list. Set the base URL before layout, and make sure the callback returns the original stylesheet URI so nested resources resolve correctly.

CSS that works in a browser but not in the PDF

Correct URL resolution does not guarantee visual parity. OpenHTMLtoPDF documents support for a reasonable subset of well-formed XML/XHTML and CSS 2.1; Flying Saucer and iText pdfHTML also have renderer-specific support. Browser-only JavaScript, client-side route changes, unsupported layout features, and dynamically injected styles may therefore be absent.

  • Prefer valid XHTML-style markup when using OpenHTMLtoPDF or Flying Saucer.
  • Inline critical styles temporarily to distinguish a retrieval problem from an unsupported CSS feature.
  • Replace JavaScript-generated content with server-rendered HTML before conversion.
  • Test flexbox, grid, filters, animations, custom properties, and complex pagination in the exact renderer and version you deploy.
  • Ensure the stylesheet response has CSS content and a successful status, rather than an HTML login page or an error document.

Common failures and precise fixes

Symptom Likely cause Fix
All relative CSS is ignored No base URI was supplied Set iText setBaseUri, or pass the document URL to the OpenHTMLtoPDF/Flying Saucer renderer.
Some links work, others point to the wrong folder Base URI is one directory too high or too low Match the base directory to the link. Check the resolved URL for css/site.css and for url(...) inside the CSS.
CSS request returns 401 or 403 Missing cookies, authorization, or user agent Use a custom retriever, FSUriResolver, or UserAgentCallback that supplies credentials securely.
HTTPS resource fails while the page works in Chrome Certificate, proxy, redirect, or firewall difference Inspect the renderer’s HTTP path, configure a trusted TLS/proxy setup, and permit redirects explicitly. Do not disable certificate validation in production.
Fonts or background images from CSS are missing Nested relative URLs were resolved against the HTML page or a local path Resolve them against the stylesheet’s own URL and return that URL from your resolver.
Layout differs despite successful CSS retrieval Unsupported browser CSS or JavaScript-dependent markup Check the renderer’s supported CSS subset and pre-render dynamic content.
Conversion hangs or is very slow Unbounded external requests, slow hosts, or redirect loops Set HTTP timeouts, cap redirects and resource size, cache approved assets, and use an allow-list.

Choosing a Java approach

Option URL and CSS control Best fit Trade-off
iText pdfHTML setBaseUri plus a resource retriever Commercial support and iText PDF features Commercial licensing; verify current terms.
OpenHTMLtoPDF Document URI plus FSUriResolver Open-source JVM projects CSS/HTML subset; browser parity is limited.
Flying Saucer UserAgentCallback, URI resolution, and base URL methods Existing XHTML/CSS pipelines Older API generations require maintenance validation.
Aspose.PDF for Java Web-page load options, CSS media and resource controls Commercial alternative with broad conversion controls Commercial licensing; verify current terms.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Reliability, performance, and security checklist

  • Fetch HTML and assets with explicit connect and read timeouts.
  • Set a maximum document size and maximum CSS, font, and image size.
  • Cache immutable stylesheets and fonts by URL and revision, but keep cache invalidation under your control.
  • Log resolved URLs, status codes, and timing without recording cookies or authorization headers.
  • Block file, loopback, link-local, and private-network URLs unless your use case explicitly requires them.
  • Use a deterministic media configuration and wait for all required resources before writing the PDF.
  • Compare a representative PDF after every renderer or stylesheet upgrade; no single renderer implements every browser feature.

Or skip the browser setup

If your goal is a dependable image or PDF capture of a live URL rather than Java-side HTML rendering, ScreenshotNeo provides a website screenshot API and MCP server. It accepts a URL in one request and can return PNG, JPEG, WebP, or PDF. Cookie and consent banners are accepted before capture, and more than 60 known consent platforms, newsletter popups, and chat widgets can be removed; each step can be disabled.

Only clean shots are billed. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and each response reports the result in X-Page-Verdict and X-Billed headers. Its MCP server exposes take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

See the ScreenshotNeo documentation for PDF parameters, CSS and JavaScript injection, selectors, device presets, authentication headers, cookies, waiting rules, resource blocking, signed links, asynchronous jobs, bulk capture, caching, and the usage API. The free plan includes 1,000 shots per month with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.

What to verify before shipping

  1. Resolve one relative stylesheet URL by hand and confirm it points to the intended host and directory.
  2. Fetch that URL with the same credentials, proxy, and TLS settings used by the converter.
  3. Check nested font and image URLs inside the stylesheet.
  4. Render a page containing print rules, a web font, a background image, and a page break.
  5. Test denied hosts, redirects, timeouts, and malformed CSS so failures are bounded and diagnosable.

Frequently Asked Questions

Should the base URI end with a slash?

Use a directory-style base ending in / when relative links are relative to that directory. A file-like final segment can change how a resolver combines paths.

Can I use an absolute stylesheet URL and still set a base URI?

Yes. An absolute href does not depend on the base, while relative links elsewhere in the document still use it.

Why does downloading CSS manually sometimes break fonts?

CSS often contains its own relative url(...) references. If you detach the stylesheet from its original URL, those nested paths lose their correct base.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Is a PDF renderer equivalent to a headless browser?

No. Libraries such as OpenHTMLtoPDF target a narrower, CSS-oriented subset and may not execute browser JavaScript or support every modern layout feature.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.