Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

RuntimeWorkerException: Invalid nested tag html found, expected closing tag body usually means XMLWorker reached a closing tag that does not match the tags still open in the input. First fix the XHTML: close tags in last-in, first-out order, use self-closing syntax for empty elements, and keep block elements out of paragraphs. Then parse with the correct character encoding. XMLWorker is not a browser and should not be expected to repair arbitrary modern HTML.

What the invalid nested tag error means

iText’s XMLWorker converts XHTML/CSS or XML flow to PDF. It processes markup while tracking which elements are open. An error such as Invalid nested tag html found, expected closing tag body indicates that the parser encountered </html> while its stack said </body> should close first. The mismatch may originate well before the tag named in the exception: a missing end tag, crossed tags, or an invalid wrapper can leave the parser in the wrong state.

For example, this markup crosses the closing order:

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
<div><p>Text</div></p>

Correct it so the most recently opened element closes first:

<div><p>Text</p></div>

The error is generally about the structure of the input rather than writing the PDF file. Repair the source before changing PDF writer settings or suppressing parser checks.

Repair the markup before parsing

Use the actual string or file passed to XMLWorker, not only the template or browser view. Templates, string concatenation, user content, and HTML cleanup steps can produce different final markup from what you expect.

  1. Capture the exact input. Log the HTML or XHTML immediately before parsing. Preserve the failing input securely if it contains private data, and reduce a copy to the smallest fragment that still throws.
  2. Inspect the tag named in the exception and the markup before it. If the parser expected body when it found html, look for an unclosed child of body, an extra closing tag, or wrappers that do not match.
  3. Close elements in nesting order. For every open element, close its children before its parent. Do not cross elements such as p and div.
  4. Make wrappers consistent. If your input includes document wrappers, include one root html element with corresponding head and body boundaries. Do not splice a second full document inside an existing body.
  5. Use XHTML syntax for empty elements. Write <br />, <hr />, and <img src="photo.png" />, rather than HTML-style unclosed empty tags. XMLWorker’s default tag factory includes processors for common elements such as br, hr, and img; the input still needs valid syntax.
  6. Keep block structure legal. Close a paragraph before starting a div, table, list, or heading. Close list items and table cells and rows in order: td or th, then tr, then any enclosing table section and table.
  7. Escape text and attributes. In text, encode a literal ampersand as &amp; and literal angle brackets as &lt; and &gt;. Check that attribute values use matching quotes and that entity references are valid.
  8. Validate separately. Run an XML/XHTML parser or validator as a preflight check. Fix its first structural error, then repeat; later parser errors can be consequences of the first one.

A browser may display malformed HTML by inferring omitted end tags or repairing nesting. That tolerance does not make the source well-formed XML, and it does not mean XMLWorker will apply the same repairs.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use XMLWorkerHelper with the right encoding

Once the input is valid, the standard iText 5 route is XMLWorkerHelper.getInstance().parseXHtml(...). The byte stream and the charset argument must agree. For UTF-8 input, convert the string to UTF-8 bytes and pass UTF-8 as the parser charset:

import com.itextpdf.text.Document;
import com.itextpdf.text.pdf.PdfWriter;
import com.itextpdf.tool.xml.XMLWorkerHelper;

import java.io.ByteArrayInputStream;
import java.io.ByteArrayOutputStream;
import java.nio.charset.StandardCharsets;

public class XmlWorkerExample {
    public static byte[] toPdf(String xhtml) throws Exception {
        ByteArrayOutputStream output = new ByteArrayOutputStream();
        Document document = new Document();
        PdfWriter writer = PdfWriter.getInstance(document, output);

        document.open();
        try {
            XMLWorkerHelper.getInstance().parseXHtml(
                writer,
                document,
                new ByteArrayInputStream(xhtml.getBytes(StandardCharsets.UTF_8)),
                StandardCharsets.UTF_8
            );
        } finally {
            document.close();
        }
        return output.toByteArray();
    }
}

This is a minimal shape for a controlled XHTML string, not a complete web-page rendering engine. Add any required CSS, font provider, or resource root through an appropriate parseXHtml overload when the document needs them. If input comes from a file or HTTP response, determine its actual encoding rather than assuming that every source is UTF-8.

Opening the document before parsing is necessary for this flow. Ensure it is closed even if conversion fails; the finally block above does that. In a larger application, preserve the original parsing exception in logs so a cleanup error does not obscure the markup failure.

Distinguish nesting errors from unknown tags

An unknown/custom tag and a known tag with invalid nesting are separate problems. A TagProcessorFactory maps tag names to processors; a missing mapping can cause an unsupported-tag problem. By contrast, setting the parser to accept unknown tags does not repair missing end tags or crossed nesting.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

For a custom element

If your source contains a custom element and the failure indicates that no processor is available, register a processor for that tag. iText’s custom-tag examples show creating or extending a processor and attaching the factory to HtmlPipelineContext with setTagFactory(factory). Choose a processor whose behavior matches the element’s content; mapping a custom element to an unrelated processor may parse without producing the intended layout.

When accepting unknown tags is appropriate

HtmlPipelineContext.setAcceptUnknown(true) is available when unknown tags can safely be accepted by the pipeline. It is not a general HTML recovery switch. Use it only for the unsupported-tag case, and keep validation and well-formed nesting checks in place.

When building the pipeline manually

For a custom pipeline, the main pieces are a CSSResolver, HtmlPipelineContext, HtmlPipeline, and PdfWriterPipeline, followed by XMLWorker and XMLParser. Configure the tag factory on the HTML context before parsing. Use this route when you need control over processors or pipeline behavior; it is more setup than the helper method and does not remove the requirement for valid input.

Choose between fixing the source and changing converters

Situation Practical path Trade-off
You control the markup and can produce stable XHTML Normalize and validate the source, then keep XMLWorker Retains the legacy pipeline, but requires disciplined input and supported layout features
The source contains custom elements Register suitable processors, or safely handle tags that can be ignored Requires explicit behavior for those elements; accepting unknown tags alone does not solve nesting
The source is browser-oriented HTML with optional end tags or modern CSS Normalize it to XHTML first; if layout demands remain unmet, evaluate migration to pdfHTML Migration can involve layout changes and compatibility work; it is not a guarantee of identical output

XMLWorker is a legacy iText 5 component designed around top-to-bottom, text-line-based conversion. iText’s comparison paper identifies pdfHTML as its successor and describes more robust handling of imperfect or invalid HTML, as well as broader HTML/CSS support. Treat that as a reason to evaluate a migration when requirements outgrow XMLWorker, not as a promise that every malformed document or legacy layout will convert unchanged.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Check the dependency and its licensing

Sonatype lists com.itextpdf.tool:xmlworker:5.5.13.6 as an XML-to-PDF artifact with CSS support and AGPL-3.0 licensing. Projects can load an older version transitively, so inspect the dependency actually resolved and deployed before applying version-specific fixes. Also verify that the license terms fit your distribution and use; do not assume that adding a newer XMLWorker dependency automatically replaces the version on the runtime classpath.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Troubleshoot common failure patterns

Symptom Likely cause What to do
Expected body, found html An earlier child element is still open, or document wrappers are mismatched Inspect the end of the body and work backward through open tags; make sure each child closes before body and html.
Error location appears far beyond the bad markup A missing or crossed tag earlier shifted the parser’s stack Reduce the captured input and validate from the start; fix the first structural error, not just the final reported tag.
Failure begins around br, hr, or img HTML-style empty-element syntax is being used in XHTML input Use a self-closing form such as <br /> or <img src="photo.png" />, with valid attributes.
Failure begins around a table, list, or heading inside a paragraph Invalid block nesting or missing paragraph/list/table closures End the paragraph first; close cells before rows, rows before tables, and list items within their list.
Only documents containing a custom element fail No matching tag processor is registered Register a processor through the tag factory, or omit the element only if losing its content is acceptable.
Text with symbols or non-English characters is corrupted or rejected Declared charset and actual bytes differ, or text/attributes contain unescaped XML characters Use a consistent encoding such as UTF-8 from source read through parser input, and escape reserved characters.
It works locally but fails after deployment A different iText/XMLWorker version or conflicting transitive dependency is loaded Inspect the resolved dependency tree and runtime classpath; align the deployed artifacts before comparing behavior.

Or skip the browser setup

If you need a visual reference for a publicly accessible page while diagnosing its generated markup, ScreenshotNeo can capture the page; a screenshot is for visual comparison and does not validate XHTML or fix XMLWorker input. Its API makes a screenshot request without setting up a browser in your application:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

See the ScreenshotNeo API documentation for request options. ScreenshotNeo accepts cookie/consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks/CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and responses identify the page verdict and billing status in headers. Its MCP server gives AI agents tools to take screenshots, get page information, and capture PDFs. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000. See ScreenshotNeo for details. Sign up for 1,000 free screenshots a month, with no card required.

FAQ

Is XMLWorker the same as iText 7’s pdfHTML?

No. XMLWorker belongs to the iText 5-era conversion path; pdfHTML is the successor product. Treat a move as a migration to evaluate against your input and layouts, not as a drop-in guarantee of identical PDFs.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Can I use XMLWorker for ordinary HTML copied from a browser?

It is safest when the input is controlled XHTML and uses features its pipeline can represent. Browser HTML may rely on optional end tags and browser recovery, so normalize and validate it before conversion.

Does setting setAcceptUnknown(true) stop nested-tag errors?

No. It concerns elements without registered processors; it does not fix tag-stack mismatches.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.