Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

To preserve Unicode text in HTML-to-PDF conversion with legacy iTextSharp/iText 5 and XML Worker, do three things: pass the HTML’s actual character encoding to the parser, register the font files the HTML needs, and use CSS family names that match those fonts. A registered font must also contain the target script’s glyphs. For Arabic and other shaping-sensitive or right-to-left text, font registration alone does not guarantee correct shaping or reading order; test the exact XML Worker version and production output.

Confirm you are using the legacy XML Worker stack

This guide is for applications using iTextSharp/iText 5 with the matching XML Worker package. It is not a guide to the newer pdfHTML API: the two conversion paths should not be treated as interchangeable. Check the versions of the core library and XML Worker package in the application you deploy before copying code or upgrading either one.

The examples use XML Worker’s font-provider approach, which is directly relevant to HTML parsing. iText’s legacy FontFactory documentation also describes registering TrueType files and directories, but that is not a reason to assume that XML Worker will automatically find every installed system font.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Why Unicode text disappears or changes

The HTML may be decoded with the wrong character set

A UTF-8 font cannot repair text that was decoded into the wrong characters. If the HTML bytes are UTF-8, pass UTF-8 to XML Worker and ensure the bytes really are UTF-8. If the source is in another encoding, identify that encoding accurately and use it consistently rather than labeling it UTF-8.

The requested font may not be available to XML Worker

Declaring a family in CSS does not make its font file available. Register the font file with an XMLWorkerFontProvider and use the family name in the HTML. The iText examples for Cyrillic and Arabic demonstrate this explicit registration approach.

The font may not contain the required characters

Font availability and glyph coverage are separate checks. A registered font can still lack Cyrillic, Arabic, or other characters in the content. Choose files whose glyph coverage matches the languages you render, and test the actual text rather than assuming that a font described as Unicode-capable contains every glyph you need.

Complex-script shaping is a separate concern

Correct font selection does not establish that every XML Worker version will shape Arabic or lay out right-to-left text correctly. iText treats right-to-left HTML as a distinct concern. Validate glyph forms, ordering, and directionality using your deployed version; newer pdfHTML documentation can provide context, but it does not prove identical behavior in legacy XML Worker.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Register multiple fonts and parse UTF-8 HTML

The following example registers one font for a general text family and another for Arabic, keeps font lookup limited to the files you register, and passes UTF-8 explicitly. Put the font files in the application’s fonts directory and deploy that directory with the program. Use fonts you are licensed to redistribute and embed.

using System;
using System.IO;
using System.Text;
using iTextSharp.text;
using iTextSharp.text.pdf;
using iTextSharp.tool.xml;
using iTextSharp.tool.xml.pipeline.html;

public class HtmlPdf
{
    public static void Convert(string outputPath)
    {
        string html = @"<!doctype html>
<html>
<head>
  <meta charset='utf-8'>
  <style>
    body { font-family: FreeSans; }
    .arabic { font-family: 'Noto Naskh Arabic'; }
  </style>
</head>
<body>
  <p>Cyrillic: Привет, мир</p>
  <p class='arabic' dir='rtl'>مرحبا بالعالم</p>
</body>
</html>";

        byte[] htmlBytes = Encoding.UTF8.GetBytes(html);
        string baseDirectory = AppDomain.CurrentDomain.BaseDirectory;
        string fontDirectory = Path.Combine(baseDirectory, "fonts");
        string outputFile = Path.GetFullPath(outputPath);

        XMLWorkerFontProvider fontProvider =
            new XMLWorkerFontProvider(XMLWorkerFontProvider.DONTLOOKFORFONTS);
        fontProvider.Register(Path.Combine(fontDirectory, "FreeSans.ttf"));
        fontProvider.Register(Path.Combine(fontDirectory, "NotoNaskhArabic-Regular.ttf"));

        using (FileStream output = new FileStream(outputFile, FileMode.Create, FileAccess.Write))
        using (MemoryStream input = new MemoryStream(htmlBytes))
        using (Document document = new Document(PageSize.A4))
        {
            PdfWriter writer = PdfWriter.GetInstance(document, output);
            document.Open();
            XMLWorkerHelper.GetInstance().ParseXHtml(
                writer, document, input, null, Encoding.UTF8, fontProvider);
            document.Close();
        }
    }
}

The example uses the FreeSans and Noto Naskh Arabic font names shown in iText’s legacy Cyrillic and Arabic examples. Make sure the files you deploy are the intended fonts and that the family names recognized from those files correspond to the names in CSS. The Arabic dir='rtl' attribute is not a guarantee of correct shaping in every XML Worker version.

Adapt it when the HTML comes from a file

Read the file bytes without silently assuming a different encoding. If you know the file is UTF-8, pass those original bytes into a stream and keep Encoding.UTF8 in ParseXHtml. If the source uses another encoding, use that actual encoding for decoding and parsing. Avoid converting already-misdecoded text to UTF-8; that preserves the wrong characters.

Keep font lookup deliberate

The provider is initialized with DONTLOOKFORFONTS and the required files are registered explicitly. This follows the approach in iText’s XML Worker performance example: do not search broadly for fonts when you can state which fonts the HTML uses. Explicit paths and deployment of the same font files make selection easier to reason about across machines.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Choose and deploy fonts for the scripts you render

Evaluate each candidate font against the actual output requirements rather than choosing only by its family name.

  • Glyph coverage: confirm it includes the characters in your source text, including punctuation and language-specific marks.
  • Deployment: ensure the file is present at the registered path in every environment where conversion runs. A font installed on a developer’s computer may not exist on a server.
  • Embedding rights: review the font’s license and embedding permissions before shipping it in an application or generated PDFs.
  • Visual fidelity: check that the font’s appearance fits the intended design and does not change line wrapping in a way that matters.
  • Script layout: verify shaping and bidirectional order in the exact legacy XML Worker version and PDF viewers relevant to your application.

iText’s Arabic example uses Noto Naskh Arabic, while its Cyrillic example illustrates a Unicode-capable font. These are examples, not a claim that either font is the right choice for every document or that registration alone solves all script-layout behavior.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Test the PDF, not just the conversion call

A conversion that completes without throwing an exception can still produce missing glyphs or incorrect order. Test representative documents in the deployed runtime and inspect both appearance and extracted text.

  1. Include samples with the scripts, punctuation, mixed-language runs, and directionality used in production.
  2. Generate the PDF using the same XML Worker package, font files, and configuration as the deployed application.
  3. Inspect the rendered pages for missing boxes, substitutions, unexpected wrapping, joining, and reading order.
  4. Extract text from the PDF and compare it with the expected characters. Visible output and extracted text are separate checks.
  5. Open the file in the PDF viewers your users rely on, especially for right-to-left or shaping-sensitive content.

If the output fails, isolate the problem: first check the source bytes and charset, then font registration and glyph coverage, then script shaping and directionality.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Troubleshoot common failures

Symptom Likely cause What to check
Question marks or corrupted characters throughout the document The HTML was decoded using an encoding different from the one used for its bytes. Inspect the source bytes and use the matching charset when parsing. Do not rely only on a meta charset declaration if the bytes contradict it.
Boxes or missing glyphs for one language The selected font is not registered, or it lacks the needed characters. Confirm the font path exists on the conversion machine, registration succeeds, the CSS family matches, and the file covers the text.
It works locally but not on the server The server lacks the font file or uses a different path or deployment layout. Deploy licensed font files with the application and register explicit paths based on the application’s runtime location.
Arabic letters appear but joining or order is wrong Font registration may be successful while shaping or bidirectional layout is not handled as required by the legacy conversion path. Test the exact XML Worker version and right-to-left markup with representative text. Do not treat a font change as proof that layout is fixed.
CSS names a font but output uses another appearance The requested family may not be registered under the name used in the HTML, or broad lookup may select an unintended font. Register the specific file, verify its recognized family name, and use that name in CSS.
Changing to pdfHTML examples causes compile or runtime errors The code belongs to a newer conversion API rather than the iTextSharp/XML Worker stack. Keep code aligned with the installed library generation, or plan and validate a separate migration instead of mixing API examples.

Or skip the browser setup

If what you need is a capture of a webpage available at a URL, rather than conversion of an arbitrary HTML string inside your C# application, ScreenshotNeo is an alternative. Its API returns screenshots or PDFs, but it does not replace XML Worker’s font-registration workflow for local HTML conversion.

For a screenshot, one GET request is enough:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

See the ScreenshotNeo API documentation for request options and PDF output. Before capture, it accepts consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be turned off. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and response headers identify the page verdict and billing status. Its MCP server provides screenshot and PDF tools for AI agents. The Free plan includes 1,000 shots a month with no card; paid plans start at $5 for 3,000 shots. Sign up for ScreenshotNeo’s free plan.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.