Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
Unicode support in HTML-to-PDF depends on several parts working together: the HTML must be decoded as UTF-8, the renderer must have fonts with the needed glyphs, those fonts must finish loading before printing, and the rendering engine must support the scripts’ shaping and text direction. UTF-8 fixes character decoding; it does not, by itself, guarantee that Chinese, Japanese, Arabic, Hebrew, or other scripts will render correctly in a PDF.
How to add Unicode support to HTML-to-PDF
Treat multilingual PDF output as a pipeline. Check each stage in order: input bytes, font availability, font readiness, renderer support, and the PDF that was actually produced. A failure at any stage can make the result look wrong even when the HTML preview appears correct.
- Encode the HTML as UTF-8. Save the source file as UTF-8. Put
<meta charset="UTF-8">near the beginning of the document’s<head>. For HTML served over HTTP, returnContent-Type: text/html; charset=UTF-8. Chrome’s guidance says the meta declaration should be entirely within the first 1,024 bytes of the document; a matching HTTP charset is also recognized. - Provide fonts with the required glyphs. Make suitable font files available in the conversion environment, configure fallbacks, and confirm the font chain covers the characters in your content. A CSS font-family name does not install a font or prove that it contains the needed glyphs.
- Wait for web fonts before printing. If the page loads fonts with CSS
@font-face, wait fordocument.fonts.readyafter the content is present and before asking the browser to generate the PDF. - Check shaping and direction support. Font coverage is separate from a renderer’s ability to lay out a script correctly. Test representative text, including mixed-direction passages, using the exact renderer and deployment environment you plan to use.
- Validate the PDF itself. Inspect its appearance and test line breaks, copy and paste, search, and embedded fonts in the target PDF reader. A correct HTML preview is not proof that the PDF is correct.
Why Unicode text can fail in a PDF
Wrong character decoding: mojibake or question marks
HTML is made of bytes that must be decoded into characters. If the file’s actual encoding and the encoding assumed by the renderer disagree, text can turn into mojibake or replacement characters before font selection even begins. Set the source encoding to UTF-8 and declare it early; for HTTP input, make the response charset agree with the bytes. A font change cannot repair text that was decoded incorrectly.
Quick wins for a faster PC:
Scan for outdated or missing drivers - takes under a minuteDriver Scan →Clear out junk files and repair common Windows errorsFree Scan →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Missing glyphs: boxes or blank characters
A correctly decoded Unicode character can still be absent from the selected font. The renderer then has no glyph to draw, which may appear as a box, a blank, or a missing-glyph symbol. The requested CSS family may not exist in the server image, or its glyph coverage may not include the target script. Configure fonts and a fallback chain in the actual conversion runtime, not just on a developer’s workstation.
#1 Best Overall
- HTML CSS Design and Build Web Sites
- Comes with secure packaging
- It can be a gift option
Incorrect shaping, ordering, or direction
Some scripts need shaping or bidirectional layout in addition to glyphs. Having every character in a font does not prove that the renderer will join, order, or position those characters correctly. This is especially important for right-to-left text and documents that mix scripts or directions.
Fonts that have not finished loading
A page can appear ready while a web font request is still in flight. Printing at that point can capture fallback text even though a later browser preview uses the intended font. Wait for the browser’s font readiness promise before PDF generation, and inspect failed font requests as well: readiness alone does not prove that every intended font loaded successfully.
Set the input encoding correctly
For a generated HTML document, use a UTF-8 file and declare the charset at the start of the head:
<!doctype html>
<html>
<head>
<meta charset="UTF-8">
<title>Multilingual document</title>
</head>
<body>
<p>English · 中文 · 日本語 · العربية · עברית</p>
</body>
</html>
The sample characters are only a basic decoding check; they do not establish that your fonts or PDF renderer support those scripts. If the HTML arrives over HTTP, make the response header’s charset match the document encoding. Chrome guidance places the meta declaration fully within the first 1,024 bytes, so avoid putting large content or other material ahead of it.
Rank #2
When a page contains multiple languages or directions, represent its actual language and direction in the markup and identify language changes where needed. Do not treat a language declaration as a substitute for fonts or script support: the available source material does not establish a complete markup recipe, and a lang value by itself does not add glyphs or renderer capability.
Choose and configure fonts for the target scripts
Choose fonts by the characters in the document, then make those font files discoverable in the process that creates the PDF. Test the complete fallback chain, not only the first family named in CSS. A browser on a developer laptop can silently use locally installed fonts that are absent from a container, server, or build image.
WeasyPrint font handling
WeasyPrint uses fonts discoverable through Pango and Fontconfig. Its reference says fonts are embedded and subset by default. It also notes that a missing Unicode code point becomes a .notdef glyph and triggers a warning in the logs. Check the conversion logs when characters disappear, and verify that the intended font files are visible to Pango/Fontconfig in the deployed environment.
Recommended Free Tools
Embedding and subsetting help put the fonts needed for the rendered document in the PDF; neither guarantees that a font covers every character or that the renderer supports every script’s layout behavior. Keep a representative multilingual PDF as a deployment check so font or image changes do not go unnoticed.
Rank #3
Wait for web fonts before browser PDF generation
In browser automation, wait for the page’s used fonts and associated layout work to settle before calling the PDF method. The browser Font Loading API exposes this readiness promise as document.fonts.ready. In Puppeteer, an illustrative sequence is:
await page.goto('file:///absolute/path/to/document.html', {
waitUntil: 'networkidle0'
});
await page.evaluate(() => document.fonts.ready);
await page.pdf({
path: 'document.pdf',
format: 'A4',
printBackground: true
});
Use a page URL and PDF settings appropriate to your application. The important ordering is to make the document available, await font readiness in the page, and only then generate the PDF. The snippet does not establish that a font request succeeded: inspect network failures and computed styles or otherwise verify that the intended font is in use. A readiness promise can settle even when a desired optional or unused font was never successfully loaded.
Puppeteer automates browsers and can generate PDFs, but the available evidence does not establish a complete per-script compatibility guarantee for any particular Chrome build. Verify your own scripts and layout on the exact browser build you deploy.
Free tools Windows power users keep installed
One-click scans. No signup required.
Test multilingual output with representative text
Build a small test document that reflects the content your application really handles. Include the target scripts, combining marks where relevant, punctuation, numerals, and mixed-direction passages. Check at least the following:
Rank #4
- Brand: Wiley
- Set of 2 Volumes
- A handy two-book set that uniquely combines related technologies Highly visual format and accessible language makes these books highly effective learning tools Perfect for beginning web designers and front-end developers
- Every character appears, with no boxes, blanks, or unexpected replacement symbols.
- Characters and marks are shaped and positioned as intended, including in words and mixed-script passages.
- Text direction, punctuation, and numerals appear in the expected order.
- Line breaks and fallback-font changes do not create collisions, clipped text, or inconsistent spacing.
- Text can be selected, copied, and searched in the resulting PDF.
- The PDF includes the expected fonts, and the result behaves in the PDF readers your users rely on.
Run this check in the actual deployment image and conversion path. A local browser preview exercises neither the server’s font installation nor every step in PDF generation.
Check renderer support before choosing an engine
Compare candidates against the scripts and directionality you need, required HTML and CSS, how fonts will be installed or loaded in your runtime, font embedding and text extraction, and deployment complexity. Reproduce the same sample document with the exact versions you intend to ship; broad claims about one renderer being best for every language are not supported by a comparable cross-renderer compatibility matrix here.
One important qualification applies to WeasyPrint: its current stable API reference lists right-to-left and bidirectional text as unsupported. Installing an Arabic- or Hebrew-capable font does not overcome that renderer limitation. If correct RTL or bidirectional layout is a requirement, test a renderer that supports the needed behavior rather than inferring support from font coverage alone. Renderer capabilities can change, so confirm against the version you deploy.
The Tool Desk
Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →PDF/A and Unicode text availability
WeasyPrint describes PDF/A-3u as a variant where the “u” indicates that text is available as Unicode. That property can matter when Unicode text availability is important for an archival output, but it is not a guarantee of correct visual glyphs, shaping, direction, or support for arbitrary HTML and CSS. Check both the standard’s text behavior and the visual output your document requires.
Best Value
Troubleshooting common multilingual PDF failures
| Symptom | Likely cause | What to check or change |
|---|---|---|
| Mojibake or question marks throughout the document | Bytes were decoded using the wrong character encoding, or the source text was already corrupted. | Confirm the HTML file is actually UTF-8, place the charset meta declaration early, and align any HTTP response charset with the bytes. Check the original text before changing fonts. |
| One script displays as boxes or blanks | The deployed font and fallback chain lack one or more glyphs, or the renderer reports a missing code point. | Make a font with the required coverage available to the conversion process, verify the fallback configuration, and inspect renderer logs. For WeasyPrint, check font discovery through Pango/Fontconfig and look for missing-glyph warnings. |
| The browser preview is right but the PDF uses a different-looking font | A web font was not ready when printing began, failed to load, or was unavailable in the PDF runtime. | Await document.fonts.ready, inspect failed font requests, and verify the font used in the actual conversion environment. |
| Arabic or Hebrew glyphs exist but text order or shaping is wrong | The rendering engine does not support the required RTL or bidirectional layout, or the document’s layout is not being handled as expected. | Test a representative mixed-direction sample. WeasyPrint’s current stable reference lists RTL/bidirectional text as unsupported; use an engine whose support you have verified for the deployed version. |
| Text looks right but cannot be searched or copied correctly | Visual rendering and Unicode text mapping are not the same check; the PDF may have a text extraction or font-encoding issue. | Test selection, copy/paste, and search in the resulting PDF, and inspect embedded fonts and Unicode text behavior rather than relying on appearance alone. |
| Only the server or container fails | The runtime lacks fonts available on a developer’s machine, or font configuration differs between environments. | Install or load the required font files in the deployed environment, verify renderer font discovery there, and run the same multilingual test PDF during deployment validation. |
Or skip the browser setup
If your task is to capture a web page as an image or PDF rather than build a custom multilingual HTML-to-PDF pipeline, ScreenshotNeo is a website screenshot API and MCP server. A single GET request can return PNG, JPEG, WebP, or PDF. It is not a substitute for validating a custom HTML renderer’s script support or font behavior.
For a one-call image capture, this cURL request saves a WebP screenshot of the target page. Create an API key and replace YOUR_API_KEY; the ScreenshotNeo documentation covers the API.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
ScreenshotNeo removes cookie and consent banners, newsletter popups, and chat widgets before capture; those cleanup steps can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and the response identifies the page verdict and billing status in headers. Its MCP server gives AI agents tools for screenshots, page information, and PDF capture. The free plan includes 1,000 shots per month with no card; paid plans start at $5 for 3,000 shots.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Sign up for ScreenshotNeo’s free 1,000 shots per month—no card required.
Frequently asked questions
Does UTF-8 alone guarantee Unicode support in a PDF?
No. UTF-8 addresses how input bytes become characters. Fonts must cover those characters, and the renderer must support the required shaping and direction.
Does embedding a font guarantee that every language renders correctly?
No. Embedding a font does not add missing glyphs or make a renderer support a script’s shaping or bidirectional layout.
Can I use WeasyPrint for Arabic or Hebrew PDFs?
Do not assume correct right-to-left or bidirectional output from font installation alone. WeasyPrint’s current stable API reference lists RTL/bidirectional text as unsupported; verify the deployed version and choose a renderer with tested support if that behavior is required.
Do these 3 things before closing this tab:
1Clear out junk files and repair common Windows errors2Scan for outdated or missing drivers - takes under a minute3Repair Windows errors before they cause bigger problemsQuick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

