Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

Most wkhtmltopdf encoding failures have one of four causes: the HTML bytes are not actually UTF-8, the document does not declare UTF-8 early enough, an HTTP response supplies a conflicting charset, or the rendering host lacks a font for the characters. Fix those in that order, then use --encoding utf-8 (or web.defaultEncoding = "utf-8" in a binding) as a fallback—not as a replacement for correctly encoded input.

Identify the failure before changing options

Encoding and font coverage are different problems. Mojibake such as é usually means bytes were decoded with the wrong character set. Empty squares, question marks, or missing Chinese, Japanese, Korean, emoji, or symbol characters can mean the selected font has no glyph, even when the text is correctly decoded.

Symptom Most likely cause First check
Every accented character is garbled Wrong source bytes, response charset, or document declaration Inspect the file bytes and HTTP Content-Type
A URL works but the saved HTML file fails The downloaded copy lost HTTP charset metadata or was re-saved in another encoding Compare the response headers and local file encoding
Only CJK text or symbols are boxes/missing Font lacks the required glyphs Install and select a font with coverage on the rendering host
Page body is correct but header/footer text is wrong Header/footer is a separate input with its own decoding Use UTF-8 header/footer HTML

Make the HTML bytes really UTF-8

A meta tag tells a parser how to interpret bytes; it cannot convert bytes that were saved in Windows-1252, ISO-8859-1, or another encoding. Configure your editor, template engine, export job, or database connection to emit UTF-8, then verify the result before invoking wkhtmltopdf.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Check a local file

  • Open the file in an editor that displays the actual encoding, not just the characters.
  • On Unix-like systems, run file input.html or an equivalent byte/encoding inspector. Treat the result as a hint and verify any ambiguous file manually.
  • Look for a UTF-8 byte-order mark only if your toolchain expects one. A BOM is not a substitute for a declaration and can confuse older tools.
  • If the file came from a database or API, verify that the conversion happened once. Double-decoding UTF-8 often produces mojibake.

Declare UTF-8 at the start of the document

Put this inside <head>, before stylesheets, scripts, or substantial text:

<meta charset="utf-8">

The older equivalent is:

<meta http-equiv="Content-Type" content="text/html; charset=utf-8">

An issue report on Unicode input specifically notes that decoding failed until a UTF-8 content-type meta declaration was added. The exact spelling is less important than placing a valid declaration early and ensuring the bytes match it.

Make HTTP headers and document declarations agree

For a URL input, wkhtmltopdf receives an HTTP response before parsing the HTML. Inspect the response’s Content-Type, including its charset parameter. A server that sends charset=iso-8859-1 while the document is UTF-8 can cause the response metadata to win during parsing.

  1. Fetch the URL headers and record the exact Content-Type.
  2. Configure the server or application to send text/html; charset=utf-8 (or the appropriate UTF-8 media type).
  3. Keep the early HTML declaration as a second line of defense.
  4. Test the URL and a downloaded copy separately. If only the copy fails, preserve its bytes and account for the charset information that HTTP supplied to the original request.

Do not “fix” a server header by changing valid UTF-8 content to a legacy encoding. Correct the metadata or convert the content deliberately and consistently.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use wkhtmltopdf’s fallback encoding option

When input does not specify an encoding, the command-line fallback is:

wkhtmltopdf --encoding utf-8 input.html output.pdf

The official usage description defines --encoding <encoding> as the default text encoding for input. It is useful for incomplete documents, but it cannot repair incorrectly encoded bytes or override a contradictory HTTP response in every build.

Bindings and libraries

In libwkhtmltox-based wrappers, set the web setting named defaultEncoding to utf-8. That setting is used when content does not specify an encoding. The exact property syntax differs by language, but the value and purpose are the same:

web.defaultEncoding = "utf-8"

Also place the UTF-8 meta element in every template. Wrapper defaults vary between versions and distribution builds, so confirm the generated command or settings object when diagnosing a production-only failure.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Fix local files, standard input, and URL inputs separately

Local HTML files

Use a file that is actually UTF-8, contains an early declaration, and is passed with the fallback option when necessary:

wkhtmltopdf --encoding utf-8 ./input.html ./output.pdf

If local output differs from the URL, compare the downloaded file byte-for-byte with the server response body. A text editor, deployment step, or shell script may have transcoded it.

Standard input

When piping HTML, ensure the producer writes UTF-8 bytes and that the HTML itself declares UTF-8:

generate_html | wkhtmltopdf --encoding utf-8 - output.pdf

If the producer writes a locale-dependent encoding, set its output encoding explicitly rather than relying on the receiving process to guess.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

URLs

For a URL, correct the server’s response header first. You can still supply --encoding utf-8, but a reliable fix requires agreement among response bytes, HTTP metadata, and the HTML declaration.

When the problem is a missing font, not encoding

If Latin text is correct but Chinese, Japanese, Korean, emoji, or another script is absent, install a font containing those glyphs on the machine that runs wkhtmltopdf. A documented Ubuntu example for missing Chinese coverage is the fonts-wqy-zenhei package; package names and availability depend on your operating system and image.

  1. Install a font with the script and symbol coverage you need.
  2. Make sure the process user can read the font files.
  3. Declare a predictable CSS fallback stack, for example font-family: "Noto Sans CJK SC", "WenQuanYi Zen Hei", sans-serif;, using names actually installed on the host.
  4. Refresh the host’s font cache when your operating system requires it, then restart any long-running rendering workers.
  5. Render a test page containing representative characters, not just ASCII.

Changing --encoding cannot create a glyph that no installed font provides. Conversely, adding a font will not correct mojibake caused by decoding the bytes incorrectly.

Headers and footers are independent encoding inputs

Command-line header and footer text can fail while the page body is perfect. Treat dynamic header/footer content as HTML and declare UTF-8 in that document:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
<!-- footer.html -->
<!doctype html>
<html><head>
  <meta charset="utf-8">
</head><body>
  <span>Résumé — 日本語 — 😀</span>
</body></html>
wkhtmltopdf --encoding utf-8 
  --footer-html footer.html 
  input.html output.pdf

An issue report describes UTF-8 footer HTML as a working approach for non-ASCII footer text. Keep the footer file’s bytes, declaration, and fonts correct just as you do for the main page.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

A repeatable diagnostic decision tree

  1. Everything is garbled: verify source bytes, the HTTP charset (for URLs), and an early UTF-8 meta declaration. Then add --encoding utf-8.
  2. Only a downloaded local file fails: compare it with the URL response and check whether the download or editor changed its encoding.
  3. Only CJK or symbols fail: inspect installed fonts and CSS fallback before changing encoding settings.
  4. Only header/footer text fails: move dynamic content into UTF-8 header/footer HTML and declare UTF-8 there.
  5. A framework integration fails: set the wrapper’s encoding option and include the UTF-8 meta element in every template; log the effective settings used by the worker.

Common errors and practical fixes

What you see Cause Fix
é instead of é UTF-8 bytes decoded as a single-byte encoding, often twice Stop the extra conversion, save as UTF-8, align HTTP and HTML declarations
Black or empty squares No glyph in the selected fonts Install a font with coverage and set a CSS fallback stack
URL succeeds; file fails Local copy differs or lacks response charset context Compare bytes and declarations; regenerate the file as UTF-8
Body works; footer fails Footer is parsed separately Use UTF-8 footer HTML and a font that covers its characters
--encoding utf-8 changes nothing Bytes are wrong, a header conflicts, or the issue is font coverage Return to byte/header/font checks instead of stacking more flags
Works on a laptop, fails in a container Different fonts, locale, wkhtmltopdf build, or wrapper settings Install required fonts in the image and record versions/settings in deployment logs

Reliability and deployment checks

  • Keep a small Unicode fixture containing accented Latin, CJK, combining marks, right-to-left text if relevant, and emoji supported by your chosen fonts.
  • Render that fixture in CI whenever the base image, wkhtmltopdf package, templates, or font packages change.
  • Log whether input was a URL, file, or stdin, the response Content-Type, the effective default encoding, and the installed font package versions.
  • Use the same rendering image or host configuration in development and production; issue reports are build- and platform-specific evidence, not universal guarantees.
  • Do not assume that successful browser display proves PDF rendering will match. Browsers may have different font fallback and HTML parsing behavior.

Or skip the browser setup

If your actual goal is a clean image or PDF of a web page rather than debugging a local wkhtmltopdf pipeline, ScreenshotNeo provides a single HTTP request. Before capture it accepts cookie/consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each cleanup step can be disabled. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and every response identifies the result with X-Page-Verdict and X-Billed headers. It also offers an MCP server for AI agents and supports PNG, JPEG, WebP, and PDF output.

cURL

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

See the ScreenshotNeo documentation for output and capture options. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 shots, and every feature is available on every plan. Create a free ScreenshotNeo account.

FAQ

Does adding <meta charset="utf-8"> convert an old-encoded file?

No. It only tells the parser how to decode the existing bytes. Convert the source to UTF-8 first.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Why do boxes remain after I set UTF-8?

Boxes normally indicate missing glyphs. Install a suitable font on the rendering host and select it with CSS.

Should I always use the HTTP-equivalent meta tag?

Either valid UTF-8 declaration is suitable. The important factors are early placement and agreement with the actual bytes and HTTP response.

Frequently Asked Questions

Does adding <meta charset="utf-8"> convert an old-encoded file?

No. It only tells the parser how to decode the existing bytes. Convert the source to UTF-8 first.

Why do boxes remain after I set UTF-8?

Boxes normally indicate missing glyphs. Install a suitable font on the rendering host and select it with CSS.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Should I always use the HTTP-equivalent meta tag?

Either valid UTF-8 declaration is suitable. The important factors are early placement and agreement with the actual bytes and HTTP response.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.