Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

If wkhtmltoimage renders Unicode as question marks or empty boxes, check two separate things: whether the text is decoded as UTF-8, and whether the runtime can find fonts containing the needed glyphs. Start by saving the HTML as UTF-8, declaring <meta charset="utf-8">, and trying --encoding UTF-8. If the characters are present but their shapes or joining are wrong, the bundled legacy Qt WebKit engine may be the limitation rather than the character encoding.

What the symptoms mean

Unicode rendering problems often look alike but have different causes. A question mark may mean that text was decoded using the wrong character set or was already lost before it reached the renderer. A box usually means the renderer has no usable glyph for that character in its available fonts. Text can also have the right characters and glyphs but still display incorrectly when a script needs shaping, such as Arabic joining or Indic character composition.

Separate these checks instead of changing settings at random. First verify the source bytes and decoding; then verify font availability and fallback; finally assess the renderer’s shaping behavior. A UTF-8 declaration addresses character decoding, not missing fonts or limitations in the browser engine.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Prepare the HTML and input data as UTF-8

Declare the document encoding

Put a charset declaration near the beginning of the document head, before content that depends on correct decoding:

<!doctype html>
<html lang="en">
<head>
  <meta charset="utf-8">
  <title>Unicode test</title>
  <style>
    body { font-family: "Noto Sans", "DejaVu Sans", sans-serif; }
  </style>
</head>
<body>
  <p>English — Ελληνικά — Русский — 中文 — العربية — हिन्दी — 日本語</p>
  <p>Emoji test: 😀</p>
</body>
</html>

Save the file itself as UTF-8. A declaration in the document cannot recover characters that were previously converted incorrectly or replaced with question marks by the program that generated the HTML. If your application receives bytes, decode them explicitly as UTF-8 before inserting the text into the document.

Set the renderer encoding explicitly

Try this command with a local HTML file:

wkhtmltoimage --encoding UTF-8 input.html output.png

A user report in a wkhtmltopdf project issue from 2018 says adding --encoding 'UTF-8' fixed that reporter’s Unicode problem. Treat it as a diagnostic step, not a guarantee for every failure: it can help with decoding, but it cannot install a missing font or correct a shaping limitation.

If the HTML is produced dynamically, preserve the same principle in the code that creates or passes it to the renderer. Do not rely on a machine’s default locale or an implicit narrow-string conversion. The libwkhtmltox documentation says settings passed to its PDF and image C bindings use UTF-8 encoded strings. That concerns the strings supplied to the bindings; it does not by itself prove that every page’s source bytes or every font is correct.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Check fonts separately from encoding

A renderer can decode a character correctly and still be unable to draw it. The Qt 4 internationalization documentation explains that language display also requires suitable JIS or Unicode fonts. Qt’s whitepaper describes combining installed fonts for multilingual text, but fallback can work only when appropriate fonts are installed and discoverable by the account running wkhtmltoimage.

  • If Latin text works but a CJK character or another script appears as a box, check font coverage before changing encoding options.
  • Use a CSS font stack with likely installed fallbacks, as in the example above. The named fonts must actually be available in the runtime environment.
  • Check the same container, server, or user account that runs the capture. Fonts visible in an interactive desktop session may not be installed or discoverable in a service environment.
  • Test the exact deployment image, not only a developer workstation. Minimal server and container environments may have fewer fonts.

A wkhtmltopdf fallback-font issue describes investigation of missing glyphs despite a UTF-8 meta declaration. That illustrates why changing the charset is not a substitute for checking glyph coverage.

Use a diagnostic fixture that distinguishes the failures

Keep a small HTML test page containing several scripts and one emoji. Use the same file, binary, runtime user, container image, and font packages as the failing production job. A useful fixture should include ordinary Latin text with an accent, a CJK character, Arabic, an Indic script, and an emoji. The example HTML above is a starting point; add the exact characters that fail in your own page.

  1. Confirm the original text. Inspect the HTML as text or bytes and verify that the characters are present and encoded as UTF-8. If the source already contains replacement question marks, the renderer cannot restore the original text.
  2. Check the document head. Make sure the UTF-8 declaration appears early, before dependent content, and that the file was saved in that encoding.
  3. Render with an explicit setting. Run wkhtmltoimage --encoding UTF-8 and record the precise binary version and command used.
  4. Classify the output. Distinguish question marks from boxes, and both from characters that appear individually but have incorrect ordering, joining, or composition.
  5. Verify fonts under the runtime identity. Confirm that the relevant font files are installed and discoverable to the same account and environment as the rendering process; use an appropriate CSS fallback stack.
  6. Compare environments. Run the same fixture in development and production with the same image, font packages, locale, and binary. A difference points toward environment configuration rather than the fixture alone.
  7. Investigate shaping last. If the text decodes correctly and the needed glyphs are available, but Arabic joining, Indic shaping, combining marks, or emoji remain wrong, investigate the legacy Qt WebKit engine bundled with the renderer.

Keep the test input and output with a short record of the command, version, runtime user, environment, and installed fonts. That makes comparisons meaningful when a server image or renderer installation changes. Do not infer that an encoding fix worked merely because a different run used a different machine or font set at the same time.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

When the command-line flag is not enough

The encoding setting is worth trying when text appears to have been decoded incorrectly. It is not a universal Unicode switch. If a required glyph is absent, the renderer still has nothing to draw; install a font with the necessary coverage and make sure it is available to the process. If a script’s glyphs exist but are not shaped or joined correctly, adding more charset declarations is unlikely to address the underlying problem.

At that point, decide whether the project needs the existing renderer’s behavior or whether another rendering path is acceptable. The project issue discusses missing glyphs and WebKit problems for emoji, but it does not establish one setting that fixes those cases generally. Test the exact scripts and symbols your application needs before relying on any renderer for production output.

Qt integration: avoid implicit byte conversion

If you call the rendering library through Qt code, make byte-to-text conversion explicit. Qt’s Qt 4 encoding guidance notes that constructing a QString from const char * can interpret the bytes as Latin-1, while QString::fromUtf8() explicitly decodes UTF-8. The general lesson for wrappers is to pass a Unicode string or an explicitly UTF-8-encoded byte sequence through the binding, rather than a locale-dependent narrow string.

This check belongs upstream of rendering: inspect the value at the point where your application turns incoming bytes into a string and again where it passes data or settings to the binding. If the string is already corrupted there, changing an HTML meta tag later will not repair it.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup

If your goal is to capture a public webpage rather than maintain a local wkhtmltoimage installation, ScreenshotNeo is a website screenshot API and MCP server. It is an alternative capture workflow, not a claim that an API flag fixes every font or shaping issue in a local renderer. Check the resulting image with the scripts your page needs.

One GET request returns an image or PDF. For example, this cURL request saves a WebP screenshot of the target page:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

See the ScreenshotNeo documentation for request options. Cookie and consent banners are accepted like a visitor and removed along with 60+ known consent platforms, newsletter popups, and chat widgets; each of those steps can be turned off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and responses include X-Page-Verdict and X-Billed headers. Its MCP server gives AI agents tools for screenshots, page information, and PDF capture. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 screenshots.

Sign up for 1,000 free screenshots a month, with no card required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Troubleshooting by symptom

What you see Likely area to check Next step
Question marks replace the intended text Input bytes, application decoding, or renderer encoding Verify the original UTF-8 bytes and explicit decoding; declare the charset and test --encoding UTF-8.
Empty boxes for particular scripts Font coverage or font discovery Install a font that contains the glyphs, check runtime-user visibility, and configure CSS fallbacks.
Some scripts work but Arabic or Indic text looks malformed Shaping behavior or renderer engine After confirming bytes and fonts, reproduce with a minimal fixture and assess whether the bundled legacy Qt WebKit engine meets the requirement.
It works on a workstation but fails in a container or server Different fonts or runtime environment Compare the exact image, font packages, locale, binary, and account used by each run.
Qt-generated text is already wrong before capture Implicit conversion from bytes to a string Use explicit UTF-8 decoding such as Qt 4’s QString::fromUtf8() rather than relying on a narrow-string conversion.

What to record before changing production

For a repeatable fix, record the input fixture, exact command, wkhtmltoimage version, runtime identity, environment image, locale, and font packages alongside the rendered result. Change one factor at a time: encoding first, then font availability and fallback, then shaping and engine suitability. This avoids mistaking a font or environment change for an encoding fix and gives you a focused regression test for the scripts your application actually uses.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.