Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
To keep Cyrillic readable in an HTML-to-PDF conversion, preserve the text as Unicode, use a font that contains every Cyrillic character in the document, and make that font available to the renderer. For WeasyPrint, configure fonts and share one FontConfiguration between CSS and PDF generation; for Puppeteer, wait for the page’s web fonts before calling page.pdf(). Then check both how the PDF looks and whether its text can be selected, copied, and searched.
Why Cyrillic turns into squares, blanks, or incorrect text
HTML can contain valid Cyrillic Unicode while the PDF renderer still fails to draw it. Rendering crosses several boundaries: the HTML must be decoded correctly, the renderer must find a font with the needed glyphs, and the PDF must preserve usable text. A browser may silently use a locally installed fallback font that is absent from the server or conversion container.
- Squares or blank glyphs: the renderer did not find a usable font containing those characters.
- Only some letters fail: the chosen bold, italic, or other font face may not cover the same code points as the regular face.
- The browser is fine but the PDF is not: the browser and converter may have different installed fonts or font-loading access.
- Text looks right but cannot be copied: check the PDF’s Unicode text mapping and the renderer’s font embedding behavior, not only its appearance.
WeasyPrint’s documentation specifically advises installing fonts and making them available to WeasyPrint when the PDF contains squares or no drawn characters: WeasyPrint font documentation.
Do these 3 things before closing this tab:
1Clear out junk files and repair common Windows errors2Fix the driver behind crashes, sound loss and screen glitches3Repair Windows errors before they cause bigger problemsStart with Unicode and a font that covers the text
Keep the HTML as UTF-8
Declare UTF-8 in the document and ensure the program reading or serving the HTML uses the same encoding. A declaration cannot repair text that has already been mis-decoded or converted through a legacy character set.
#1 Best Overall
<!doctype html>
<html lang="ru">
<head>
<meta charset="utf-8">
<title>Пример</title>
</head>
<body>
<p>Здравствуйте! Проверка русского текста.</p>
</body>
</html>
If the input comes from a file, database, or HTTP response, verify how that input is decoded before passing it to the renderer. Correct HTML markup is not enough if the text has become replacement characters earlier in the pipeline.
Select and test the complete font stack
Choose a family with coverage for the Cyrillic characters actually used. Keep a fallback family in the CSS stack, and confirm that the regular, bold, and italic faces used in the document all cover the target script. A family name alone does not guarantee that every face or weight has the same glyph coverage.
body {
font-family: "Your Cyrillic-Capable Font", sans-serif;
}
Use a real font installed in the conversion environment, or provide a reachable font resource. For repeatable server-side output, packaging a local font with the application avoids dependence on a remote font server or a developer’s workstation.
Generate a PDF with WeasyPrint
WeasyPrint embeds fonts in PDF files and subsets them by default to the glyphs used in the PDF, according to its documentation: WeasyPrint font handling. When using CSS @font-face, configure the font and pass the same FontConfiguration to both the CSS object and HTML.write_pdf(). The following example uses a local font file; adjust the path and family to match the installed font.
Rank #2
from pathlib import Path
from weasyprint import CSS, HTML
from weasyprint.text.fonts import FontConfiguration
html = """<!doctype html>
<html lang="ru">
<head>
<meta charset="utf-8">
<style>
@font-face {
font-family: "CyrillicFont";
src: url("file:///app/fonts/cyrillic-font.ttf") format("truetype");
font-weight: 400;
font-style: normal;
}
body { font-family: "CyrillicFont", sans-serif; }
</style>
</head>
<body>
<p>Здравствуйте! Проверка русского текста.</p>
</body>
</html>"""
font_config = FontConfiguration()
css = CSS(string="", font_config=font_config)
HTML(string=html, base_url=Path("/app").as_uri()).write_pdf(
"output.pdf",
stylesheets=[css],
font_config=font_config,
)
For an external stylesheet that contains @font-face, construct it with the same configuration, for example CSS(filename="styles.css", font_config=font_config), and supply that CSS object and configuration to write_pdf(). The font URL must resolve from the renderer’s environment; use a valid local file URL or an accessible URL. See the official WeasyPrint example and font notes.
WeasyPrint reports missing glyphs through warnings and displays the missing-glyph placeholder when neither the selected face nor its fallbacks provides a character. Treat those warnings as conversion errors worth investigating, rather than assuming that a generated PDF means every character rendered.
Generate a PDF with Puppeteer
Puppeteer’s page.pdf() renders using print CSS, so print-specific rules and the web fonts loaded by the page both affect the PDF. Wait for the document’s fonts to be ready before generating output, and select print or screen media intentionally. Puppeteer documents these PDF behaviors in its Page.pdf() API reference.
import puppeteer from 'puppeteer';
const browser = await puppeteer.launch();
try {
const page = await browser.newPage();
await page.goto('https://example.com/report', { waitUntil: 'networkidle0' });
await page.evaluate(() => document.fonts.ready);
await page.pdf({
path: 'report.pdf',
format: 'A4',
printBackground: true,
});
} finally {
await browser.close();
}
Replace the example URL with the page you control. If the page uses a remote web font, make sure the conversion environment can reach its URL and that the page’s CSS selects the intended family in print. For deterministic output, bundle the font with the page or ensure that the remote resource is reliably available before PDF generation.
Verify appearance and searchable text
- Open the PDF in a viewer. Inspect Cyrillic text in regular, bold, and italic styles and on the pages where those styles occur.
- Copy a sentence. Paste it into a plain-text editor and check that the characters remain Cyrillic rather than turning into blanks, question marks, or unrelated symbols.
- Search for a known word. A visually correct page is not sufficient if the text layer is missing or mapped incorrectly.
- Review renderer output. Check font-load failures, inaccessible resource errors, and missing-glyph warnings.
For archival output, WeasyPrint documents PDF/A-3u; its “u” variant indicates that PDF text is available as Unicode. This is a relevant option when Unicode text availability is part of the archival requirement, not a substitute for validating the actual document. See WeasyPrint PDF/A documentation.
Troubleshooting by symptom
Every Cyrillic character is a square or blank
Install a font with the required glyphs or supply it through a working @font-face. Confirm the renderer process—not only your interactive browser—can read the file or fetch the URL. In WeasyPrint, check the documented font availability guidance and warnings.
Some letters or styles are wrong
Check the exact characters and font face in use. A regular face may work while bold or italic falls back to a face without coverage. Define the needed weight and style faces explicitly, retain a fallback family, and inspect the rendered pages that use each face.
Recommended Free Tools
The browser page looks correct, but the PDF does not
Assume the environments differ until proven otherwise: the browser may have a local font the server-side converter lacks. Package or install the font where conversion happens. With Puppeteer, wait for document.fonts.ready and check print CSS, since page.pdf() uses print styles.
Rank #4
- Funny saying for any front-end developer, web developer, computer programmer, computer systems engineer, mobile app developer, software developer, or code lover who likes to code, make funny programming jokes, and take memorable photos.
- Wear it proudly at International Programmers' Day, school, coding classes, or coding communities! It also makes a funny present for a computer programming lover friend.
- Lightweight, Classic fit, Double-needle sleeve and bottom hem
A remote font is ignored
Verify the font URL from the renderer’s network context, including access restrictions and redirects. If reliable access cannot be guaranteed, use a packaged local font and point @font-face to it.
The text looks right but copying or search fails
Inspect the PDF’s text layer and the renderer’s Unicode/font behavior. WeasyPrint embeds and subsets fonts by default, but visual appearance alone does not establish that extraction works. Test copied text and consider PDF/A-3u when Unicode availability is an archival requirement.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Choose the renderer around font control and verification
For this problem, compare renderers by how they load fonts, whether you can control fallbacks and shaping, what print CSS behavior they provide, whether the PDF contains searchable Unicode text, and what diagnostics they expose for missing glyphs. The key operational decision is whether the exact font files and styles used in production are available to the conversion process; do not infer that from a local browser preview.
Free tools Windows power users keep installed
One-click scans. No signup required.
Or skip the browser setup
For a page screenshot rather than a searchable text PDF, ScreenshotNeo can return an image or PDF from one request. It is a website screenshot API and MCP server for developers. Its clean-shot flow accepts cookie or consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, with response headers identifying the page verdict and billing status. AI agents can use its MCP server’s take_screenshot, get_page_info, and capture_pdf tools. It does not replace a font-configured HTML-to-PDF workflow when Unicode text extraction or archival PDF/A output is required.
Example cURL request for a page capture (replace the URL and provide your API key):
Best Value
- Programming Language Lover Code Apparel. App or Web Design and Development Expert Funny Dress. Best Valentines Idea For Coding Lover. HTML Code or Meaning Costume
- Funny I Know HTML - How To Meet Ladies Computer Programmer Quotes
- Lightweight, Classic fit, Double-needle sleeve and bottom hem
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp
See the ScreenshotNeo API documentation for output options and request parameters. One thousand screenshots a month are free without a card; paid plans start at $5 for 3,000. Sign up for ScreenshotNeo’s free 1,000 screenshots a month, with no card required.
FAQ
Does declaring UTF-8 alone prevent missing Cyrillic glyphs?
No. UTF-8 preserves the character data; a font with the needed glyphs must also be available to the renderer.
Can I use a web font instead of installing a system font?
Yes, if the conversion process can reach and load it. For WeasyPrint’s @font-face path, use a shared FontConfiguration for the CSS and PDF generation.
Does a generated PDF prove the text is searchable?
No. Open it, copy a Cyrillic sentence, and search for a known word to validate the text layer separately from appearance.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

