Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

To create a searchable PDF with wkhtmltopdf, convert HTML that contains real text, then open the PDF and verify that you can select and search for a phrase. A successful conversion only proves that a PDF was written; it does not prove that its words are searchable. If your source is a scan or consists of page images, wkhtmltopdf will render those images but will not recognize their words. OCR is needed to add machine-readable text.

What makes a PDF searchable?

A PDF can look like a document while containing only pictures of its pages. A reader can see those pictures, but there are no underlying words to select or find. A searchable PDF contains text that a PDF reader can locate and, typically, select and copy.

wkhtmltopdf converts HTML page objects into PDF using Qt WebKit. When the HTML contains actual text, the generated PDF can retain text for search and selection. Searchability is not guaranteed for every input or configuration, so validate the file you create rather than treating a zero exit status or a newly written PDF as proof.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • HTML with text: This is the appropriate input for a text-searchable output. Test the generated PDF in a reader.
  • Scanned or image-only pages: These contain pixels, not text for wkhtmltopdf to preserve. Run OCR on the pages or use an OCR workflow, then check the result.

Convert HTML to PDF

Convert a local HTML file

Install a wkhtmltopdf build suitable for your operating system and distribution, then run this command from a shell:

wkhtmltopdf input.html output.pdf

Replace input.html with the path to your HTML file and output.pdf with the destination path. The source file should contain normal HTML text for the words you want to search. If the file references stylesheets, fonts, images, or other resources, make sure those resources are available to the renderer in the environment where the command runs.

Convert a web page

The command accepts a URL as a page input as well as a local HTML file. For example:

wkhtmltopdf https://example.com/ output.pdf

Use a URL you are authorized to access. A web page may load its content or resources differently from a local file, and the resulting PDF depends on what the renderer can load. Check that the PDF contains the intended page content before relying on its text.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

More than one page object

The command-line interface supports page objects, covers, tables of contents, and options scoped to a page or to the overall document. The exact options you need depend on the document; consult the installed binary’s help and test the features in the same environment used for deployment. Do not assume every package exposes every feature: patched and unpatched Qt builds differ, and some multi-input features require patched Qt.

Verify that the PDF is searchable

  1. Open the output PDF in a reader. Confirm that the expected pages and content are present.
  2. Select a phrase. Try to highlight a few words in the body text. If the reader selects only a large rectangle or the whole page as an image, the page may be image-only.
  3. Search for a distinctive phrase. Use the reader’s Find or Search command with a phrase that appears in the HTML. Confirm that the reader finds it in the expected location.
  4. Check representative pages. For a multi-page document, inspect pages with different content and layouts; one searchable page does not prove that every page contains selectable text.
  5. Use text extraction as a second check if needed. An extraction tool can help establish whether text is present, but still inspect the document in a reader: extraction alone does not tell you whether the page layout or content is correct.

If selection and search work for the intended text, the practical test has passed for the pages checked. If they do not, determine whether the input was real HTML text, an image, or a mix of the two before changing conversion settings.

Scanned pages and OCR

wkhtmltopdf is an HTML-to-PDF renderer, not an optical character recognition system. If a page is a scan, the renderer can place its image in a PDF, but it cannot infer the words from the pixels. To make scanned material searchable, apply OCR to produce a text layer and then verify that the output supports selection and search. Depending on the workflow, OCR may happen before or after PDF generation; the important point is that OCR, not HTML rendering alone, supplies the recognized words.

A document can also mix live HTML text and images. In that case, text portions may be searchable while text embedded in images is not. Check the particular regions readers need to find; a successful search in the HTML-generated headings does not establish that captions, diagrams, or scanned inserts have been recognized.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Choose and check the wkhtmltopdf build

Version and platform

The project’s downloads page identifies version 0.12.6 as its stable series and gives June 11, 2020 as its release date. That is a dated project statement, not confirmation that a package is currently available or recommended for every system. Check the project’s current download information and select a package for your operating system and distribution. Avoid assuming one generic binary will behave identically everywhere.

Before deployment, check the installed binary’s version and help output, then run a test conversion using the exact options and inputs your application needs. A distribution-provided package may be built against unpatched Qt; the project documents feature differences between patched and unpatched builds. Its source also errors when a build using unpatched Qt is asked to process more than one input document. If your workflow needs multiple page objects, covers, a table of contents, or another build-sensitive capability, verify that exact operation on the target machine.

Dependencies and fonts

The project’s downloads FAQ cautions that “static” refers to Qt linking, not to every dependency. Other system packages may still be required. Installed fonts, fontconfig, and freetype affect runtime behavior and how text is rendered. A conversion that works on a developer workstation may produce different results or fail in a deployment image with different libraries or fonts.

  • Match the package to the target operating system and distribution.
  • Check the installed binary’s version and help output rather than relying on a package name alone.
  • Install and test the fonts the document needs in the actual runtime environment.
  • Exercise all required page-object and document options in that environment before release.

Security: do not pass untrusted HTML directly

The wkhtmltopdf project warns: “Do not use wkhtmltopdf with any untrusted HTML – be sure to sanitize any user-supplied HTML/JS, otherwise it can lead to complete takeover of the server on which it is running!” Treat user-provided HTML and JavaScript as unsafe input. Sanitize it before conversion, and do not assume that requesting a PDF makes hostile markup harmless. This is especially important when the converter runs on a server with access to sensitive files or services.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup

If your actual job is to capture a web page visually—not to create a text-searchable PDF—ScreenshotNeo offers a website screenshot API and MCP server. A screenshot image should not be treated as a substitute for verifying searchable text, and the product information does not establish that its PDF output is searchable. For an API capture, one GET request can return an image; the following example saves a WebP response. See the ScreenshotNeo API documentation for setup and options.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp

ScreenshotNeo removes cookie and consent banners, newsletter popups, and chat widgets before capture; each of those steps can be turned off. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and response headers report the page verdict and billing status. Its MCP server provides take_screenshot, get_page_info, and capture_pdf tools for AI agents and MCP clients. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 shots.

Sign up for ScreenshotNeo’s free plan to try visual page capture.

Troubleshooting

The PDF exists, but Find returns no text

First determine whether the source HTML actually contains text or whether the visible content is an image. If it is image-only or scanned, use OCR; changing HTML-to-PDF options does not recognize words inside pixels. If the input contains text, check a phrase directly in the HTML and test the generated PDF by both selection and search.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Only some content is searchable

The document may combine HTML text with image-based text. Check the affected page and region, not just a heading elsewhere. OCR is needed for words embedded in scans or pictures.

A feature or multi-input conversion fails

Check the installed build and its help output. Patched versus unpatched Qt affects supported features, and the project’s source documents an error for more than one input document when built against unpatched Qt. Test a build that supports the required operation on the deployment platform.

The output differs across machines

Compare operating system packages, system libraries, and installed fonts. The project’s downloads FAQ notes that a static package can still depend on other system packages, and font availability can affect rendering. Reproduce the conversion in the target environment instead of assuming workstation results will carry over.

The converter processes unsafe content

Do not feed untrusted HTML or JavaScript directly to wkhtmltopdf. Sanitize user-supplied content before conversion, following the project’s security warning.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Practical checklist

  • Use HTML containing real text for content that must be searchable.
  • Use OCR for scanned or image-only words.
  • Run wkhtmltopdf input.html output.pdf for a basic local-file conversion.
  • Check selection and search in a PDF reader, including representative pages.
  • Verify build features, dependencies, and fonts on the target platform.
  • Sanitize untrusted HTML and JavaScript before conversion.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.