Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
urllib3 downloads HTML; it does not turn that HTML into a PDF. Use it to retrieve the page, then pass the response text to a renderer such as WeasyPrint or xhtml2pdf. The key details are checking the HTTP status, decoding the response using its declared character set when available, and supplying the original page URL so relative CSS, images, and fonts can resolve.
What urllib3 does—and what it does not do
urllib3 is an HTTP client. It can request a page and return its response body, but it does not run a browser or implement PDF layout. A PDF renderer must interpret the HTML and its styles and produce the file. urllib3 documents its pool and request workflow in its User Guide.
This distinction matters because downloading a page source is not the same as printing the page as a browser would. A renderer may support only part of modern browser CSS, and HTML that depends on client-side JavaScript may not contain the final content in the downloaded response.
Convert a page with urllib3 and WeasyPrint
WeasyPrint is a practical choice when CSS layout, web fonts, images, and external stylesheets matter. Its Python API accepts HTML as a string, a base URL for resolving relative resources, and a destination for the resulting PDF. The example below checks for an HTTP error before rendering and falls back to UTF-8 if the response does not declare a charset.
#1 Best Overall
import urllib3
from weasyprint import HTML
url = "https://example.com/page"
http = urllib3.PoolManager()
response = http.request("GET", url)
try:
if response.status >= 400:
raise RuntimeError(f"HTTP {response.status} while fetching {url}")
content_type = response.headers.get("content-type", "")
charset = "utf-8"
for part in content_type.split(";")[1:]:
name, sep, value = part.strip().partition("=")
if sep and name.lower() == "charset":
charset = value.strip().strip('"'')
break
html_text = response.data.decode(charset, errors="replace")
HTML(string=html_text, base_url=url).write_pdf("page.pdf")
finally:
response.release_conn()
Install the Python packages in the environment you will run the script from. WeasyPrint may also require platform libraries, so follow its installation guidance for your operating system in WeasyPrint First Steps. The API supports URLs, files, file objects, and in-memory strings; when using a string, base_url gives relative references their context.
Why the base URL matters
If the downloaded HTML contains <img src="/assets/mark.png"> or a relative stylesheet link, that path is not meaningful by itself after the source becomes an in-memory string. Setting base_url to the page URL lets the renderer resolve those references against the page’s location. Without it, the PDF may contain text but omit styles, images, or fonts.
Charset handling
The response body is bytes. Decoding every page as UTF-8 can garble characters if a server declares another encoding. The example checks the HTTP Content-Type charset and uses UTF-8 only as a fallback. For unusual or inconsistent sites, inspect the actual response headers and HTML rather than assuming that every declaration is correct.
Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchPC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Rank #2
Use xhtml2pdf as an alternative renderer
If you want a direct pisa.CreatePDF call or a mostly Python-based pipeline, xhtml2pdf is another option. It supports HTML5, CSS 2.1, and some CSS 3, but complex modern layouts should be tested against the PDF you need. Its API documents the destination stream, base path, encoding, callbacks, and resource policy at the Python API reference.
from xhtml2pdf import pisa
with open("page.pdf", "wb") as output:
result = pisa.CreatePDF(
html_text,
dest=output,
path="https://example.com/page",
encoding="utf-8",
raise_exception=True,
)
Here, html_text is the decoded HTML from the urllib3 example. The path supplies a base for resource paths; for more control over resource lookup, xhtml2pdf provides a link_callback and resource_policy. See its advanced usage guide for the documented pattern of writing an HTML string to a binary PDF file and checking conversion status.
Choose the renderer for your page and operating environment
| Consideration | WeasyPrint | xhtml2pdf |
|---|---|---|
| CSS and layout | Prefer it when CSS layout, web fonts, images, and external stylesheets are important; validate the rendered result. | Documents HTML5, CSS 2.1, and some CSS 3 support; verify complex modern CSS against your output needs. |
| Relative resources | Pass base_url for an HTML string. Its fetcher can be replaced for additional controls. |
Pass a base path or use link_callback to map resource locations. |
| Authentication and fetch control | The default URL fetcher does not provide advanced cookies or authentication. A custom URL fetcher can add headers, cookies, authentication, or timeouts. | Use a callback and resource policy to control how linked resources are resolved and fetched. |
| Runtime and batch considerations | For many documents, the Python API avoids repeated process startup costs compared with invoking a command for each file. | No comparable speed figure is established in the cited documentation; test your own pages and deployment environment. |
Neither choice has a universal speed advantage established by the cited official documentation. Layout fidelity, dependency installation, resource access, and throughput should be evaluated using representative pages from your own workload.
Make CSS, images, fonts, and links resolve correctly
- Set the resource base: use the requested page URL as WeasyPrint’s
base_urlor xhtml2pdf’spath. For custom URL mapping, use the renderer’s fetcher or callback. - Check the downloaded document: a server-rendered HTML response can be converted directly; content inserted only by browser JavaScript may not be present in that response.
- Check asset access separately: an asset may require cookies, authorization, or special headers even if the HTML request succeeds. WeasyPrint’s default fetcher does not supply advanced authentication; configure a custom fetcher where required.
- Decide how to handle missing assets: renderers may produce a PDF even when a stylesheet, font, or image fails to load. For production workflows, capture warnings or validate required assets and fail explicitly when omissions are unacceptable.
- Distinguish HTML links from PDF links: a base URL helps resolve resources used while rendering. It does not guarantee that every HTML feature or browser behavior will be reproduced in the PDF.
Restrict resource access for untrusted HTML
HTML rendering can trigger additional requests for stylesheets, images, and other resources. User-supplied markup could point at local files, internal services, or private network addresses. Do not give an unrestricted renderer access to arbitrary resources.
Quick wins for a faster PC:
Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Clear out junk files and repair common Windows errorsFree Scan →Scan for outdated or missing drivers - takes under a minuteDriver Scan →- WeasyPrint: use a custom URL fetcher that permits only approved URL schemes and hosts when rendering untrusted content. Its documentation describes replacing the URL fetcher and resource handling at First Steps.
- xhtml2pdf: use the available host and resource restrictions, such as
--allow-host,--resource-root, or--no-remote, as appropriate to your deployment. The CLI documentation also describes private-network behavior and its opt-in at the CLI reference. - Application policy: do not enable broad filesystem or network access just to make one missing image appear. Allow only the resource locations the conversion job actually needs.
Troubleshoot common conversion failures
| Symptom | Likely cause | What to do |
|---|---|---|
| The script reports an HTTP error | The server returned an error status, or the requested URL redirects to a failing destination. | Check the status and final URL behavior before invoking the renderer. Keep the explicit status check so an error page is not silently saved as the intended document. |
| Characters display incorrectly | The response was decoded using the wrong charset. | Read the response’s declared charset where present and inspect the HTML declaration if the server headers and content disagree. |
| Images, fonts, or CSS are missing | Relative URLs have no base, resources require credentials, or the renderer cannot fetch them. | Set base_url or path; then check resource URLs and authentication. For WeasyPrint, use a custom fetcher if headers, cookies, or authentication are required. |
| The page looks different from a browser | The renderer does not support a layout feature the page uses, or the final page depends on client-side JavaScript. | Check whether the fetched HTML contains the final content, simplify or adapt the print CSS, or test the other renderer with the same representative page. |
| Conversion fails on a resource URL | A resource policy blocks that host or path, or unsafe remote access is being denied. | Review allowed hosts, schemes, and resource roots. Expand access only to an approved location rather than disabling restrictions wholesale. |
| Batch jobs take longer than expected | Repeated process startup or repeated resource fetching may add overhead. | For WeasyPrint, prefer its Python API for many documents where appropriate; measure throughput on your own document corpus rather than relying on an undocumented speed comparison. |
Performance, reliability, and operating costs
Conversion time depends on the HTML, CSS, number and size of remote assets, renderer, and execution environment. The official documentation cited here does not provide a comparable benchmark establishing a speed winner between WeasyPrint and xhtml2pdf. Benchmark with the pages, asset availability, and concurrency levels your application will actually use.
For repeated WeasyPrint conversions, using the Python API avoids starting a separate command-line process for every document. Reusing an urllib3 PoolManager for repeated HTTP requests also keeps the retrieval layer organized around a pool. This does not remove renderer work or guarantee a particular throughput.
Reliability depends on treating retrieval and rendering as distinct stages: reject unsuccessful HTTP responses, preserve encoding, provide a resource base, and decide whether missing assets are warnings or job failures. Network timeouts and retries should be configured to suit the application rather than allowing a batch to hang indefinitely; renderer support for resource-fetch timeouts and authentication can be extended with a custom WeasyPrint fetcher.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Or skip the browser setup
If you need a clean image or PDF capture of a live URL rather than a local Python conversion pipeline, ScreenshotNeo provides a one-request screenshot API and an MCP server. Its capture options include PDF output, custom headers and cookies, and waiting for a selector, delay, or network idle.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com/page -o shot.webp
See the ScreenshotNeo API documentation for request parameters and output options. Cookie and consent banners are accepted and removed before the shot, along with supported newsletter popups and chat widgets; these steps can be turned off. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and response headers identify the page verdict and billing status. Its MCP server lets AI agents use screenshot tools. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 shots.
Sign up for ScreenshotNeo’s free plan to try 1,000 screenshots a month with no card.
Best Value
Frequently Asked Questions
Does urllib3 itself create a PDF?
No. urllib3 retrieves the HTTP response; a renderer such as WeasyPrint or xhtml2pdf must create the PDF.
Can this method capture content added by JavaScript?
Not if that content is absent from the HTML response urllib3 downloads. The conversion uses retrieved source, not a browser’s executed page state.
Free tools Windows power users keep installed
One-click scans. No signup required.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

