Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
For a large HTML document that uses JavaScript, web fonts, images, or modern CSS, render it with headless Chromium through Puppeteer. Navigate with an explicit readiness condition, wait for application data, images, and fonts, apply print CSS, and write the PDF to a file or stream. Use Chrome’s --print-to-pdf for a simple published URL, and choose WeasyPrint when the document is mostly static and paged CSS matters more than browser behavior.
Choose the engine before you write code
The right converter depends on what your HTML needs at render time.
| Engine | Best fit | JavaScript | Readiness control | Output options |
|---|---|---|---|---|
| Puppeteer with Chromium | Applications with client-side data, web fonts, images, and modern browser CSS | Yes | Navigation waits, selectors, network-idle waits, font and image checks, custom scripts | File path, byte data, or a readable stream |
| Chrome headless CLI | A controlled, already-published URL | Yes, using Chrome’s normal page loading | Very limited command-line control | PDF file |
| WeasyPrint | Static or server-rendered HTML where CSS Paged Media is central | Do not assume browser JavaScript behavior | Primarily CSS and document input | PDF bytes or a file through its API |
Puppeteer’s Page.pdf() generates PDF using the print CSS media type. Its Page.createPDFStream() API returns a readable stream, which is useful when the next service can consume chunks instead of holding the complete artifact in application memory.
The Tool Desk
Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Prepare the HTML for printing
Make print CSS explicit
Put print-only rules in your document or stylesheet. Hide navigation, buttons, sticky controls, and other interactive chrome. Define paper size and margins with @page, then tell Puppeteer to honor that declared size with preferCSSPageSize: true.
<style>
@page { size: A4; margin: 16mm 14mm 18mm; }
@media print {
nav, .toolbar, .chat-widget, .no-print { display: none !important; }
a { color: inherit; text-decoration: none; }
.avoid-break { break-inside: avoid; }
-webkit-print-color-adjust: exact;
print-color-adjust: exact;
}
</style>
Puppeteer uses print media by default, so test the page with print emulation rather than assuming the screen layout will be preserved. Use color-adjust properties when exact background and text colors are important; printing can otherwise alter colors.
#1 Best Overall
- PORTABLE SCANNER FOR USE ON-THE-GO — The fastest and lightest mobile single-sheet-fed compact document scanner in its class¹
- QUICK DOCUMENT SCANNING ― This Epson ultra-fast scanner scans a single page as quickly as 5.5 seconds²; Windows and Mac compatible
- VERSATILE PAPER HANDLING ― Portable scanner scans documents up to 8.5 x 72 in; Also easily digitizes receipts and ID cards to make accounting, bookkeeping, and organizing simpler
- INTUITIVE, HIGH-SPEED SOFTWARE — Epson ScanSmart Software³ is a smart tool allowing you to easily scan, review, and save; Stay organized easily with the help of this Epson scanner
- EASY SETUP — USB-powered connect to your computer for quick and simple scanning; No batteries or external power supply required to operate portable document scanner; Standard Connectivity: USB 2.0
Make asset URLs resolvable
- Use absolute HTTPS URLs or a stable base URL for images, stylesheets, and fonts.
- Ensure the rendering identity can access private assets through headers, cookies, or an authenticated session.
- Set explicit image dimensions where possible so late layout shifts do not move content between pages.
- Do not rely on a short fixed delay as proof that asynchronous data is complete.
Reliable Puppeteer conversion
Install Puppeteer in a Node.js project with npm install puppeteer. The following script navigates, waits for application readiness, waits for fonts and images, applies print settings, and writes a PDF. Replace the URL and readiness selector with values from your application.
const puppeteer = require('puppeteer');
async function htmlToPdf() {
const browser = await puppeteer.launch({ headless: true });
const page = await browser.newPage();
try {
await page.goto('https://example.com/large-report', {
waitUntil: 'networkidle2',
timeout: 120000
});
await page.waitForSelector('[data-report-ready="true"]', {
timeout: 120000
});
await page.evaluate(async () => {
if (document.fonts) await document.fonts.ready;
const images = Array.from(document.images);
await Promise.all(images.map(image => {
if (image.complete) return Promise.resolve();
return new Promise(resolve => {
image.addEventListener('load', resolve, { once: true });
image.addEventListener('error', resolve, { once: true });
});
}));
});
await page.emulateMediaType('print');
await page.pdf({
path: 'large-report.pdf',
format: 'A4',
printBackground: true,
preferCSSPageSize: true,
displayHeaderFooter: false,
margin: { top: '16mm', right: '14mm', bottom: '18mm', left: '14mm' },
timeout: 120000
});
} finally {
await page.close();
await browser.close();
}
}
htmlToPdf().catch(error => {
console.error(error);
process.exitCode = 1;
});
The waitUntil: 'networkidle2' condition is only one signal: an application can finish its network activity before rendering data, and long-polling applications may never become idle. A readiness selector emitted by your own application is stronger. The explicit font and image check covers assets that are present in the DOM but not yet usable. Puppeteer’s PDF guide states that, by default, Page.pdf() waits for fonts to load; keeping the explicit check makes the rest of the document’s readiness rules visible and testable.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Important PDF options
pathwrites directly to a file. Omit it when your code will consume returned bytes or a stream.formatselects a standard paper size such as A4 or Letter. Usewidthandheightfor a custom sheet.preferCSSPageSizegives your@pagedeclaration priority over format or dimensions.printBackground: truepreserves CSS backgrounds and colored panels.landscape: truechanges orientation for wide tables or charts.marginsets top, right, bottom, and left printable margins.pageRangeslimits output to selected pages when you do not need the entire document.displayHeaderFooter, together with header and footer templates, adds repeating metadata; keep template CSS small and test page numbers on the target paper size.
Use a stream when the document is large
A file path is simple and usually the least surprising choice for a worker that can write to temporary or durable storage. If your HTTP response, object-storage client, or queue accepts chunks, use createPDFStream() so your application can process the readable stream instead of first assembling the whole PDF value.
const puppeteer = require('puppeteer');
async function streamPdf(res) {
const browser = await puppeteer.launch({ headless: true });
const page = await browser.newPage();
try {
await page.goto('https://example.com/large-report', {
waitUntil: 'networkidle2', timeout: 120000
});
await page.waitForSelector('[data-report-ready="true"]', { timeout: 120000 });
await page.evaluate(() => document.fonts ? document.fonts.ready : Promise.resolve());
const pdfStream = await page.createPDFStream({
format: 'A4', printBackground: true, preferCSSPageSize: true
});
for await (const chunk of pdfStream) res.write(Buffer.from(chunk));
res.end();
} finally {
await page.close();
await browser.close();
}
}
A stream is an interface-level memory strategy, not a promise of a particular percentage reduction. Chromium still has to lay out the page, and a very long document can require substantial renderer memory. Measure representative files in your deployment environment.
Bound work in a production service
- Timeouts: set navigation, selector, and PDF timeouts. Return a diagnostic error instead of leaving a worker occupied indefinitely.
- Concurrency: cap simultaneous browser pages and jobs. A queue with a small worker pool is safer than launching an unlimited browser per request.
- Isolation: render untrusted HTML in an isolated browser context or separate worker, restrict network access where appropriate, and do not expose internal credentials to page scripts.
- Cleanup: close pages, contexts, and browsers in
finallyblocks, including error paths. - Temporary storage: write to a uniquely named temporary file, verify it, then move it to durable storage.
- Observability: record URL, document identifier, timings, browser errors, console messages, and the readiness condition that succeeded or timed out.
The official Puppeteer documentation does not publish a universal maximum HTML size, page-count limit, or memory ceiling. Establish service limits by testing the largest real reports you expect, including image-heavy and font-heavy cases.
Rank #2
- FAST SPEEDS - Scans color and black and white documents a blazing speed up to 16ppm (1). Color scanning won’t slow you down as the color scan speed is the same as the black and white scan speed.
- ULTRA COMPACT – At less than 1 foot in length and only about 1. 5lbs in weight you can fit this device virtually anywhere (a bag, a purse, even a pocket).
- READY WHENEVER YOU ARE – The DS-640 mobile scanner is powered via an included micro USB 3. 0 cable allowing you to use it even where there is no outlet available. Plug it into you PC or laptop and you are ready to scan.
- WORKS YOUR WAY – Use the Brother free iPrint&Scan desktop app for scanning to multiple “Scan-to” destinations like PC, Network, cloud services, Email and OCR. (2) Supports Windows, Mac and Linux and TWAIN/WIA for PC/ICA for Mac/SANE drivers. (3)
- OPTIMIZE IMAGES AND TEXT – Automatic color detection/adjustment, image rotation (PC only), bleed through prevention/background removal, text enhancement, color drop to enhance scans. Software suite includes document management and OCR software. (4)
Validate the PDF instead of trusting a successful process exit
After conversion, check that the file exists and is non-empty. For important reports, inspect the expected page count and search for key text. Verify that representative images and non-system fonts appear. A successful browser call can still produce a blank page, an early-loading state, or missing assets.
Free tools Windows power users keep installed
One-click scans. No signup required.
- Check the file size is greater than zero.
- Open the PDF with a parser or viewer and confirm the expected number of pages.
- Search extracted text for a title, report identifier, and final section heading.
- Compare a sample of pages visually, especially pages containing tables, charts, page breaks, and background colors.
- Keep browser and page logs for failed jobs so the next run identifies whether navigation, readiness, assets, or PDF writing failed.
Chrome headless for a simple published URL
When no HTML injection, custom headers, application-state wait, or stream handling is required, Chrome’s command-line mode is the shortest path:
chrome --headless --print-to-pdf https://developer.chrome.com/
Chrome saves the target page as output.pdf, according to the official headless command-line reference. Use Puppeteer when you need to wait for a selector, set print options, inject content, configure headers or cookies, add headers and footers, or control the response pipeline.
When WeasyPrint is the better choice
WeasyPrint is a print-layout engine for static or server-rendered HTML. Its API documents PDF hyperlinks, bookmarks, attachments, and CSS Paged Media features including @page selectors, page size, bleed, marks, named pages, page counters, running elements, and footnotes. It also documents limitations in some generated-content features. Do not choose it expecting full browser JavaScript execution.
from weasyprint import HTML
HTML('report.html', base_url='.').write_pdf('report.pdf')
Use it when the server has already produced the complete document and your priority is deterministic paged CSS. Use Chromium when the browser must execute application code or load modern web components before printing. The WeasyPrint API reference lists the supported paged-media behavior and limitations.
PC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchRank #3
- STAY ORGANIZED – Easily convert your paper documents into digital formats like searchable PDF files, JPEGs, and more.Power Consumption : 2.5W or less (Energy Saving Mode: 0.7W). Suggested Daily Volume : 500 scans..Does it contain liquid: no
- CONVENIENT AND PORTABLE –lightweight and small in size, you can take the scanner anywhere from home offices, classrooms, remote offices, and anywhere in between
- HANDLES VARIOUS MEDIA TYPES – Digitize receipts, business cards, plastic or embossed cards, reports, legal documents, and more
- FAST AND EFFICIENT – No technical hurdles or complicated setups here; easily scan both sides of a document at the same time, in color or black-and-white, at up to 12 pages-per-minute, and with a 20 sheet automatic feeder
- BROAD COMPATIBILITY – Works with both Windows and Mac devices, be it laptop or computer
Common failures and fixes
Fonts are missing or substituted
Confirm that the font URL is reachable from the browser, that the response has the correct content type, and that the page waits for document.fonts.ready. Check that the font is not blocked by authentication or cross-origin policy.
Images are blank
Wait for every image’s load or error event, use resolvable URLs, and log failed requests. An image error should fail validation when the image is mandatory rather than silently producing an incomplete report.
The PDF captures a loading shell
Replace a guessed delay with an application-owned readiness selector or a script that verifies the data has rendered. Keep a generous navigation and selector timeout for slow environments.
Pages break through tables or cards
Add print rules such as break-inside: avoid to units that must remain together, use explicit heading and table styles, and test at the actual paper size and margins. Some very tall elements cannot fit on one page; design them to split safely.
Colors or backgrounds differ
Enable printBackground and use -webkit-print-color-adjust: exact where exact color reproduction is required. Verify the result in a PDF viewer and on the intended printer.
Rank #4
- IRIScan Express, portable scanner : scans color and black and white documents a blazing speed up to 8ppm simplex. Color scanning won’t slow you down as the color scan speed is the same as the black and white scan speed.
- IRIScan Express mobile scanner is powered via an included micro USB 2. 0 cable allowing you to use it even where there is no outlet available. Plug it into you PC or laptop and you are ready to scan. USB cable provided. AC Adapter not provided and not needed.
- IRIScan flatbed scanner uses a simplex scanning mode allows for quick and straightforward scanning of single-sided documents. IRIScan with its full portable features is the ideal document scanners for computers.
- IRIScan document scanner : Versatile scanning capabilities, including scanning to Word, PDF, and Excel formats with companion software provided Readiris OCR
- Receipt scanner and card scanner with Additional features include scanning business cards directly to Outlook, photo scanning, and receipt scanning for efficient document management
The worker runs out of memory
Reduce concurrency, avoid retaining HTML and PDF buffers simultaneously, stream output where practical, resize oversized source images, and split an extremely long report into deliberate sections. There is no documented universal Chromium memory limit, so use measurements from your own workload to set limits.
The process hangs
Set navigation and PDF timeouts, investigate pages with persistent connections, and ensure cleanup runs in finally. Capture console, request, and page-error logs before retrying.
Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server that can return a screenshot or PDF from one GET request. Its cleanup steps accept cookie and consent banners before capture and remove more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be turned off. Bot checks and CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and the response identifies the result with X-Page-Verdict and X-Billed headers.
For a URL capture, the documented call is:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
The same endpoint supports PDF output; select the PDF response option documented at ScreenshotNeo’s API documentation rather than inventing browser orchestration in your service. Equivalent client examples are:
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
ScreenshotNeo also exposes an MCP server for Claude, Cursor, and other MCP clients, with take_screenshot, get_page_info, and capture_pdf tools. Its 63 options include full-page capture with lazy images loaded, CSS-selector element capture, dark mode, device presets and arbitrary viewports, retina scale, PDF paper size, margins, landscape and page ranges, custom CSS and JavaScript, click and hide selectors, waits for selectors, delays or network idle, request and resource blocking, custom headers, cookies, user agents and Authorization, timezone and geolocation, transparent backgrounds, image resizing, chosen cache TTLs, signed links, asynchronous jobs with signed webhooks, bulk capture for 100 URLs per call, a usage API, an OpenAPI specification, and compatibility with parameter names used by other screenshot APIs.
The Free plan includes 1,000 screenshots per month with no card. Paid plans are Starter $5 for 3,000, Growth $15 for 15,000, Pro $39 for 60,000, Scale $99 for 250,000, and Business $249 for 1,000,000; yearly billing gives two months free, and every feature is on every plan. Create a free ScreenshotNeo account to start with 1,000 screenshots a month and no card.
Best Value
- Scanner type: Document
- Connectivity technology: USB
- With Auto Scan Mode, the scanner automatically detects what you're scanning
- Digitize documents and images
Frequently Asked Questions
Can a generated PDF preserve links and bookmarks?
Yes. WeasyPrint’s documented API includes PDF hyperlinks and bookmarks. Verify the result in your target viewer when using another engine.
Do these 3 things before closing this tab:
1Repair Windows errors before they cause bigger problems2Scan for outdated or missing drivers - takes under a minute3Clear out junk files and repair common Windows errorsShould I split one very long report into multiple PDFs?
Split it when your measured renderer memory, timeout, or review workflow requires a bound on job size; keep each section’s headings and identifiers explicit so the parts can be reassembled or audited.
Is network-idle waiting enough for every web application?
No. Applications with delayed rendering, long polling, or client-side data hydration need an application-specific readiness selector or validation script.
The Bottom Line
Use Puppeteer and Chromium for dynamic, asset-heavy HTML; make readiness and print CSS explicit, stream or write output according to your memory budget, and validate the artifact. Use Chrome CLI for a straightforward public URL and WeasyPrint for static, paged documents.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.
Quick wins for a faster PC:
Clear out junk files and repair common Windows errorsFree Scan →Scan for outdated or missing drivers - takes under a minuteDriver Scan →Repair Windows errors before they cause bigger problemsFix Now →

