What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
The most practical way to build an HTML-to-PDF converter in Node.js is to render the document in an isolated Chromium worker with Puppeteer, then call page.pdf(). Put a validation layer in front of it, make print CSS explicit, wait for fonts and images, and treat every submitted HTML string or URL as untrusted input.
Use a controlled rendering pipeline
A reliable converter is more than one browser call. Separate it into five stages so limits, security checks and operational failures are visible:
- API layer: accept a server-side template plus data whenever possible. If callers must submit HTML, accept only the MIME types and features your product needs.
- Validation: cap payload bytes, nesting depth, CSS and asset sizes, page count, output bytes and render time. Sanitize user-authored markup.
- Renderer worker: run Puppeteer or Playwright in a low-privilege, isolated process. Set content or navigate only to an approved origin, wait for deterministic readiness, fonts and images, then generate the PDF.
- Response: return
application/pdfwith an intentionalContent-Disposition, or enqueue a job and store the result for a later download when documents are large. - Operations: apply queue backpressure and concurrency limits, recycle unhealthy browsers, remove temporary files, and emit structured logs and metrics.
Templates are safer for multi-tenant products because callers supply data rather than executable page code. A free-form HTML endpoint should be considered a browser execution service, not a text formatter.
Choose the rendering engine
| Engine | Best fit | Trade-offs |
|---|---|---|
| Puppeteer | Chromium fidelity and JavaScript-heavy pages | Browser process cost plus sandbox and network hardening |
| Playwright | Chromium rendering with a broader browser-automation toolset | Similar worker, isolation and resource concerns |
| wkhtmltopdf | Simple CLI deployments and legacy WebKit-compatible layouts | Older rendering engine; verify modern CSS and JavaScript compatibility |
| PDFKit | Structured, data-driven documents with direct programmatic layout | Not an HTML/CSS renderer; you position content yourself |
Puppeteer and Playwright generate PDFs with the print CSS media type. Puppeteer documents that Page.pdf() waits for fonts by default. If your source is already an HTML/CSS page and fidelity matters, a browser engine is usually the shortest path. Choose PDFKit when you want a coordinate-driven document model and do not need to render arbitrary HTML.
Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minuteWindows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstall#1 Best Overall
Build a minimal Node.js converter
1. Install the application
Use a current Node.js LTS release, create a project, and install Express and Puppeteer:
mkdir html-pdf-api
cd html-pdf-api
npm init -y
npm install express puppeteer
Puppeteer downloads a compatible Chromium during installation. In a restricted build environment, make sure that browser is available to the runtime image and that the process has enough shared memory or uses the flag shown below.
2. Add the conversion endpoint
This runnable example accepts JSON such as {"html":"<h1>Invoice</h1>"}, renders it with print styles, and returns a downloadable PDF:
import express from 'express';
import puppeteer from 'puppeteer';
const app = express();
app.use(express.json({ limit: '1mb' }));
app.post('/convert', async (req, res, next) => {
if (!req.body || typeof req.body.html !== 'string' || req.body.html.length === 0) {
return res.status(400).json({ error: 'html must be a non-empty string' });
}
const browser = await puppeteer.launch({
headless: true,
args: ['--disable-dev-shm-usage']
});
try {
const page = await browser.newPage();
await page.setContent(req.body.html, { waitUntil: 'networkidle0' });
await page.emulateMediaType('print');
const pdf = await page.pdf({
format: 'A4',
printBackground: true,
preferCSSPageSize: true,
tagged: true,
timeout: 30000
});
res
.type('application/pdf')
.set('Content-Disposition', 'attachment; filename="document.pdf"')
.send(pdf);
} catch (err) {
next(err);
} finally {
await browser.close();
}
});
app.use((err, req, res, next) => {
if (res.headersSent) return next(err);
res.status(500).json({ error: 'pdf_render_failed' });
});
app.listen(3000, () => {
console.log('HTML-to-PDF API listening on http://localhost:3000');
});
Set "type": "module" in package.json or convert the imports to CommonJS. Start it with node server.js, then post JSON to http://localhost:3000/convert. The example launches a browser per request to stay easy to understand. A production service should keep a bounded browser pool, create a fresh page for each job, and recycle workers after crashes or a defined number of jobs.
Rank #2
3. Make readiness deterministic
networkidle0 is useful for pages whose requests eventually settle, but analytics, WebSockets and long polling can prevent it from completing. For controlled templates, prefer a readiness marker:
await page.setContent(html, { waitUntil: 'domcontentloaded' });
await page.waitForSelector('[data-pdf-ready]', { timeout: 10000 });
await page.evaluate(() => document.fonts.ready);
Use a bounded delay only when a third-party widget cannot expose a marker. Always retain a hard render timeout so a page cannot occupy a worker indefinitely. Wait for image completion when documents contain remote assets:
await page.evaluate(async () => {
const images = Array.from(document.images);
await Promise.all(images.map(image => image.complete
? Promise.resolve()
: new Promise(resolve => {
image.addEventListener('load', resolve, { once: true });
image.addEventListener('error', resolve, { once: true });
})));
});
Design print CSS for predictable pages
Put page geometry in @page and enable preferCSSPageSize when the document’s CSS should control the paper size:
@page {
size: A4;
margin: 16mm 14mm 18mm;
}
html, body {
margin: 0;
font-family: "Inter", Arial, sans-serif;
color: #111;
}
.invoice-header,
.invoice-total,
table tr {
break-inside: avoid;
}
h1, h2, h3 {
break-after: avoid;
}
.brand-panel {
print-color-adjust: exact;
}
Use break-before, break-after and break-inside to keep headings, invoice blocks and table rows together. Apply print-color-adjust: exact only where background colors are essential because it can increase output size. Embed or preload the exact fonts used by the document; a missing font changes line wrapping and pagination. Decide explicitly whether external images, web fonts and JavaScript are allowed. Self-hosted assets make output reproducible.
Free tools Windows power users keep installed
One-click scans. No signup required.
Rank #3
Accepting URLs or user HTML safely
Sanitize markup
User HTML is executable input in a browser context. Sanitize it before rendering, remove event-handler attributes, reject dangerous URL schemes, and never treat a submitted string as a trusted application template. Enforce maximum nesting, CSS size, image dimensions, page count, render duration and output bytes. If your product can use templates, expose a template identifier and validated data instead of accepting arbitrary markup.
Defend against SSRF
A converter that fetches a URL is a server-side network client. Prefer an identifier or an allowlisted host over a complete user URL. Resolve DNS and block loopback, link-local, private, metadata and other internal ranges. Re-check the destination after redirects, prevent protocol changes, and restrict outbound egress from the renderer. Do not allow browser jobs to reach cloud credentials or internal administration panels.
Isolate Chromium
Run workers as a low-privilege user in a separate container or sandbox with a read-only filesystem, no cloud credentials and tightly restricted egress. Chromium’s sandbox and Site Isolation are defensive layers, not replacements for application-level validation. Avoid logging raw HTML or generated PDFs by default; encrypt stored results, use short retention and scrub temporary files.
Scale the service without losing reliability
Synchronous versus asynchronous jobs
Keep a synchronous endpoint for small documents with a strict deadline. For large files or busy tenants, return 202 Accepted with a job identifier, expose status, and store the finished PDF in object storage for a later download. Apply per-tenant quotas and queue backpressure so one customer cannot consume every browser slot.
Rank #4
Concurrency and recycling
Limit pages and browser processes according to available CPU and memory rather than accepting unlimited parallel requests. Recycle a browser after crashes, repeated timeouts or a controlled number of jobs. Terminate jobs that exceed memory, page-count, duration or output-size limits.
Observability and upgrades
Return stable error classes such as invalid_html, blocked_url, timeout, renderer_crash and output_too_large. Record job duration, queue delay, page count, output bytes and renderer version without storing document contents. Keep a fixture corpus covering long tables, RTL text, Unicode, charts, headers and footers, then compare PDF text, page count and rasterized snapshots whenever Chromium, fonts or CSS dependencies change. Include the renderer version in job metadata so a layout change can be traced.
Troubleshoot common failures
| Symptom | Likely cause | Fix |
|---|---|---|
| Request never finishes | Network-idle is blocked by polling, WebSockets or an unresolved asset | Use domcontentloaded plus a readiness selector, allow only required requests, and enforce a hard timeout |
| Fonts or line breaks differ | Font failed to load or a different font version is installed | Self-host or preload fonts, await document.fonts.ready, and pin the renderer image |
| Backgrounds are missing | Print CSS disables backgrounds | Set printBackground: true and use print-color-adjust: exact selectively |
| Rows split across pages | No print break rules | Apply break-inside: avoid to rows or blocks and test unusually long content |
| Browser crashes under load | Too many concurrent pages, large images or insufficient shared memory | Lower concurrency, cap asset dimensions, recycle workers and use --disable-dev-shm-usage where appropriate |
| Internal services are reachable | URL fetching permits private addresses or redirects | Allowlist hosts, block private ranges before and after redirects, and restrict container egress |
| Downloads have the wrong behavior | Missing or incorrect response headers | Send application/pdf and an explicit inline or attachment Content-Disposition |
Performance, cost and fidelity decisions
- Browser cost: Chromium startup and memory dominate small jobs, so reuse a bounded pool instead of launching unlimited browsers.
- Fidelity: Browser engines execute modern CSS and JavaScript; wkhtmltopdf may require layout changes because it uses an older WebKit engine.
- Determinism: Self-host fonts and images, pin browser versions, and avoid time-dependent scripts in templates.
- Large documents: Queue them, cap page count and output bytes, and store results outside the API process.
- Testing: Compare both semantic output (text and page count) and rendered snapshots; a successful HTTP response does not prove the layout is correct.
Or skip the browser setup
ScreenshotNeo provides a website screenshot API and MCP server; its PDF endpoint can capture a URL without you operating Chromium workers. One GET request returns a PDF or image, and the response identifies the page verdict and billing status with X-Page-Verdict and X-Billed headers.
For a one-call PDF capture, see the ScreenshotNeo API documentation:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
In this example the output filename is shot.webp; request PDF output using the documented PDF parameter for your integration. The same service also supports full-page capture with lazy images loaded, CSS-selector element capture, dark mode, device presets or custom viewports, retina scale, paper size and margins, page ranges, custom CSS and JavaScript, click and wait actions, hidden selectors, blocked requests or resource types, custom headers, cookies, user agents, authorization, timezone and geolocation, transparent backgrounds, resizing, selectable cache TTLs, signed links, asynchronous jobs with signed webhooks, bulk capture of up to 100 URLs per call, a usage API and an OpenAPI specification. Parameter names used by other screenshot APIs are accepted to ease migration.
Cookie and consent banners, newsletter popups and chat widgets are removed before the shot. Bot checks, blank pages, failed loads and timeouts are not billed, and cache hits are not billed either. An MCP server exposes take_screenshot, get_page_info and capture_pdf to Claude, Cursor and other MCP clients. The Free plan includes 1,000 shots each month with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account to try it without a card.
Frequently Asked Questions
Should a converter accept complete user URLs?
Only when there is a strong product reason. An allowlisted host or internal document identifier is safer because a complete URL turns the renderer into an SSRF-capable network client.
When is PDFKit a better choice than Chromium?
Choose PDFKit when the document is structured data and you want to place every element programmatically. Choose a browser engine when existing HTML and CSS, responsive layout or client-side JavaScript are part of the source.
Quick wins for a faster PC:
Scan for outdated or missing drivers - takes under a minuteDriver Scan →Repair Windows errors before they cause bigger problemsFix Now →What should be included in a renderer upgrade test?
Keep fixtures for long tables, RTL and Unicode text, charts, headers and footers, then compare extracted text, page count and rasterized pages after changing Chromium, fonts or CSS dependencies.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

