What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

Short answer: PDFKit does not provide a documented method that parses an HTML string and lays it out as HTML/CSS. Calling doc.text('<h1>Hello</h1>') writes the tag characters as text; it does not create a heading. To use PDFKit, translate the content into PDFKit text, image, table and drawing operations, then pipe the document stream to a writable destination and call doc.end(). If you need browser-style HTML/CSS rendering, use an HTML-to-PDF renderer instead.

Can you pass an HTML string directly to PDFKit?

No. PDFKit’s documented text API accepts strings for methods such as doc.text(), but it is a PDF-generation library, not a browser layout engine. It does not document a general HTML-string parser, CSS cascade, DOM layout system or JavaScript runtime. The text API documentation is at pdfkit.org/docs/text.html.

For example, this code produces a PDF containing the literal characters <h1>Hello</h1>:

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
const PDFDocument = require('pdfkit');
const fs = require('node:fs');

const doc = new PDFDocument();
doc.pipe(fs.createWriteStream('literal-tags.pdf'));
doc.text('<h1>Hello</h1>');
doc.end();

PDFKit will apply its own text wrapping, fonts, margins and paragraph options to that string. It will not infer that <h1> should be larger, bold or separated from a paragraph.

The normal PDFKit document flow in Node.js

PDFKit creates a readable Node.js stream. It does not save a file automatically. The supported flow is to create a PDFDocument, pipe it to a writable stream, add content with PDFKit methods, and call doc.end() to finalize the document. See the official guide at pdfkit.org/docs/getting_started.html.

Minimal runnable example

  1. Create a project and install PDFKit:

    mkdir pdfkit-example
    cd pdfkit-example
    npm init -y
    npm install pdfkit
  2. Save this as make-pdf.js:

    const fs = require('node:fs');
    const PDFDocument = require('pdfkit');
    
    const doc = new PDFDocument({ margin: 50 });
    const output = fs.createWriteStream('output.pdf');
    
    doc.pipe(output);
    doc.fontSize(20).text('Hello from PDFKit');
    doc.moveDown();
    doc.fontSize(12).text('This layout was built with PDFKit operations, not HTML parsing.');
    doc.end();
    
    output.on('finish', () => {
      console.log('Wrote output.pdf');
    });
  3. Run it with node make-pdf.js. Wait for the writable stream’s finish event before treating the file as complete.

For HTTP responses, pipe to the response instead of a file and set headers before writing:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
const http = require('node:http');
const PDFDocument = require('pdfkit');

http.createServer((req, res) => {
  if (req.url !== '/report.pdf') {
    res.writeHead(404).end('Not found');
    return;
  }

  res.writeHead(200, {
    'Content-Type': 'application/pdf',
    'Content-Disposition': 'inline; filename="report.pdf"'
  });

  const doc = new PDFDocument();
  doc.pipe(res);
  doc.fontSize(18).text('Report');
  doc.fontSize(11).text('Generated as a PDF stream.');
  doc.end();
}).listen(3000);

Three practical ways to handle an HTML string

1. Treat it as text when markup is not needed

If the string is untrusted or you only need its readable words, remove or decode markup before passing the result to PDFKit. A regular expression is only a small, controlled-input solution; it is not a complete HTML parser and can mishandle scripts, malformed markup and entities.

const PDFDocument = require('pdfkit');
const fs = require('node:fs');

function basicTextFromHtml(html) {
  return html
    .replace(/<script[sS]*?</script>/gi, '')
    .replace(/<style[sS]*?</style>/gi, '')
    .replace(/<[^>]+>/g, '')
    .replace(/s+/g, ' ')
    .trim();
}

const html = '<h1>Hello</h1><p>World</p>';
const doc = new PDFDocument();
doc.pipe(fs.createWriteStream('text-only.pdf'));
doc.text(basicTextFromHtml(html));
doc.end();

This intentionally loses headings, links, lists, colors and layout. Use a real HTML parser if the input can be complex or user supplied, and sanitize it before any conversion.

2. Map known HTML elements to explicit PDFKit operations

For templates you control, define a small supported subset and map each element yourself. This gives predictable output and keeps the PDF independent of a browser.

const fs = require('node:fs');
const PDFDocument = require('pdfkit');

const doc = new PDFDocument({ margin: 54 });
doc.pipe(fs.createWriteStream('mapped.pdf'));

function renderArticle({ title, paragraphs, bullets }) {
  doc.font('Helvetica-Bold').fontSize(22).text(title);
  doc.moveDown(0.6);

  doc.font('Helvetica').fontSize(11);
  for (const paragraph of paragraphs) {
    doc.text(paragraph, { paragraphGap: 8, lineGap: 2 });
  }

  doc.moveDown(0.4);
  for (const bullet of bullets) {
    doc.text(`• ${bullet}`, { indent: 12, paragraphGap: 4 });
  }
}

renderArticle({
  title: 'A mapped document',
  paragraphs: ['The h1 became a bold title.', 'Each p became a PDFKit text call.'],
  bullets: ['A list item', 'Another list item']
});

doc.end();

A production mapper normally handles headings, paragraphs, links, lists, images, tables, page breaks and inline emphasis separately. Keep the mapping explicit: for example, call font() and fontSize() for a heading, image() for an image, and drawing methods for rules or boxes. PDFKit’s feature overview describes text, images, tables and vector drawing at pdfkit.org.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

3. Render HTML with a browser-style HTML-to-PDF tool

Choose this route when the source already depends on CSS layout, responsive rules, web fonts, flexbox, grid, JavaScript, external assets or print-specific styles. An HTML renderer can load the document and produce a PDF from its rendered layout; PDFKit cannot do that by itself.

A service called pdfkitt advertises accepting an HTML string or live URL in its API documentation (pdfkitt.dev/docs). The documentation surfaced for this topic does not establish its fidelity, runtime, security model, pricing or suitability, so evaluate those items before adopting it.

Choosing between PDFKit and an HTML renderer

Requirement PDFKit HTML-to-PDF renderer
Input model Explicit PDFKit calls for text, images, tables and drawing HTML/CSS document and, depending on the product, a browser-like runtime
Existing HTML template Must be parsed and mapped; tags are not interpreted Usually the native input model
CSS layout and JavaScript Not provided by the documented API Capability varies; verify it for the renderer you select
Deployment Node stream and PDFKit package May require a browser process, service or other runtime; verify operational requirements
Control Direct control over drawing and generated content Control comes from HTML/CSS and renderer options
Security surface Your own input parsing and asset handling Also consider HTML, JavaScript, network requests and renderer isolation

There is no universal winner. PDFKit is a good fit for reports whose structure you can model as PDF operations. An HTML renderer is the appropriate category when visual fidelity to an existing web page is the requirement. Compare font and asset handling, page-break controls, accessibility, privacy, runtime footprint and cost for the specific renderer; no general performance or quality figure is established here.

Handling common HTML features in a PDFKit mapper

Headings and paragraphs

Assign each heading level a font, size and spacing, then call text() for paragraph content. Use continued: true only when you deliberately join runs with different styles; otherwise separate calls are easier to reason about.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Links

PDFKit can draw link annotations, but your mapper must preserve the target URL and apply the annotation explicitly. Do not assume that an <a> tag passed to text() becomes a clickable link.

Images

Resolve each permitted image source, validate its type and size, and pass the resulting file path or buffer to PDFKit’s image operation. Remote fetching, redirects and untrusted data need timeouts, allowlists and size limits.

Lists and tables

Convert list items into indented text or drawn bullets. For tables, calculate column widths, draw borders and advance the vertical cursor row by row; decide how a row behaves when it reaches a page boundary. PDFKit exposes primitives, not automatic browser table layout.

Page breaks and headers

Track the current vertical position and create a new page before content collides with the bottom margin. Repeat headers explicitly on each page. Test long words, empty paragraphs, very tall images and rows that are taller than one page.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Reliability and security checklist

  • Always end the document. A missing doc.end() leaves the stream unfinished and commonly produces a zero-byte or unreadable file.
  • Handle stream errors. Listen for errors on both the PDF document and destination stream; handle disk-full, permission and client-disconnect failures.
  • Control input size. Limit HTML length, image dimensions, number of elements and total output time to prevent memory or CPU exhaustion.
  • Sanitize untrusted markup. Do not execute scripts or fetch arbitrary URLs while parsing or resolving assets. Use an allowlist for protocols, hosts and local files.
  • Make fonts and assets deterministic. Package the fonts you are licensed to use and use stable paths rather than relying on a developer workstation.
  • Test generated files. Open PDFs in more than one viewer and test text extraction, links, page count, Unicode, images and page boundaries.
  • Separate generation from delivery. For large jobs, write to a temporary file or object store and return a job result rather than holding an entire document in application memory.

Troubleshooting PDFKit HTML conversions

Symptom Likely cause Fix
HTML tags appear in the PDF The string was passed to text(), which treats it as text Strip markup for plain text, map supported elements, or use an HTML renderer
Output file is empty or truncated The document was not ended, or the process exited before the stream finished Call doc.end() and wait for the destination’s finish event
Styles are missing CSS is not interpreted by PDFKit Translate the styles into PDFKit calls or select an HTML/CSS renderer
Images fail to load Bad path, unsupported data, network failure or an unsafe URL policy Validate and log the resolved source, use supported image data, and enforce explicit fetch limits
Text overlaps or runs off the page Manual positioning ignored wrapping, font metrics or page boundaries Use PDFKit’s text layout options, measure content, and implement a page-break policy
HTTP response hangs The response is not receiving the PDF stream or an exception interrupted generation Pipe the document to the response, set headers first, and attach error handlers
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup

If your actual goal is a clean capture of a live, rendered web page rather than a hand-built PDF from PDFKit operations, ScreenshotNeo provides a website screenshot API and MCP server. One GET request returns a PNG, JPEG, WebP or PDF. It accepts cookie and consent banners before capture and removes more than 60 known consent platforms, newsletter popups and chat widgets; bot checks, blank pages, timeouts, failed loads and cache hits are not billed, and response headers identify the page verdict and billing status. AI agents can call its take_screenshot, get_page_info and capture_pdf MCP tools.

Here is the documented one-call pattern (the complete option list and parameter names are in the ScreenshotNeo docs):

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

The same request in Python is:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)

And in Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot failed: ${res.status}`);
require('node:fs').writeFileSync('shot.webp', Buffer.from(await res.arrayBuffer()));

ScreenshotNeo includes full-page capture with lazy images loaded, element selection, device and viewport controls, retina scale, PDF paper and margin settings, custom CSS and JavaScript, waits, request blocking, headers, cookies, user agent, timezone, geolocation, resizing, caching, signed links, asynchronous jobs, webhooks, bulk capture, usage reporting and an OpenAPI specification. Every feature is on every plan: 1,000 shots per month are free with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.

Frequently asked questions

Does PDFKit support SVG, and does that mean it supports HTML?

No. PDFKit documents SVG path syntax for vector geometry. That is drawing path data, not HTML parsing or CSS layout. See pdfkit.org/docs/vector.html.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Can I use a DOM parser and then feed the result to PDFKit?

Yes, but the parser only gives you a tree to inspect. You still need code that maps nodes and styles to PDFKit operations; parsing alone does not create layout.

When should I keep PDFKit instead of switching tools?

Keep it when you need programmatic, repeatable PDFs and can define the document as text, images, tables and drawings. Switch categories when pixel-level reproduction of an existing HTML/CSS page is the primary requirement.

Why is my PDF valid but visually different from the web page?

PDFKit is not running the page’s browser layout, CSS or JavaScript. Visual differences are expected unless your mapper explicitly reproduces the required structure and styling.

Frequently Asked Questions

Can PDFKit execute JavaScript embedded in an HTML string?

No. PDFKit is not a browser runtime and does not execute HTML scripts. Remove that dependency, map the resulting data into PDFKit calls, or use a renderer that explicitly supports JavaScript.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Is calling doc.end() optional when piping to a file?

No. Piping establishes the destination, but doc.end() signals that PDF generation is complete and allows the writable stream to finish.

Will an HTML anchor automatically remain clickable in the PDF?

No. A string containing an anchor tag is treated as text. Your mapper must create a PDF link annotation explicitly.

The Bottom Line

There is no documented HTML-string renderer in PDFKit. Use text() and the rest of PDFKit’s drawing APIs after translating your markup, or choose an HTML-to-PDF renderer when browser layout is the requirement.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.