Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

To export selected pages from an existing PDF in Node.js, use pdf-lib: load the source file, copy the desired pages into a new document, then save that document. The key detail is that copyPages takes zero-based page indices, while people usually number PDF pages starting at 1. For example, pages 1, 3, and 5 correspond to indices [0, 2, 4].

Export selected pages with pdf-lib

pdf-lib is a pure-JavaScript option for selecting and assembling PDF pages in a Node.js program. Its PDFDocument.copyPages(srcDoc, indices) API copies the requested page objects from one document into another. Add the returned pages to the destination document in order, then call save() to get the output bytes.

Install the package in your project:

npm install pdf-lib

Save the following as extract-pages.mjs and run it with Node.js:

import { readFile, writeFile } from 'node:fs/promises'
import { PDFDocument } from 'pdf-lib'

const inputPath = 'input.pdf'
const outputPath = 'selected-pages.pdf'

// Page numbers as a person would count them: 1, 3, and 5.
const pageNumbers = [1, 3, 5]

const inputBytes = await readFile(inputPath)
const source = await PDFDocument.load(inputBytes)
const pageCount = source.getPageCount()

// pdf-lib indices start at 0, so subtract 1 from each page number.
const indices = pageNumbers.map((pageNumber) => pageNumber - 1)

for (const index of indices) {
  if (!Number.isInteger(index) || index < 0 || index >= pageCount) {
    throw new RangeError(
      `Requested page ${index + 1}; input has ${pageCount} pages.`
    )
  }
}

const output = await PDFDocument.create()
const pages = await output.copyPages(source, indices)
for (const page of pages) {
  output.addPage(page)
}

const outputBytes = await output.save()
await writeFile(outputPath, outputBytes)
console.log(`Wrote ${outputPath} with ${pages.length} pages.`)

The output contains the selected pages in the same order as pageNumbers. For other uses, consult the pdf-lib project documentation and its API reference for copyPages.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Validate page numbers before copying

The example validates after converting page numbers to indices. Each requested page must be an integer between 1 and the source document’s page count, inclusive. This prevents a bad request from reaching the copy operation and gives callers a useful error that relates to the page numbering they provided.

If page numbers arrive from a UI, HTTP request, or database, validate their type and range at the input boundary. Do not silently round values such as 2.7 or accept strings without deciding explicitly whether your application should parse them.

Keep the requested order

Pass indices in the exact order you want in the output. For example, requesting pages 5, 1, and 3 means indices [4, 0, 2]; adding the returned page objects sequentially creates an output in that order. If you need to repeat a page, duplicate indices may be useful, but confirm that repeated pages are acceptable in your particular workflow.

Export a range or build a reusable selection function

For a contiguous range, construct the zero-based indices. Pages 4 through 7 become [3, 4, 5, 6]. Because page ranges are easy to get wrong at their boundaries, it is useful to put conversion and validation in a small function:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
function pageRange(firstPage, lastPage, pageCount) {
  if (!Number.isInteger(firstPage) || !Number.isInteger(lastPage)) {
    throw new TypeError('Page numbers must be integers.')
  }
  if (firstPage < 1 || lastPage < firstPage || lastPage > pageCount) {
    throw new RangeError(`Invalid page range ${firstPage}-${lastPage}.`)
  }
  return Array.from(
    { length: lastPage - firstPage + 1 },
    (_, offset) => firstPage - 1 + offset
  )
}

const indices = pageRange(4, 7, source.getPageCount())
const pages = await output.copyPages(source, indices)
for (const page of pages) output.addPage(page)

The conversion is inclusive at both ends: pageRange(4, 7, count) selects pages 4, 5, 6, and 7. An empty selection should generally be rejected by your application instead of producing an unexpectedly empty PDF.

What the copied PDF does—and does not—preserve

Copying page objects is not necessarily the same as cloning every document-level feature of the original. If your files use forms, outlines or bookmarks, annotations, encryption, or metadata, verify the result with representative input PDFs and the viewers or downstream systems your users rely on. The cited API documents page copying, but does not promise identical preservation of every such feature in every document.

This matters especially when a selected page refers to document structures beyond its visible page content. A basic text-and-image PDF may behave as expected, while an interactive or specially structured file can require separate handling or a different workflow. Treat fidelity as a tested requirement, not an assumption based solely on a successful save().

save() returns the complete output as bytes. The example writes those bytes to a file with Node’s promise-based filesystem API. In a web service, you may instead send the bytes in an HTTP response or write them to object storage; account for the memory use of loading and saving the complete documents when processing large files.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use qpdf when a command-line PDF tool fits better

qpdf’s command-line documentation describes --pages for selecting pages from one or more input PDFs. For a single input and pages 1, 3, and 5, run:

qpdf input.pdf --pages . 1,3,5 -- selected-pages.pdf

Here, the dot identifies the primary input as the page source; the comma-separated list uses human-style page numbers. qpdf also documents page ranges, reverse ordering, selection across files, and password handling for encrypted inputs. In normal mode, document-level information comes from the primary input; using --empty starts a new output and changes metadata behavior. Check the qpdf documentation for exact syntax and behavior for your installed version and use case.

A Node.js application can run qpdf with child_process, but it must ensure the executable is installed and available, validate all user-supplied arguments, and avoid building a shell command by concatenating untrusted input. Prefer an argument-array form such as execFile or spawn so arguments are passed separately. qpdf can be a good fit for a server image that already includes native PDF tooling; pdf-lib avoids that external executable dependency and keeps the work in-process.

Choose the approach for your deployment

Consideration pdf-lib qpdf
Deployment Pure JavaScript dependency; no separate PDF executable is required. Requires the qpdf native executable to be installed and discoverable.
Selection syntax Zero-based index array passed to copyPages. Command-line page selection syntax, including ranges.
Multiple input files Pages can be copied into a destination document; the application manages source loading and selection. Official CLI documentation explicitly covers selection from one or more inputs.
Operational concerns Runs in the Node.js process; save() returns document bytes. Adds process startup, executable discovery, argument validation, and native deployment concerns.
Feature fidelity Test forms, outlines, annotations, encryption, and metadata that matter to your files. Test the same document features with your actual inputs and required output behavior.

PDFKit is not the default choice for extracting pages from an existing PDF. Its official getting-started guide focuses on creating a PDF document and piping generated output to a writable stream; that workflow is useful for generating PDFs, but does not provide the existing-document page-copy workflow needed here.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Troubleshoot common failures

  • “Cannot find package ‘pdf-lib’” or module resolution errors: Install pdf-lib in the project where the script runs. If using import, run an .mjs file as shown, or configure your project for ES modules. Do not assume a dependency installed in a different project directory is available.
  • Input file not found: The example uses paths relative to the process’s current working directory. Confirm where the process starts, or pass an absolute path or resolve paths relative to the script deliberately.
  • Page index out of range: Convert one-based numbers to zero-based indices exactly once. Page 1 is index 0; the last valid index is pageCount - 1. Validate every requested page against the loaded document’s actual page count.
  • Output order is wrong: Check the order of the index array and append returned pages in sequence. Sorting indices will change a deliberately custom order.
  • The output opens but interactive details differ: Test forms, annotations, outlines, metadata, encryption, and any other document-level behavior that your workflow requires. A successful save alone does not establish complete fidelity.
  • Encrypted input cannot be loaded: Decide how your application obtains and handles authorization for that file. For command-line workflows, qpdf documents password options; do not log passwords or expose them in process listings or error output.
  • Large files consume too much memory or take too long: Loading and saving through pdf-lib in this example keeps document data in memory. Measure with files representative of your workload, limit concurrent jobs, and consider an established native PDF tool if its deployment model better fits the service.
  • qpdf works locally but not in production: Verify that the executable is installed in the production runtime and that the application can locate it. Pass validated arguments separately rather than interpolating user input into a shell string.

Or skip the browser setup

ScreenshotNeo is a website screenshot API, not a tool for extracting selected pages from an existing PDF. It is relevant if the input you have is a web page and you need to capture that page as an image or PDF instead. One GET request can return a screenshot or PDF; see the ScreenshotNeo API documentation for request options.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

ScreenshotNeo removes cookie and consent banners, newsletter popups, and chat widgets before capture; each cleanup step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing status. Its MCP server offers screenshot tools for AI agents, and the free plan includes 1,000 shots per month without a card; paid plans start at $5 for 3,000 shots.

Learn about ScreenshotNeo, or sign up for 1,000 free screenshots a month with no card.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.