Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minutePC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
To export selected PDF pages in Ruby, open the source with HexaPDF, create a new document, import the pages you want, and write the new file. Ruby arrays are zero-based, so source pages 1, 3, and 5 are indexes 0, 2, 4; the array order becomes the output order.
The basic workflow is short, but production code must also validate page bounds, handle encrypted or malformed files, and account for document-level features such as outlines, forms, attachments, links, and metadata.
Use HexaPDF for a Ruby-native export
HexaPDF’s documented merging pattern is to open a source document, create a target document, import source pages into that target, and write the result. For selected-page export, iterate over an explicit index list.
Recommended Free Tools
Install the gem
Add HexaPDF to your application:
gem install hexapdf
In a Bundler project, add gem "hexapdf" to your Gemfile, then run bundle install.
#1 Best Overall
Minimal runnable script
require "hexapdf"
input_path = "input.pdf"
output_path = "selected.pdf"
selected = [0, 2, 4] # source pages 1, 3, and 5
source = HexaPDF::Document.open(input_path)
target = HexaPDF::Document.new
selected.each do |index|
target.pages << target.import(source.pages[index])
end
target.write(output_path, optimize: true)
After it runs, selected.pdf contains three pages in this order: original page 1, original page 3, and original page 5. Reordering the array changes the result; [4, 0, 2] produces pages 5, 1, and 3.
Make page selection safe and predictable
Convert reader-friendly page numbers
People normally specify pages using one-based numbers. Convert them to Ruby indexes at the boundary of your program, then validate before touching the PDF.
require "hexapdf"
input_path = "input.pdf"
output_path = "selected.pdf"
requested_pages = [1, 3, 5] # one-based numbers from a user or API request
source = HexaPDF::Document.open(input_path)
page_count = source.pages.count
indexes = requested_pages.map do |page_number|
unless page_number.is_a?(Integer) && page_number.between?(1, page_count)
raise ArgumentError, "page #{page_number.inspect} is outside 1..#{page_count}"
end
page_number - 1
end
raise ArgumentError, "select at least one page" if indexes.empty?
target = HexaPDF::Document.new
indexes.each do |index|
target.pages << target.import(source.pages[index])
end
target.write(output_path, optimize: true)
puts "Wrote #{output_path} (#{indexes.length} pages)"
Decide how duplicates should work
The import loop follows the list literally. If your input is [1, 1, 2] (converted to indexes [0, 0, 1]), the output contains page 1 twice followed by page 2. If duplicates are not allowed, reject them explicitly:
if requested_pages.uniq.length != requested_pages.length
raise ArgumentError, "duplicate page numbers are not allowed"
end
Accept ranges without losing order
For a user string such as 1,3-5,9, parse each token into one-based numbers, preserve token order, and then convert to indexes. Reject descending or malformed ranges rather than silently changing the request.
def parse_pages(spec)
spec.split(",").flat_map do |token|
case token
when /A(d+)z/
[$1.to_i]
when /A(d+)-(d+)z/
first_page = Regexp.last_match(1).to_i
last_page = Regexp.last_match(2).to_i
raise ArgumentError, "descending range: #{token}" if last_page < first_page
(first_page..last_page).to_a
else
raise ArgumentError, "invalid page token: #{token}"
end
end
end
requested_pages = parse_pages("1,3-5,9")
Apply the same bounds check after parsing. Consider a maximum selection length if requests come from untrusted clients.
Rank #2
Preservation limits you should test
The simple import approach preserves page contents, but it may not correctly carry every document-level structure. Pay particular attention when the source includes:
- named destinations and bookmarks (outlines);
- internal or external links whose destinations depend on removed pages;
- interactive AcroForm fields and their appearance streams;
- file attachments and embedded files;
- optional content layers;
- encryption, permissions, or password protection;
- metadata, document-level JavaScript, and other catalog entries.
Open the generated file in the viewers your users actually use and inspect these features. When they matter, use HexaPDF’s more advanced import or CLI capabilities rather than assuming a page-only copy is a complete document migration. A successful write only proves that a PDF file was produced, not that every behavior survived.
Free tools Windows power users keep installed
One-click scans. No signup required.
Encrypted sources
Protected PDFs may require a password or may prohibit operations under their permissions. Decide whether your application accepts a password, rejects encrypted input, or delegates decryption to an approved workflow. Never log passwords, and return a controlled error instead of exposing a library backtrace to an end user.
Output integrity checks
Write to a temporary path, close or finish the operation, then rename it into place. Check that the file exists and is non-empty, and optionally open the result again with HexaPDF in a validation step. For high-value documents, render or parse each output page in a second PDF consumer as part of quality assurance.
HexaPDF command-line extraction
If Ruby does not need to manipulate the pages in memory, the HexaPDF CLI can perform a merge with a --pages selection:
Rank #3
hexapdf merge input.pdf --pages 1,3,5 selected.pdf
The CLI manual defines 1-e as the default all-pages range and supports page selection per input. Page-specification grammar can vary with the installed release, so run hexapdf help merge on the deployment machine and verify the exact syntax for ranges and multiple inputs.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Calling the CLI from Ruby safely
require "open3"
input = "input.pdf"
output = "selected.pdf"
stdout, stderr, status = Open3.capture3(
"hexapdf", "merge", input, "--pages", "1,3,5", output
)
unless status.success?
warn stderr
raise "hexapdf merge failed (exit #{status.exitstatus})"
end
puts stdout unless stdout.empty?
Pass each argument separately, as above. Do not interpolate a user-provided filename into a shell command. Check that the executable is installed, capture standard error, apply a timeout in your job runner, and handle non-zero exit status.
PDFtk as an external alternative
PDFtk’s cat operation uses one-based page references and preserves the order in which references appear:
pdftk A=input.pdf cat A1 A3 A5 output selected.pdf
From Ruby, treat PDFtk as an external dependency: verify it is available on every target platform, pass an argument array without shell interpolation, capture its exit status, and define behavior for encrypted inputs. Packaging and licensing considerations belong in your deployment review.
CombinePDF as another Ruby library
CombinePDF exposes a pages collection and can assemble selected entries:
Rank #4
require "combine_pdf"
pdf = CombinePDF.load("input.pdf")
out = CombinePDF.new
[0, 2, 4].each do |index|
out << pdf.pages[index]
end
out.save("selected.pdf")
This demonstrates page access and saving, but verify the current gem’s behavior for the features your files use. Do not assume that forms, annotations, outlines, attachments, encryption, or metadata receive the same treatment as ordinary page content.
Choosing an implementation
| Approach | Indexing | Deployment | Best fit | Main caution |
|---|---|---|---|---|
| HexaPDF Ruby API | Ruby zero-based indexes in code | One gem | Application-controlled selection, validation, and ordering | Document-level structures need testing or advanced handling |
| HexaPDF CLI | CLI page specification, commonly one-based | Executable plus runtime | Simple jobs and existing command pipelines | Confirm installed page-range grammar and process failures |
| PDFtk | One-based references such as A1 |
External executable | Mature command-line workflows | Platform availability, encryption, and safe argument passing |
| CombinePDF | Ruby zero-based collection indexes | One gem | Small Ruby scripts and experiments | Confirm preservation guarantees for your PDF features |
Troubleshooting common failures
“undefined method” or missing constant
Install the gem and require it in the same environment that runs the script. With Bundler, execute through bundle exec ruby your_script.rb so the locked dependency is used.
Wrong pages or an off-by-one result
Ruby arrays start at zero; human page labels start at one. Convert exactly once. Keep external requests one-based, validate against source.pages.count, and subtract one immediately before indexing.
Index out of range
The requested page does not exist. Report the valid range, such as 1..12, and reject the request before importing. Do not let a nil page flow into the target document.
Do these 3 things before closing this tab:
1Repair Windows errors before they cause bigger problems2Scan for outdated or missing drivers - takes under a minute3Clear out junk files and repair common Windows errorsOutput opens but bookmarks or forms are missing
That is a document-level preservation issue, not necessarily a page-content failure. Reproduce with a representative file, inspect HexaPDF’s advanced import options, and verify the result in your target viewer. If the requirement cannot be met, communicate the limitation rather than silently shipping a degraded file.
Best Value
CLI works locally but not in production
Check the executable path, installed version, operating-system package, working-directory permissions, and service-account environment. Log the exit code and sanitized stderr, and enforce a process timeout.
Large files or slow jobs
Selection still requires opening and parsing the source. Run large exports in a background job, limit concurrent workers, use temporary storage with enough space, and delete temporary files after success or failure. Measure memory and elapsed time with your actual PDFs; file size alone does not predict complexity.
Or skip the browser setup
If your real goal is to obtain a clean visual capture of selected web content rather than manipulate an existing PDF, ScreenshotNeo provides a single-request screenshot API. It accepts cookie and consent banners as a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each cleanup step can be disabled. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing status. Its MCP server exposes take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients.
See the ScreenshotNeo documentation for all options. A direct call looks like this:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python:
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
Every feature is available on every plan. The Free plan includes 1,000 screenshots per month without a card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.
Operational checklist
- Define whether callers use one-based page numbers or Ruby indexes.
- Validate every page before importing and set a maximum request size.
- Preserve the requested order deliberately; do not sort unless that is the product requirement.
- Choose a policy for duplicates, encrypted files, and empty selections.
- Test outlines, links, forms, attachments, layers, metadata, and permissions when they matter.
- Write atomically, verify output integrity, and clean up temporary files.
- For external CLIs, pass argument arrays, capture stderr, check exit status, and enforce timeouts.
Frequently Asked Questions
Can I export pages in a different order?
Yes. The import loop follows the selection array, so arrange that array in the exact order required for the new PDF.
Should page numbers in an API request be zero-based?
Human-facing APIs are usually clearer with one-based numbers. Convert to zero-based indexes internally and validate the original numbers against the source page count.
The Tool Desk
Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Does selecting pages automatically preserve bookmarks and form fields?
No guarantee should be assumed. Test document-level features and use advanced import handling when those structures are required.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

