To download a file initiated by a page in Puppeteer, configure Chrome to allow downloads to a known writable directory before clicking the page control. Listen for the download lifecycle, wait for it to finish, then verify the resulting file before closing the browser. The exact Puppeteer integration depends on the Puppeteer and browser versions you use, so check the matching API documentation.
Choose the right download workflow
There are two different situations that are often described as “downloading a file with Puppeteer.” Choose the one that matches how the site delivers the file:
| Situation | Use this approach | What to verify |
|---|---|---|
| A click, form submission, or other page action initiates a browser download. | Configure browser download behavior and a destination directory, trigger the action, and wait for browser download events. | The download reached its completed state and the expected file is present and usable. |
| You already know the file URL and can retrieve it directly from Node.js. | Make an HTTP request from Node rather than asking the browser to download it, if the request does not depend on browser state or a page action. | The response succeeded and the saved content is complete; preserve required authentication and request details. |
A direct request can be simpler, but it is not interchangeable with a page-triggered download when the site relies on a click, browser cookies, or other authenticated browser state. The browser-managed method below is for the first case. Puppeteer’s Page API provides createCDPSession() to attach a Chrome DevTools Protocol (CDP) session for protocol-level operations: Puppeteer Page API.
Install Puppeteer and confirm the browser is available
The puppeteer package downloads a compatible Chrome during installation. puppeteer-core installs the library without downloading a browser, so you must provide an appropriate browser installation and configuration yourself. If a package manager blocked install scripts, the Puppeteer documentation gives npx puppeteer browsers install as a manual browser-installation command. See the Puppeteer installation guide for setup details.
Quick wins for a faster PC:
Clear out junk files and repair common Windows errorsFree Scan →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →#1 Best Overall
Install the full package with:
npm i puppeteer
If you use puppeteer-core, follow its browser setup instructions and ensure your script points to a compatible browser. Puppeteer can control Chrome or Firefox and runs headless by default, but the CDP download workflow below is specifically based on Chrome’s DevTools Protocol. Do not assume identical protocol support in Firefox or across all browser modes.
Configure Chrome, then trigger and verify the download
The Chrome DevTools Protocol Browser domain documents Browser.setDownloadBehavior to set file-download behavior. The current protocol reference lists deny, allow, allowAndName, and default. For allow and allowAndName, a downloadPath is required. The protocol also documents download lifecycle events and browser-context selection; check the protocol reference matching your Chrome version because the linked reference is the current tot version: Chrome DevTools Protocol Browser domain.
- Create a known destination. Choose a directory your process can write to and pass its absolute path. The example creates a dedicated directory under the operating system’s temporary directory.
- Attach a CDP session and configure downloads. Use the page’s
createCDPSession()method and set behavior toallowwith the destination path. - Listen before clicking. Register for
Browser.downloadWillBeginandBrowser.downloadProgressbefore activating the control; fast downloads can otherwise begin before your listener exists. - Wait for a terminal event. Resolve when the matching download reports
completed; reject oncanceledor after a deadline. - Check the result. Confirm that the file exists and is nonempty, then perform any format-specific validation your application needs. Close the browser only after the download has completed.
The example below uses allow, so the file retains its downloaded filename in the destination directory. It assumes the page is already at the relevant site and that #download is the site-specific selector for the control. Replace both the URL and selector with the ones for your page.
const fs = require('node:fs/promises');
const os = require('node:os');
const path = require('node:path');
const puppeteer = require('puppeteer');
async function downloadFromPage() {
const downloadPath = await fs.mkdtemp(path.join(os.tmpdir(), 'puppeteer-download-'));
const browser = await puppeteer.launch({ headless: true });
let page;
try {
page = await browser.newPage();
await page.goto('https://example.com/files', { waitUntil: 'domcontentloaded' });
const client = await page.createCDPSession();
await client.send('Browser.setDownloadBehavior', {
behavior: 'allow',
downloadPath,
eventsEnabled: true,
});
const timeoutMs = 60_000;
const downloadFinished = new Promise((resolve, reject) => {
let guid;
const timer = setTimeout(() => {
cleanup();
reject(new Error(`Download did not finish within ${timeoutMs} ms`));
}, timeoutMs);
function cleanup() {
clearTimeout(timer);
client.off('Browser.downloadWillBegin', onBegin);
client.off('Browser.downloadProgress', onProgress);
}
function onBegin(event) {
guid = event.guid;
}
function onProgress(event) {
if (event.guid !== guid) return;
if (event.state === 'completed') {
cleanup();
resolve(event);
} else if (event.state === 'canceled') {
cleanup();
reject(new Error('The browser canceled the download'));
}
}
client.on('Browser.downloadWillBegin', onBegin);
client.on('Browser.downloadProgress', onProgress);
});
await page.click('#download');
const completed = await downloadFinished;
const names = await fs.readdir(downloadPath);
if (names.length === 0) {
throw new Error('Download reported completion, but the directory is empty');
}
const files = await Promise.all(names.map(async name => {
const fullPath = path.join(downloadPath, name);
const stat = await fs.stat(fullPath);
return { name, fullPath, size: stat.size };
}));
const nonemptyFiles = files.filter(file => file.size > 0);
if (nonemptyFiles.length === 0) {
throw new Error('No nonempty downloaded file was found');
}
console.log({ downloadPath, completed, files: nonemptyFiles });
return { downloadPath, files: nonemptyFiles };
} finally {
await browser.close();
}
}
downloadFromPage().catch(error => {
console.error(error);
process.exitCode = 1;
});
This example is an implementation pattern, not a universal tested script. In particular, the CDP command, event exposure, and event payloads should be checked against the installed Puppeteer and Chrome versions. If your version does not support this integration as written, consult its documentation and use the matching supported API rather than assuming a convenience method exists.
Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minutePC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Handle multiple downloads deliberately
The example tracks one download, which suits a single click expected to produce one file. If one action can start multiple downloads, track each guid from Browser.downloadWillBegin and resolve only after the expected set reaches a terminal state. Match events to the intended action or URL where possible; otherwise an unrelated download in the same browser context could be mistaken for the result. Avoid relying on directory contents alone when the directory may contain files from earlier runs.
Use unique directories and clean up safely
A unique per-run directory prevents stale files from looking like successful output and reduces collisions between concurrent jobs. Retain the directory when downstream code needs the file; remove it only after the file has been consumed or moved. In a long-running service, define a retention policy so temporary downloads do not accumulate indefinitely.
Rank #3
Diagnose PDF navigation versus a file download
A PDF link can either return an attachment that Chrome downloads or navigate to a document that the browser displays in its PDF viewer. A viewer navigation is not necessarily a browser-triggered attachment download, so enabling download behavior alone may not save the PDF. Determine whether the server response is an attachment or whether the browser is displaying a document before choosing the workflow.
Puppeteer’s Page API documents that headless shell does not support navigation to a PDF document. If a script uses page.goto() to open a PDF and fails in headless shell, that limitation may be relevant; it is distinct from clicking a link that initiates a download. The same API notes that goto() may not throw for valid HTTP error statuses in headless shell, so inspect the response status where applicable: Puppeteer Page API.
When direct retrieval is a better fit
If you have a stable file URL and can reproduce the necessary request outside the browser, a Node.js HTTP request can avoid browser download configuration. This route is conditional, not universally preferable: if access depends on a browser login or session, you must preserve the required cookies, headers, or other authentication. Also verify the HTTP status and saved file rather than treating a completed request as proof of a valid document.
Use the browser-managed route when the page action matters or you need to exercise the actual browser session. Use direct retrieval when the URL and request are sufficient and you can correctly reproduce authentication. The available documentation establishes the browser controls and events, but the right choice depends on the site’s behavior and access requirements.
Troubleshoot common failures
- The browser will not launch. Check whether you installed
puppeteerorpuppeteer-core. The former downloads a compatible Chrome during installation; the latter does not. If dependency install scripts were blocked, install a browser withnpx puppeteer browsers installand verify your browser configuration. - The click completes but no file appears. Confirm that download behavior was configured before the click, that the destination path is absolute and writable, and that the selector activated the intended control. Check whether the page navigated to a viewer instead of initiating an attachment download.
- The script hangs waiting for completion. Add a deadline, as in the example. Check whether a download-start event arrived, whether the browser context is the one being monitored, and whether the action opened a new page or initiated a different request. Handle cancellation as a terminal failure rather than waiting forever.
- The event reports completion but the expected filename is missing. Do not assume the suggested filename or event metadata maps to the final path identically in every protocol/browser version. Inspect the destination directory and validate the file; for stricter workflows, correlate the start event, its
guid, and the completion event. - A PDF navigation fails in headless shell. The Page API documents the PDF-navigation limitation for headless shell. Distinguish direct navigation to a PDF from clicking a link that causes an attachment download, and select an approach compatible with the actual response and browser mode.
- An HTTP error did not reject
goto(). In headless shell, valid HTTP error statuses may not causegoto()to throw. Check the returned response status where available and stop before treating the page action as a successful setup. - The file exists but is incomplete or invalid. Wait for the protocol’s completed state, not merely the appearance of a filename. Then check file size and, when necessary, parse or otherwise validate the expected format before passing it to downstream code.
- The download is canceled. Treat the protocol’s
canceledstate as a failure, preserve enough context to diagnose it, and retry only when the cause is understood and retrying is safe for the site.
Performance, reliability, and cost considerations
Browser automation carries the overhead of launching and controlling a browser, but it is often necessary when the download depends on real page behavior or session state. Reusing a browser for multiple tasks may reduce repeated startup work, while separate contexts or per-task directories help isolate cookies and files; isolation requirements depend on the application. Do not close the browser while a download is still active.
For reliable automation, use explicit deadlines, unique writable directories, event-driven completion, cancellation handling, and post-download validation. Do not use an arbitrary sleep as the only completion test: network time and file size vary, and a visible filename can appear before a download is finished. These are engineering safeguards; actual timing and resource use depend on the site, file, host, and browser configuration.
Free tools Windows power users keep installed
One-click scans. No signup required.
Or skip the browser setup
If your goal is a clean visual capture of a page rather than saving a file the site downloads, ScreenshotNeo provides a screenshot API. A single GET request returns a PNG, JPEG, WebP, or PDF capture; it is not a replacement for retrieving an arbitrary attachment from a page.
For example, use cURL to capture a page as WebP:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
See the ScreenshotNeo API documentation for request options. Cookie banners, popups, and chat widgets are removed before the shot; bot checks, blank pages, and failed loads are never billed. An MCP server lets AI agents take screenshots, and 1,000 screenshots a month are free with no card; paid plans start at $5 for 3,000. Sign up for free ScreenshotNeo screenshots.
Sources and version scope
The protocol details here follow the Chrome DevTools Protocol tot Browser reference accessed September 29, 2026. Puppeteer and browser APIs change over time, so verify command and event support against the versions you actually deploy. The installation, Page API, and protocol references linked above are the documentation basis for the setup and limitations described here.
Frequently Asked Questions
Does Puppeteer download a browser when I install it?
The full puppeteer package downloads compatible Chrome during installation; puppeteer-core does not download a browser.
Recommended Free Tools
Can Puppeteer download a PDF that opens in Chrome’s viewer?
It depends on whether the link triggers an attachment download or navigates to a PDF viewer. Those are different browser behaviors, and headless shell has a documented PDF-navigation limitation.
Can I use this CDP workflow in Firefox?
The workflow uses Chrome DevTools Protocol commands and events. Do not assume it works unchanged in Firefox; check the API and browser support for your installed versions.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




