Driver FixRecommendedSound, Wi-Fi or graphics acting up? Check drivers firstFind missing or outdated drivers fast.Check DriversOctober DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsPC HealthRecommendedCrashes, freezes, slowdowns? Check your PC nowSpot repairable issues before they interrupt work.Check PC×
Skip to content
browser automation

How to Fetch a PDF with Puppeteer and Upload It Directly to Google Drive

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

There are two different jobs hidden in this request: retrieving a PDF that a site already serves, or printing the currently rendered page into a new PDF. Puppeteer can navigate, wait for the request that contains an existing PDF, and generate PDFs with page.pdf(); it is not a general programmatic browser-download API. After you have validated PDF bytes, Google Drive API v3 can upload them with a simple media, multipart, or resumable request.

This guide shows both paths in Node.js, keeps the PDF in memory, explains authentication and stream compatibility, and includes recovery steps for redirects, signed URLs, permissions and failed loads.

Choose the PDF workflow first

What you need Use Result
A PDF file the site already serves Observe or trigger the PDF request, then retrieve its response bytes with an HTTP client The original PDF file
A PDF representation of rendered HTML Navigate with Puppeteer and call page.pdf() A newly generated PDF

Puppeteer’s official Files guide states: “Currently, Puppeteer does not offer a way to handle file downloads in a programmatic way.” That limitation concerns browser download handling. It does not prevent you from reading a response body or generating a PDF. Treat these as separate operations rather than expecting a download event to become a Drive upload automatically.

Prerequisites and Drive authentication

  • Node.js with Puppeteer and the Google APIs client installed: npm install puppeteer googleapis.
  • A Google Cloud project with the Google Drive API enabled.
  • Credentials appropriate to your deployment (for example, OAuth user credentials or a service account) and a scope that permits creating files.
  • A destination understood by your credentials. A service account, for example, has its own Drive identity; it will not automatically write into a person’s My Drive unless the folder is shared with that identity or you use an appropriate delegated setup.

The Drive client is initialized with google.drive({version: 'v3', auth}). The exact credential flow, scopes, ownership and shared-drive behavior depend on your application, so keep those settings in environment variables or your deployment secret store.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
import { google } from 'googleapis';

const auth = new google.auth.GoogleAuth({
  keyFile: process.env.GOOGLE_APPLICATION_CREDENTIALS,
  scopes: ['https://www.googleapis.com/auth/drive.file']
});
const drive = google.drive({ version: 'v3', auth });

Branch A: retrieve an existing PDF

Use this branch when the page eventually requests a real PDF. Navigation may involve redirects, cookies, a signed URL or an authorization header. Observe the response, verify its status and content type, then fetch the final URL with the request context required by that site. The example below watches responses while loading a page and retrieves the first successful PDF response.

import puppeteer from 'puppeteer';
import { google } from 'googleapis';

const pageUrl = process.env.PAGE_URL;
if (!pageUrl) throw new Error('Set PAGE_URL');

const auth = new google.auth.GoogleAuth({
  keyFile: process.env.GOOGLE_APPLICATION_CREDENTIALS,
  scopes: ['https://www.googleapis.com/auth/drive.file']
});
const drive = google.drive({ version: 'v3', auth });

const browser = await puppeteer.launch({ headless: true });
try {
  const page = await browser.newPage();
  let pdfResponse;
  page.on('response', response => {
    const type = (response.headers()['content-type'] || '').toLowerCase();
    if (!pdfResponse && response.ok() && (type.includes('application/pdf') || response.url().toLowerCase().endsWith('.pdf'))) {
      pdfResponse = response;
    }
  });

  await page.goto(pageUrl, { waitUntil: 'networkidle2', timeout: 60000 });
  // If a click starts the request, perform it here instead of relying only on page.goto().
  // await page.click('#download-pdf');

  const deadline = Date.now() + 30000;
  while (!pdfResponse && Date.now() < deadline) {
    await new Promise(resolve => setTimeout(resolve, 250));
  }
  if (!pdfResponse) throw new Error('No successful PDF response was observed');

  const sourceUrl = pdfResponse.url();
  const response = await fetch(sourceUrl, {
    redirect: 'follow'
    // Add headers or cookies here when the source requires them.
  });
  if (!response.ok) throw new Error(`PDF fetch failed: ${response.status} ${response.statusText}`);
  const contentType = (response.headers.get('content-type') || '').toLowerCase();
  const bytes = Buffer.from(await response.arrayBuffer());
  if (!contentType.includes('application/pdf') && !bytes.subarray(0, 5).equals(Buffer.from('%PDF-'))) {
    throw new Error('The response was not a PDF');
  }

  const result = await drive.files.create({
    requestBody: { name: 'fetched-document.pdf', mimeType: 'application/pdf' },
    media: { mimeType: 'application/pdf', body: bytes },
    fields: 'id,name,webViewLink'
  });
  console.log(result.data);
} finally {
  await browser.close();
}

In production, the second fetch() must reproduce whatever makes the URL accessible: cookies, an authorization header, a user agent, or a short-lived signed URL. A browser response can be successful while a separate unauthenticated request receives an HTML login page. Check both HTTP status and the PDF signature (%PDF-) before sending data to Drive.

When the PDF request is triggered by an action

Install the response listener before clicking or submitting. You can also use Puppeteer’s request/response waiting methods to target a known URL pattern. Keep the wait bounded; a page may retry, return a bot check, or never request a PDF at all.

Branch B: generate a PDF from the rendered page

Use page.pdf() when the desired document is a printout of the DOM after scripts, fonts and images have loaded. It returns a Uint8Array. Puppeteer uses print CSS media by default; call page.emulateMediaType('screen') first if the screen stylesheet is what you need.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #2
Sale
The Google Workspace Bible: [14 in 1] The Ultimate All-in-One Guide from Beginner to Advanced | Including Gmail, Drive, Docs, Sheets, and Every Other App from the Suite
  • The Google Workspace Bible: [14 in 1] The Ultimate All in One Guide from Beginner to Advanced Including Gmail, Drive, Docs, Sheets, and Every Other App from the Suite
  • ABIS BOOK
import puppeteer from 'puppeteer';
import { google } from 'googleapis';

const pageUrl = process.env.PAGE_URL;
const folderId = process.env.DRIVE_FOLDER_ID;
if (!pageUrl) throw new Error('Set PAGE_URL');

const auth = new google.auth.GoogleAuth({
  keyFile: process.env.GOOGLE_APPLICATION_CREDENTIALS,
  scopes: ['https://www.googleapis.com/auth/drive.file']
});
const drive = google.drive({ version: 'v3', auth });
const browser = await puppeteer.launch({ headless: true });
try {
  const page = await browser.newPage();
  await page.goto(pageUrl, { waitUntil: 'networkidle2', timeout: 60000 });
  await page.emulateMediaType('screen'); // Remove this line for print media.
  const pdf = await page.pdf({
    format: 'A4',
    printBackground: true,
    preferCSSPageSize: true,
    margin: { top: '16mm', right: '16mm', bottom: '16mm', left: '16mm' }
  });

  const metadata = {
    name: 'rendered-page.pdf',
    mimeType: 'application/pdf',
    ...(folderId ? { parents: [folderId] } : {})
  };
  const result = await drive.files.create({
    requestBody: metadata,
    media: { mimeType: 'application/pdf', body: Buffer.from(pdf) },
    fields: 'id,name,webViewLink'
  });
  console.log(result.data);
} finally {
  await browser.close();
}

Wait for the state that matters to your page rather than assuming networkidle2 means every application render is complete. For example, wait for a report selector, a chart container, or a known API result before calling page.pdf(). If the page is authenticated, establish that session before navigation and ensure the resulting content is permitted to be stored.

Streaming with createPDFStream()

page.createPDFStream() returns a Web ReadableStream<Uint8Array>. The Google Node.js client documents media.body as accepting a Node Readable stream. Those interfaces are not automatically interchangeable in every Node.js and client version. Convert or adapt the Web Stream explicitly, or use the byte-buffer approach above when the file size is manageable.

const webStream = await page.createPDFStream({ format: 'A4', printBackground: true });

// Node.js versions with Readable.fromWeb can adapt a Web ReadableStream:
import { Readable } from 'node:stream';
const nodeStream = Readable.fromWeb(webStream);

await drive.files.create({
  requestBody: { name: 'streamed-page.pdf', mimeType: 'application/pdf' },
  media: { mimeType: 'application/pdf', body: nodeStream },
  fields: 'id,name'
});

Confirm that your installed Node.js version and Google client accept this adapter. Do not promise zero-memory operation merely because a method is called “stream”; buffering can still occur in the browser, adapter or HTTP layer.

Choose the Drive upload type

Upload Use when Metadata
Simple media A small content-only upload where creation metadata is unimportant Sent separately or omitted
Multipart You want the filename, MIME type or folder metadata and content in one request Included with media
Resumable An interrupted-transfer recovery path or large-transfer handling matters Included while the transfer proceeds in sessions

The files.create examples use multipart-style metadata plus media because a useful name and MIME type normally matter. Select resumable upload when your reliability requirements justify session management; do not rely on an unverified file-size threshold.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Validation, permissions and cleanup

  • Reject non-2xx source responses and HTML masquerading as a PDF.
  • Use a deterministic filename and, when needed, a shared Drive folder ID in requestBody.parents.
  • Return the created file ID and link only after Drive confirms creation.
  • Always close the browser in a finally block. Set navigation and response deadlines so a stalled site does not leave Chromium processes running.
  • Log request status and Drive error details without logging cookies, authorization headers or private document contents.

Troubleshooting

No PDF response is observed

The site may generate the document client-side, use a blob URL, require a click, or return a viewer page. Attach the listener before the action, wait for the relevant selector, and inspect response URLs and content types. If no remote PDF exists, switch to page.pdf().

The fetched URL returns HTML or a 401/403

Carry the browser’s required cookies, authorization and user-agent context, or use the original response body when appropriate. Signed URLs can expire between observation and refetch; retrieve promptly and follow redirects.

The PDF is blank or missing images

Wait for the application’s loaded state, fonts and image selectors. Check that lazy content was rendered before printing, and choose print or screen media deliberately. A successful HTTP request does not guarantee that client-side content has finished drawing.

Drive returns 401, 403 or “File not found”

Verify the credential, scope, target folder sharing and shared-drive settings. The identity creating the file must have permission to write to the selected location.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Stream type errors appear

createPDFStream() produces a Web Stream, while the Node client expects a Node Readable. Use Readable.fromWeb() where supported, otherwise adapt with a compatible bridge or upload a Buffer.

The process hangs

Add finite timeouts to navigation and waits, handle rejected promises, and close the browser in finally. Avoid waiting forever for a selector that may never appear because of a bot check or an application error.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup

If your actual goal is a clean visual capture rather than an original remote PDF or a print document, ScreenshotNeo provides a one-call screenshot API. It removes cookie banners, newsletter popups and chat widgets before capture; bot checks, blank pages and failed loads are not billed; and its MCP server lets AI agents take screenshots.

For a screenshot request, see the parameter options in the ScreenshotNeo documentation:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

The free plan includes 1,000 screenshots per month with no card, and paid plans start at $5 for 3,000 screenshots. This is a screenshot workflow, not a replacement for retrieving a source PDF or generating a PDF with Puppeteer. Sign up for the free ScreenshotNeo plan.

Best Value
Google Drive Reference and Cheat Sheet: The unofficial cheat sheet reference for Google Drive
  • hole punched
  • high quality card stock
  • 4 pages
  • made in USA
  • keyboard shortcuts

Frequently Asked Questions

Can Puppeteer upload a browser download directly to Drive?

Not through a built-in programmatic download handler. Observe or retrieve the PDF response yourself, or generate a PDF with page.pdf(), then call Drive files.create.

Should I use page.pdf() for a PDF link?

No. If the server already provides a PDF, fetch and validate that resource. Use page.pdf() when you want a new PDF of the rendered page.

Which Drive upload mode is required?

Simple media suits small content-only transfers, multipart combines metadata and content, and resumable is intended for interruption recovery or large-transfer handling.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Read next

Recommended PC Tool
Recommended PC Tool
PC Slower Than It Used to Be?Free scan - under a minute
Outdated Drivers Are Slowing You DownFree scan - exact matches

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.