Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use Puppeteer to collect each anchor’s browser-resolved URL, filter and deduplicate those URLs, then navigate to them one at a time and save a screenshot. A reliable batch script also chooses a page-readiness signal, gives every file a safe unique name, records failures per URL, and closes the browser even if the job stops unexpectedly.

How the link-to-screenshot workflow works

There are four separate jobs: open the page that contains the links, extract URLs, decide which URLs belong in the batch, and visit and capture each selected page. Keeping those steps distinct makes it easier to control scope and recover from individual failures.

  1. Navigate to the source page and wait until its links are present.
  2. Read a[href] anchors in the page and use each anchor’s href property, which resolves relative links against the document URL.
  3. Normalize, deduplicate, and filter the URLs before navigating to them.
  4. For each URL, navigate with an appropriate readiness strategy and save a viewport or full-page image under a filesystem-safe name.

The example below processes pages sequentially using one Puppeteer page. That is a sensible default for a batch: it limits simultaneous browser work and makes the output order predictable.

Runnable Puppeteer example

This example uses JavaScript ES modules and Node.js. Install Puppeteer in a project first with npm install puppeteer; use a Node.js version supported by the Puppeteer release you install. Save the code as capture-links.mjs and run it with node capture-links.mjs.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
import puppeteer from 'puppeteer';
import {mkdir} from 'node:fs/promises';

const startUrl = 'https://example.com';
const outDir = './screenshots';

const browser = await puppeteer.launch();
try {
  await mkdir(outDir, {recursive: true});
  const page = await browser.newPage();
  await page.goto(startUrl, {waitUntil: 'domcontentloaded'});

  const links = await page.$$eval('a[href]', anchors =>
    anchors.map(anchor => anchor.href)
  );

  const urls = [...new Set(links)]
    .filter(url => /^https?:$/.test(new URL(url).protocol));

  for (const [index, url] of urls.entries()) {
    try {
      await page.goto(url, {waitUntil: 'networkidle2', timeout: 30_000});
      const fileName = `${String(index + 1).padStart(4, '0')}.png`;
      await page.screenshot({path: `${outDir}/${fileName}`, fullPage: true});
      console.log(`Saved ${url} -> ${fileName}`);
    } catch (error) {
      console.error(`Skipped ${url}:`, error.message);
    }
  }
} finally {
  await browser.close();
}

The first navigation waits only for domcontentloaded, because the immediate task is to discover anchors rather than capture the source page. Each destination waits for networkidle2 before capture. The timeout bounds how long one slow destination can hold up the loop; the inner try/catch logs that URL and continues to the next one.

Page.screenshot() can save image data to a path or return image data, depending on the options. In this script, path writes each PNG to disk and fullPage: true requests a full-document capture rather than just the current viewport. Puppeteer’s screenshot options default fullPage to false; the API also has options such as image type and quality. See the official Page.screenshot API, ScreenshotOptions, Page API, and Puppeteer screenshot guide.

Choose which links to capture

Keep absolute HTTP and HTTPS pages

Inside the browser, anchor.href gives an absolute, browser-resolved URL, including when the HTML contains a relative link such as /pricing. The protocol filter keeps ordinary web pages and skips schemes such as mailto:, tel:, and javascript:, which are not pages to navigate to. The new URL(url) check parses each candidate before it is accepted.

Restrict the run to the source origin when appropriate

If the goal is to capture pages belonging to one site, filter by origin. This avoids following links to external websites, which may make a batch unexpectedly large or take it outside the scope of a site audit. Add this after extraction, using the source URL as the boundary:

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
const sourceOrigin = new URL(startUrl).origin;
const urls = [...new Set(links)]
  .filter(url => /^https?:$/.test(new URL(url).protocol))
  .filter(url => new URL(url).origin === sourceOrigin);

Origin includes scheme, hostname, and port. A different subdomain is a different origin, so decide explicitly whether links such as docs.example.com belong in your run.

Remove fragments or apply other normalization

Fragments identify a position or state within a URL and are not sent to the server as part of the HTTP request. If your capture should treat /guide#install and /guide#upgrade as the same destination, remove the fragment before deduplication:

const normalized = links.map(value => {
  const url = new URL(value);
  url.hash = '';
  return url.href;
});
const urls = [...new Set(normalized)]
  .filter(value => /^https?:$/.test(new URL(value).protocol));

Do not strip query parameters indiscriminately: they can select different content or application states. Likewise, URL normalization should reflect what counts as “the same page” for your task rather than silently changing destinations.

Wait for the right kind of readiness

Use DOM content readiness for fast, simple pages

waitUntil: 'domcontentloaded' returns after the initial HTML document has been parsed. It can be a good choice for a static page where that is enough for a useful screenshot. It does not guarantee that images, client-rendered content, or other late-loading elements are ready.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use network idle as a heuristic, not a guarantee

waitUntil: 'networkidle2' is useful for pages that settle after a small amount of network activity. It may be unsuitable for sites with analytics, ads, WebSockets, long polling, or other persistent requests: they can prevent a clean idle point or make the wait unrepresentative of visual readiness. The example’s 30-second navigation timeout prevents an unbounded wait, but a timeout should be handled as a failed capture rather than treated as proof the page is ready.

Wait for a known page element when possible

For an application with a reliable content marker, waiting for that selector can be more meaningful than guessing from network traffic:

await page.goto(url, {waitUntil: 'domcontentloaded', timeout: 30_000});
await page.waitForSelector('main article', {timeout: 15_000});
await page.screenshot({path: filePath, fullPage: true});

Choose a selector that appears only when the content you need is usable. A selector that exists in a hidden template or appears before the page finishes rendering may still produce an incomplete shot. Puppeteer’s Page API includes goto, waitForSelector, waitForNavigation, waitForNetworkIdle, evaluate, and screenshot methods; combine them according to the destination’s behavior rather than applying one readiness rule to every site.

Choose screenshot extent and file names

Viewport or full document

Omit fullPage or set it to false for a screenshot of the current viewport. Set fullPage: true when the image should cover the full document. Long pages can create very tall files, so viewport capture may be easier to inspect or store when only the first screen matters.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use stable, safe names

The example writes 0001.png, 0002.png, and so on. Index-based names avoid illegal filesystem characters and prevent two URLs with awkward or identical-looking paths from overwriting each other. Keep a separate URL-to-filename log or manifest if you need to map images back to their source URLs; a readable slug or a hash can also be added to the index when that traceability matters.

Select an image encoding when needed

The file extension in the example matches Puppeteer’s PNG output. The screenshot API supports a type option and a quality option for lossy formats. Set the encoding deliberately if storage size or compatibility matters, and make the filename extension match the format you request.

Batch reliability, speed, and scope

  • Sequential reuse: One page reused in a loop is straightforward and avoids the extra memory and network load of many pages running at once.
  • Isolation: Open a separate page when URLs need different cookies or isolated page state. Separate pages are also the building block for concurrency, but parallel capture increases resource use and can burden the destination site.
  • Failures: Keep the per-URL error handler so a navigation or screenshot failure does not abort all later work. For a production audit, write failures to a report with the URL and error message instead of relying only on console output.
  • Cleanup: The outer finally closes the browser whether the loop finishes or an unexpected error escapes. Without it, a failed job can leave browser processes running.
  • Scope: A page’s current DOM supplies the anchors. Links injected only after interaction or scrolling may not be present when extraction runs; if the target site reveals links dynamically, first perform the interaction or loading steps needed to make those anchors appear.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Troubleshooting common failures

No screenshots are saved

Check that the starting page loaded, that it contains anchors with href attributes, and that the protocol filter retains their schemes. Log links.length and urls.length after extraction to distinguish an empty source page from a filtering issue. If links appear after client-side rendering, wait for the relevant selector before collecting anchors.

The batch hangs or reports navigation timeouts

A destination may be slow, inaccessible, or never reach the selected network-idle condition. Retain a finite timeout; replace networkidle2 with domcontentloaded plus a known selector when that better represents usable content. The per-URL catch lets the remaining URLs proceed.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The screenshot is blank or incomplete

The chosen readiness condition may occur before the content is rendered, or the page may require interaction. Wait for a content-specific selector, or perform the required click or scroll before capture. A full-page option changes the capture extent; it does not ensure that lazy content has loaded.

External pages appear unexpectedly

Apply the same-origin filter before the loop. If subdomains should be included, define that policy explicitly rather than assuming an origin check includes them.

Files overwrite each other or cannot be written

Keep the index-based naming scheme and ensure the output directory is writable. If you build names from URLs, sanitize them and retain a unique component; raw URL strings can contain characters unsuitable for filenames.

Or skip the browser setup

If you want an API call rather than managing a Puppeteer browser, ScreenshotNeo takes a URL and returns a screenshot or PDF. The example below saves a PNG response for each URL in the same sequential loop. See the ScreenshotNeo documentation for request options and response details.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
import requests

urls = ['https://example.com']
for index, url in enumerate(urls, start=1):
    response = requests.get(
        'https://api.screenshotneo.com/v1/shot',
        params={'access_key': 'YOUR_API_KEY', 'url': url, 'format': 'png'},
        timeout=90,
    )
    response.raise_for_status()
    with open(f'screenshots/{index:04}.png', 'wb') as image:
        image.write(response.content)

ScreenshotNeo removes cookie and consent banners, newsletter popups, and chat widgets before capture; bot checks, blank pages, failed loads, timeouts, and cache hits are not billed. Its MCP server gives AI agents tools for screenshots, page information, and PDF capture. The Free plan includes 1,000 shots per month with no card, and paid plans start at $5 for 3,000 shots.

Create a free ScreenshotNeo account to use the monthly free allowance without a card.

Frequently Asked Questions

Does Puppeteer capture links that open in a new tab?

The extraction step reads anchor URLs from the source page; it does not need to click the anchors or open their target tabs. It then navigates the reused page directly to each URL.

Can I save a PDF instead of an image with Puppeteer?

Yes. Puppeteer has a page PDF capability; use it when the deliverable is a printable document rather than a screenshot image.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.