Driver FixRecommendedSound, Wi-Fi or graphics acting up? Check drivers firstFind missing or outdated drivers fast.Check DriversOctober DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsSlow PC?RecommendedPC slow today? Run a repair scan before it gets worseResolve common Windows issues and optimize system performance.Scan Now×
Skip to content
browser automation

How to Crawl a Website and Capture Screenshots of Every Page

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

To capture a defensible record of a website, first define what “every page” means, build a URL queue from the sitemap and internal links, normalize and deduplicate it, then visit each permitted URL with a browser such as Playwright. Save a full-page screenshot, final URL, status, timestamp and error details in a manifest. The result is complete for the scope and discovery sources you document—not proof that no undiscovered URL exists.

Define “every page” before you crawl

A website rarely has a finite, universally agreed list of pages. Your report should state the exact boundary you are testing:

  • Starting host: for example, example.com. Decide whether www.example.com and other subdomains are included.
  • URL rules: include or exclude paths, file extensions, languages, query parameters and fragments. Fragments (the part after #) normally identify a position in one document rather than a separate page.
  • Authentication: say whether private, account or paywalled areas are excluded. Do not bypass login controls.
  • Resource limits: set maximum URLs, depth, elapsed time and request rate so an accidental calendar or search loop cannot run forever.
  • Capture policy: choose viewport screenshots or full scrollable pages, image format, viewport size, color scheme and whether redirects count under the requested or final URL.

Use the phrase “all discovered, in-scope URLs” in the final report. A sitemap and a successful browser run cannot establish that a site has no other URLs.

Build the URL inventory

Start with the sitemap

Fetch /robots.txt and look for a sitemap declaration, then check common sitemap locations such as /sitemap.xml. A sitemap index can point to several child sitemaps. Parse each XML file and add its URL entries to your queue. Google describes sitemaps as a discovery aid: “A sitemap helps search engines discover URLs on your site, but it doesn’t guarantee that all the items in your sitemap will be crawled and indexed.” Treat that list as input, not a completeness guarantee.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
#1 Best Overall
Sale
Samsung T7 Portable SSD 1TB Titan Gray, USB 3.2 Gen 2, Up to 1,050MB/s
  • MADE FOR THE MAKERS: Create; Explore; Store; The T7 Portable SSD delivers fast speeds and durable features to back up any endeavor; Build your video editing empire, file your photographs or back up your blogs all in an instant
  • SHARE IDEAS IN A FLASH: Don’t waste a second waiting and spend more time doing; The T7 is embedded with PCIe NVMe technology that brings fast read and write speeds up to 1,050/1,000 MB/s¹, making it almost twice as fast as the T5
  • ALWAYS MAKE THE SAVE: Compact design with massive capacity; With capacities up to 4TB, save exactly what you need to your drive – from large working files to game data and everything in between
  • ADAPTS TO EVERY NEED: Whether using a PC or mobile phone, count on the T7 for extensive compatibility²; It’s a true team player when it comes to heavy-duty application usage or file-saving
  • HI RESOLUTION VIDEO RECORDING: Record Ultra High Resolution (4K 60fs) videos directly onto the T7 Portable SSD with your favorite camera or mobile devices; Supports iPhone 15 Pro Res 4K at 60fps video and more³

Follow internal links

Visit known pages and extract anchor links. Resolve relative links against the page URL, retain meaningful query parameters, and keep only links that satisfy your host and path policy. JavaScript-rendered navigation may appear only after the page runs, so link discovery should happen in the browser after an appropriate wait.

Use both sources

A sitemap-only crawl misses pages omitted from the file; a link-only crawl misses orphaned pages that no page links to. A combined queue gives you a more explainable inventory and lets the manifest record whether each URL came from a sitemap, a link, or both. It still does not guarantee complete coverage.

Normalize, deduplicate and queue safely

Normalize URLs before inserting them into a set:

  • Lowercase the scheme and hostname.
  • Remove default ports and fragments.
  • Resolve dot segments such as /a/../b.
  • Apply one deliberate trailing-slash policy; do not silently treat both forms as identical if the server distinguishes them.
  • Preserve query parameters that change content (for example, language or product identifiers). Exclude tracking parameters only when you have explicitly decided they do not represent separate pages.
  • Store both the requested URL and the final URL after redirects.

Reject obvious non-HTML downloads unless they are intentionally in scope. Add loop guards for unbounded query combinations, calendar dates, faceted filters and session URLs. Keep a reason for every exclusion so another person can audit the boundary.

Respect access guidance and server capacity

Read robots.txt and honor applicable crawl restrictions for the crawler you operate. Robots.txt communicates access preferences and can help manage traffic; it is not a privacy or security boundary and does not reliably keep a URL out of search results. Google’s documentation describes Google’s crawlers, so apply those details to a private capture as guidance rather than a universal legal rule.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #2
Sandisk 2TB Extreme Portable SSD, Up to 1050MB/s, USB-C, USB 3.2 Gen 2, IP65 Water and Dust Resistance, Updated Firmware, External Solid State Drive, SDSSDE61-2T00-G25
  • Get NVMe solid state performance with up to 1050MB/s read and 1000MB/s write speeds in a portable, high-capacity drive(1) (Based on internal testing; performance may be lower depending on host device & other factors. 1MB=1,000,000 bytes.)
  • Up to 3-meter drop protection and IP65 water and dust resistance mean this tough drive can take a beating(3) (Previously rated for 2-meter drop protection and IP55 rating. Now qualified for the higher, stated specs.)
  • Use the handy carabiner loop to secure it to your belt loop or backpack for extra peace of mind.
  • Help keep private content private with the included password protection featuring 256‐bit AES hardware encryption.(3)
  • Easily manage files and automatically free up space with the SanDisk Memory Zone app.(5). Non-Operating Temperature -20°C to 85°C

Use a modest concurrency (often one to a few pages at a time), a delay between requests when the server is busy, and a bounded retry policy. Stop or slow down on repeated 429, 500 or 503 responses. Never attempt to defeat authentication, bot checks or other access controls.

Run a Playwright crawl and capture

Install and prepare

Use a current Node.js installation, create a project, install Playwright and download its browser:

mkdir site-capture && cd site-capture
npm init -y
npm install playwright
npx playwright install chromium

The example below reads newline-separated seed URLs from seeds.txt, discovers links while visiting pages, captures each URL once, and writes a JSON manifest. It is intentionally conservative: one browser context, two concurrent workers, a 30-second navigation timeout and two retries.

Complete crawler

const { chromium } = require('playwright');
const fs = require('fs/promises');
const path = require('path');

const START_HOST = 'example.com';             // change this
const seeds = (await fs.readFile('seeds.txt', 'utf8'))
  .split(/r?n/).map(x => x.trim()).filter(Boolean);
const outDir = 'shots';
await fs.mkdir(outDir, { recursive: true });

function normalize(raw, base) {
  const u = new URL(raw, base);
  u.hash = '';
  u.protocol = u.protocol.toLowerCase();
  u.hostname = u.hostname.toLowerCase();
  if ((u.protocol === 'https:' && u.port === '443') ||
      (u.protocol === 'http:' && u.port === '80')) u.port = '';
  return u.href;
}
function inScope(u) {
  const x = new URL(u);
  return (x.hostname === START_HOST || x.hostname.endsWith('.' + START_HOST)) &&
         (x.protocol === 'https:' || x.protocol === 'http:');
}
function fileName(u) {
  return Buffer.from(u).toString('base64url').slice(0, 180) + '.png';
}

const queue = [...new Set(seeds.map(x => normalize(x, x)).filter(inScope))];
const seen = new Set();
const manifest = [];
const browser = await chromium.launch();
const context = await browser.newContext({ viewport: { width: 1440, height: 900 } });

async function capture(requested) {
  const page = await context.newPage();
  const started = new Date().toISOString();
  let record = { requestedUrl: requested, startedAt: started };
  for (let attempt = 1; attempt <= 3; attempt++) {
    try {
      const response = await page.goto(requested, { waitUntil: 'networkidle', timeout: 30000 });
      const finalUrl = page.url();
      const status = response ? response.status() : null;
      const name = fileName(requested);
      await page.screenshot({ path: path.join(outDir, name), fullPage: true });
      const links = await page.locator('a[href]').evaluateAll(as => as.map(a => a.href));
      for (const link of links) {
        try { const n = normalize(link, finalUrl); if (inScope(n) && !seen.has(n)) queue.push(n); } catch {}
      }
      record = { ...record, finalUrl, status, file: name, result: 'captured' };
      break;
    } catch (err) {
      record = { ...record, result: attempt === 3 ? 'failed' : 'retrying', error: String(err) };
      if (attempt < 3) await new Promise(r => setTimeout(r, attempt * 1500));
    }
  }
  manifest.push(record);
  await page.close();
}

const workers = Array.from({ length: 2 }, async () => {
  while (queue.length) {
    const u = queue.shift();
    if (!u || seen.has(u)) continue;
    seen.add(u);
    await capture(u);
  }
});
await Promise.all(workers);
await fs.writeFile('manifest.json', JSON.stringify({ scope: START_HOST, manifest }, null, 2));
await browser.close();

Run it with node crawl.js. Replace START_HOST, put one or more permitted URLs in seeds.txt, and inspect shots/ plus manifest.json. The fullPage: true option captures the complete scrollable document instead of only the current viewport; it does not discover additional URLs.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #3
SSK Portable SSD 250GB External Solid State Hard Drive USB C Up to 1050MB/s
  • Capacity Display Variance: 250GB external ssd often appears as around 232GB on Windows. MacOS can show full 250 GB capacity. This is binary calculation difference and doesn’t affect SSD hard drive actual physical storage
  • 1050 MB/s Speed: Instantly access to your files with blazing-fast 10Gbps external SSD read up to 1050MB/s and write up to 1000MB/s. LED Light indicates USB SSD instant activity
  • Data Security: Solid state drives S.M.A.R.T. health diagnostics​ and adaptive TRIM optimizing data block management ensures consistent write speeds and extends the longevity of the portable SSD
  • USB-C & USB-A Cable: Both cables featuring rapid USB 3.2 Gen2, this USB SSD effortlessly bridges devices, enabling seamless cross-platform file transfers and backup between computers, smartphones, tablets and iPhone
  • Always Fast: No slowdowns for large file transfers. With SLC caching (25% of current available capacity allocated as high-speed cache), this external SSD delivers steady 10Gbps for transfers within the cache capacity

When to change the script

  • For pages that load content after network idle, wait for a known selector or add a bounded delay.
  • For login-protected work you own, create a context with authorized storage state rather than embedding passwords in source.
  • For responsive review, run separate queues with different viewport sizes; a single image cannot represent every device.
  • For very long pages, consider element captures or PDF output because full-page PNG dimensions can become unwieldy.

Make coverage auditable

For every attempted URL, retain:

  • requested and final URL;
  • discovery source (sitemap, link, seed or combination);
  • timestamp, HTTP status and screenshot filename;
  • retry count, error text and final outcome (captured, skipped or failed);
  • the applied scope, exclusions, viewport and browser version.

Report totals for queued, attempted, captured, failed and skipped URLs. List failures rather than silently dropping them. A “100% captured” statement should mean 100% of the URLs in your documented queue, not every URL that might exist.

Choose a discovery and capture strategy

Approach Strength Limitation Best use
Sitemap only Fast declared inventory May be stale or incomplete; JavaScript-only routes are absent Sites with a carefully maintained sitemap
Internal links only Finds pages exposed through navigation and rendered links Misses orphaned pages and blocked states Exploring a public navigation graph
Combined sitemap + links Most explainable starting coverage Still cannot prove completeness Audited inventories and migrations

Likewise, choose viewport capture when you need a consistent screen state, and full-page capture when reviewers need all scrollable content. Full-page images are taller and may require more storage and slower review.

Or skip the browser setup

ScreenshotNeo is the quickest hosted option for a one-off or automated capture: it removes cookie/consent banners, newsletter popups and chat widgets before capture, and only clean shots are billed. Bot checks, blank pages, timeouts, failed loads and cache hits cost nothing, with the outcome shown in X-Page-Verdict and X-Billed headers. Its MCP server lets Claude, Cursor and other MCP clients use take_screenshot, get_page_info and capture_pdf.

One GET request returns PNG, JPEG, WebP or PDF. The API supports full-page lazy-image loading, CSS-selector element capture, dark mode, 12 device presets plus custom viewports, retina scale, PDF paper/margins/landscape/page ranges, custom CSS and JavaScript, clicks, selector or network-idle waits, ad/tracker/request blocking, headers/cookies/user agents/Authorization, timezone and geolocation, transparent backgrounds, resizing, chosen-TTL caching, signed image links, asynchronous jobs with signed webhooks, bulk capture of up to 100 URLs per call, a usage API and an OpenAPI specification. Common parameter names used by other screenshot APIs are accepted to ease migration.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

cURL

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)

Node.js

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`${res.status} ${await res.text()}`);
require('fs').writeFileSync('shot.webp', Buffer.from(await res.arrayBuffer()));

See the ScreenshotNeo documentation for option names and response headers. Plans include 1,000 shots per month free with no card, Starter $5 for 3,000, Growth $15 for 15,000, Pro $39 for 60,000, Scale $99 for 250,000 and Business $249 for 1,000,000; yearly billing provides two months free, and every feature is on every plan. Create a free ScreenshotNeo account to start with 1,000 screenshots a month and no card.

Rank #4
Sale
Seagate 2TB Portable Hard Drive | USB 3.0 (STGX2000400)
  • Easily store and access 2TB to content on the go with the Seagate Portable Drive, a USB external hard drive
  • Designed to work with Windows or Mac computers, this external hard drive makes backup a snap just drag and drop
  • To get set up, connect the portable hard drive to a computer for automatic recognition no software required
  • This USB drive provides plug and play simplicity with the included 18 inch USB 3.0 cable
  • The available storage capacity may vary.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Troubleshoot common failures

Queue is unexpectedly small

Check that the sitemap URL was fetched successfully, XML namespaces were handled, and your host policy is not excluding www or a required subdomain. Confirm that link extraction runs after client-side navigation renders.

Many 403 or 429 responses

Reduce concurrency, add delay, identify your crawler honestly, and follow the site’s access guidance. Do not rotate identities to evade controls. If you own the site, allow the capture host explicitly or use an authorized authenticated context.

Blank or partially rendered screenshots

Replace an overly broad networkidle wait with a page-specific selector, then wait briefly for lazy content. Capture after scrolling if the site loads images only when they approach the viewport. Record the condition so the run can be reproduced.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Navigation timeouts

Retry a small, fixed number of times, then mark the URL failed. Check DNS, TLS, redirects and third-party resources. Do not classify a timeout as a successful capture.

Best Value
Sale
Sandisk 1TB Extreme Portable SSD, Up to 2000MB/s Transfer Speeds-New Model
  • NEARLY 2X FASTER THAN OUR PREVIOUS GENERATION(8) – move 1,000 high-res photos in under 60 seconds(6) with up to 2000MB/s transfer speeds(2).
  • IP65 RATING AND UP TO 3M DROP PROTECTION(3) – protects against spills and drops.
  • POCKET-SIZED – fits easily in pockets and small bags.
  • SPACE TO OWN YOUR AI CONTENT – speed and capacity to download your high-res clips and photo edits.
  • 256-BIT AES ENCRYPTION(4) – helps keep private files secure with password protection.

Duplicate files or missing redirects

Use a normalized requested URL as the stable filename key, while storing the final URL separately. This preserves evidence that multiple requests converged on one destination.

Huge images or memory pressure

Use a smaller viewport, capture a key element, split the job, or emit PDFs. Keep workers bounded and close each page after capture.

Operational checklist

  1. Write the host, subdomains, paths, authentication and query policy.
  2. Read robots.txt and identify sitemap files.
  3. Parse sitemaps and crawl rendered internal links into one normalized queue.
  4. Set concurrency, timeout, retry and maximum-URL limits.
  5. Capture with a fixed viewport and fullPage: true when the whole document is required.
  6. Save a manifest with requested/final URLs, statuses, timestamps, filenames and errors.
  7. Review failures and report exclusions; never imply coverage beyond the documented queue.

Frequently Asked Questions

Does a sitemap contain every page on a site?

No. It is a discovery aid and can be stale, incomplete or unindexed; combine it with rendered internal-link discovery and document the resulting scope.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Can I crawl pages behind a login?

Only when you are authorized. Supply an approved session to your browser context and state clearly that authenticated URLs are included; never bypass access controls.

What does Playwright’s fullPage option change?

It expands the screenshot to the page’s full scrollable document. It does not discover links or other URLs.

How should I prove that a crawl was complete?

You cannot prove that no undiscovered URL exists. You can demonstrate complete processing of a stated queue by retaining discovery inputs, normalization rules, exclusions, attempts and failures.

Quick Recap

SaleBestseller No. 4
Seagate 2TB Portable Hard Drive | USB 3.0 (STGX2000400)
Seagate 2TB Portable Hard Drive | USB 3.0 (STGX2000400)
This USB drive provides plug and play simplicity with the included 18 inch USB 3.0 cable; The available storage capacity may vary.
$119.99
SaleBestseller No. 5
Sandisk 1TB Extreme Portable SSD, Up to 2000MB/s Transfer Speeds-New Model
Sandisk 1TB Extreme Portable SSD, Up to 2000MB/s Transfer Speeds-New Model
IP65 RATING AND UP TO 3M DROP PROTECTION(3) – protects against spills and drops.; POCKET-SIZED – fits easily in pockets and small bags.
$253.00

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Read next

Recommended PC Tool
Recommended PC Tool
PC Slower Than It Used to Be?Free scan - under a minute
Crashes, No Sound, or Screen Glitches?Free driver scan

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.