October DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsWindows FixRecommendedWindows errors stealing your time? Find the fix fastScan stability, cleanup and performance issues.Fix NowOctober DealsAmazon USDeal season is back - check today's better picksAmazon US: current deals, useful picks and tech finds.See Picks×
Skip to content
Browsertrix

How to Capture Multiple Levels of a Website

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

To capture multiple levels of a website, start a crawler at a page, set the link depth and URL scope, impose a page limit, then review the discovered URLs before saving them. A crawler follows links; a screenshot tool captures pages. Choose the output—screenshots, linked offline pages, or a WARC archive—according to what you need to do with the result.

What “multiple levels” means

Link depth is the number of link hops from the seed page. If the homepage is the seed, it is depth 0; a page linked directly from it is depth 1; a page linked from that page is depth 2. Your crawler’s convention for displaying the seed may differ, so check its documentation when setting a depth limit.

Link depth is not the same as folder depth. A URL such as example.com/products/widgets/ has a deeper-looking path than the homepage, but that path does not prove how many clicks away the page is. Crawlers may expose both settings; use link depth to control navigation from your seed and path filters to restrict URL areas. Screaming Frog documents crawl depth, URL limits, and related configuration in its SEO Spider configuration guide. WebsiteArchiver documents link-depth and page-limit controls in its crawler guide.

How to plan a bounded crawl

  1. Choose a seed URL. Start at the homepage if you want the site’s main navigation, or at a section page if you only need that part of the site. Depth is counted from this starting point.
  2. Set the scope. Decide whether to include just the starting host, subdomains, a specific directory, or external domains. A broad scope can pull in unrelated pages, while a narrow one may omit pages linked outside the selected area.
  3. Set link depth and a hard page limit. A depth cap bounds how many link hops are followed; a page cap limits the total number of URLs processed. Where available, add per-depth or path limits. These controls help contain large sites, query-string variations, tag pages, and other repetitive URLs.
  4. Decide whether to use sitemap discovery. An XML sitemap can supplement pages found by following links. Browsertrix supports regular sitemaps and sitemap indexes, while still applying crawl scope and limits; see its common options documentation. Sitemap discovery does not replace scope or page caps.
  5. Choose static downloading or browser rendering. Static downloading does not run JavaScript and may be faster for conventional pages. Use browser rendering when content or links appear only after scripts run, or when the capture depends on browser session state. Confirm the resulting pages, because dynamic behavior and access requirements vary by site.
  6. Review the discovered URLs. Remove out-of-scope pages and excessive variants before the download or capture phase if the tool provides a review step. WebsiteArchiver documents a discovery-and-review workflow with options to untick unwanted pages.
  7. Select the output format. Choose screenshots for visual records, linked offline pages for browsing saved pages, or WARC for web-archive workflows. These outputs are not interchangeable.

Choose an output that matches the job

Output Useful for What it does not replace
Full-page screenshots Visual review, page comparisons, and records of how pages appeared at capture time. An offline site with working page-to-page links or a replayable web archive.
Linked offline pages Reading and navigating a downloaded collection of pages. A faithful visual record of every page state or a WARC archive.
WARC archive Web-archiving and replay-oriented workflows. Browsertrix and Screaming Frog document WARC-related capture options. A simple set of screenshot images or necessarily convenient everyday offline browsing.

Browsertrix’s documentation describes WARC/WACZ-related output and screenshot modes; WebsiteArchiver describes website download workflows; Screaming Frog documents local website archives in hierarchical or WARC format. Check the relevant tool’s current documentation for the exact output options and behavior.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Tools that document multi-page capture

These examples have documented features relevant to crawling or archiving. This is not a hands-on performance ranking; choose based on platform, scope controls, rendering needs, and intended output.

  • Browsertrix Crawler: Its documentation describes a browser-based crawler, sitemap parsing, configurable behavior, initial-viewport, full-page and thumbnail screenshots, and WARC/WACZ-related capture output. Documentation is for versions 1.0.0 and above. Start with the Browsertrix Crawler documentation and its common options.
  • WebsiteArchiver: Its documentation describes a macOS crawler with discovery and review before download, link depth and page caps, static and browser engines, domain controls, and a robots.txt option. Its documentation says the free version is limited to depth 1 and a page count capped by remaining free items; confirm current terms and limits before relying on them. See its crawler documentation.
  • Screaming Frog SEO Spider: Its documentation covers crawl-depth and per-depth URL limits, JavaScript rendering, screenshots, and local website archives in hierarchical or WARC format. Product limits and licensing can change; consult its current configuration guide.
  • WebCapture: The Chrome Web Store listing describes bulk crawling and full-page screenshots, sitemap discovery, URL-pattern filters, page limits, and delays. Those are listing claims, not an independent performance assessment.

How to capture every page you can discover

“Every page” is only meaningful after you define the site boundary and discovery rules. A crawler cannot be assumed to find pages that are not linked within scope or included in a sitemap it reads. Pages can also be omitted by page limits, rendering differences, access requirements, or failures. Use a bounded crawl, then inspect the URL list and the crawl’s failure or progress report rather than treating the configured depth as proof of completeness.

For a large site, begin with a representative section and modest depth and page limits. Review the discovered URL patterns—especially query strings, pagination, tags, and duplicated paths—before raising the cap. If you need broader coverage, expand scope or limits deliberately and rerun, rather than starting with an unbounded crawl.

Static versus browser-based capture

A static downloader retrieves page resources without executing JavaScript. This can be adequate for conventional pages, but it may miss content or links that appear only after client-side scripts run. Browser rendering uses a browser engine and is the relevant choice when script-dependent content, interaction, or session cookies matter. It can also take longer, and the page may still behave differently depending on access and session state.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

For logged-in pages, confirm how the tool handles cookies or browser sessions and capture only the pages available in the session you intend to use. Do not assume that a public crawl will include private or session-dependent material. After capture, open representative results and verify that the expected content and links are present.

Or skip the browser setup

If you already have the URLs to capture, ScreenshotNeo can return a screenshot or PDF with one GET request. It is a website screenshot API and MCP server for developers; it does not replace a crawler that discovers a site’s links. Feed it URLs from your crawl when you need page images. See the ScreenshotNeo website and API documentation.

cURL

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Replace the example URL and supply your API key. ScreenshotNeo can remove cookie/consent banners, newsletter popups, and chat widgets before capture; each of those steps can be turned off. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and responses identify the page verdict and billing status in headers. Its MCP server gives AI agents tools for screenshots, page information, and PDF capture. The free plan includes 1,000 shots per month with no card; paid plans start at $5 for 3,000. Sign up for the free plan.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Troubleshooting a multi-level capture

  • Expected pages are missing: Check whether they are linked from the seed or available in a sitemap the crawler reads. Then review domain and path scope, link-depth limits, URL filters, and the page cap.
  • The crawl stops too soon: Confirm that the setting is link depth rather than folder depth, and check whether a total or per-depth URL limit was reached.
  • Pages load without their content or links: The site may depend on JavaScript. Try browser rendering and inspect the saved result; do not expect a static download to execute scripts.
  • Duplicate-looking URLs consume the limit: Query parameters and repeated URL patterns can create many variants. Add path or URL-pattern limits where supported and review the discovered list before continuing.
  • A page works in your browser but not in the capture: It may depend on login, cookies, or other session state. Check the tool’s session handling and verify that the capture uses the intended access state.
  • Some pages fail or appear incomplete: Use the crawler’s progress and failure reporting, and retry options where available. WebsiteArchiver documents progress/failure reporting and retries; inspect representative results after any retry.
  • The saved result is hard to use: Reconsider the deliverable. Screenshots are images, linked downloads are for offline browsing, and WARC is an archive format; select the format for the actual task.

Performance, reliability, and cost considerations

Crawl size depends on the reachable links, sitemap inputs, scope, depth, URL rules, and page cap; there is no universal page count for a given depth. Browser rendering and session-dependent pages can require more work than static downloads. Set limits first, review the URL set, and use failure reports to distinguish an incomplete capture from a successfully completed one. Product availability, plan limits, and features can change, so verify current vendor documentation before choosing a tool or planning a production workflow.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Be deliberate about crawl behavior: keep the scope relevant, respect site access controls and applicable requirements for your use, and use a robots.txt option where appropriate. The cited product documentation describes such controls, but it does not establish a jurisdiction-wide legal rule for every capture.

Frequently Asked Questions

Can a crawler guarantee that it captures every page on a website?

No. Coverage depends on discoverable links or sitemap entries, configured scope and limits, rendering, access, and page failures.

Can I use screenshots as an offline website copy?

Screenshots preserve page appearance as images, but they do not provide the linked offline browsing or archive replay that other output formats are intended to support.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Read next

Recommended PC Tool
Recommended PC Tool
Outdated Drivers Are Slowing You DownFree scan - exact matches
Windows Errors? Fix Them Before They SpreadFree repair scan

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.