Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
Configure a website screenshot archive by defining a bounded set of URLs, choosing a recurrence that matches how quickly the site changes, selecting the right screenshot mode, monitoring each crawl, and preserving the result as a portable WACZ. Use viewport images for the first screen, fullPage for the complete rendered page, and thumbnail for compact visual indexing. Repair missed pages or interactions with an interactive capture, then verify that the exported archive replays offline.
1. Define what the archive must prove
A screenshot archive is useful only when its capture boundary matches its purpose. Write a one-sentence preservation objective before opening a crawler, such as “preserve the public pricing pages each month” or “record the landing page as visitors saw it on launch day.” That sentence determines the seeds, URL scope, recurrence, screenshot mode and review effort.
Choose the capture boundary
- Single page: Use one explicit seed URL when the evidence is limited to a page or a time-sensitive state.
- Bounded section: Seed a directory and restrict traversal to the required paths. Exclude search, calendar, tag, sort and other parameter patterns that can generate effectively unlimited URLs.
- Changing site: Define the exact pages to revisit and schedule separate recurring crawls when different sections change at different rates.
Start with explicit seed URLs rather than a broad domain crawl. Treat query strings, redirects, subdomains and authenticated areas as separate scope decisions. A crawler that is allowed to follow every generated URL can spend its budget on filters, session links or infinite calendars instead of the pages you intend to preserve.
Record context before capture
Give the collection a descriptive name and record its owner, purpose, seed list, exclusions, recurrence, login profile (if any), screenshot modes and quality notes. Context distinguishes a successful archive from a technically complete crawl that captured the wrong edition, locale or logged-out state.
#1 Best Overall
- Easily store and access 2TB to content on the go with the Seagate Portable Drive, a USB external hard drive
- Designed to work with Windows or Mac computers, this external hard drive makes backup a snap just drag and drop
- To get set up, connect the portable hard drive to a computer for automatic recognition no software required
- This USB drive provides plug and play simplicity with the included 18 inch USB 3.0 cable
- The available storage capacity may vary.
2. Set a recurrence that reflects change
A one-time preservation job and a time series have different configurations. Browsertrix documents recurring schedules of daily, weekly or monthly. There is no universal “correct” interval: select the shortest interval that answers your question without producing unnecessary duplicates or review work.
| Site behavior | Practical starting schedule | What to review |
|---|---|---|
| Launch, campaign or frequently edited news pages | Daily, then shorten temporarily during the event | URL expansion, changed hero assets and consent states |
| Product, documentation or marketing site | Weekly | Navigation changes, redirects and newly generated paths |
| Stable reference content | Monthly or one-time | Whether the interval still catches meaningful revisions |
These intervals are workflow choices, not retention guarantees. Keep the schedule, timezone and run identifier with each export so a later reviewer can distinguish a missed run from an unchanged page.
3. Choose the screenshot mode deliberately
Browsertrix Crawler documents three screenshot modes through its --screenshot option. The documented viewport for view and thumbnail is 1920×1080.
| Mode | Output | Use it when | Important limitation |
|---|---|---|---|
view |
PNG of the initially visible viewport | The evidence is what appears on the first screen | Below-the-fold content is not represented |
fullPage |
PNG of the full rendered page | You need one image of the complete page | Long pages can be large and still do not prove every interaction |
thumbnail |
JPEG thumbnail of the initially visible viewport | You need a compact visual index | It is a viewport thumbnail, not a full-page record |
Modes can be combined. For a review archive, combining view and fullPage gives a quick first-screen comparison plus a complete-page image. For a large collection, add thumbnail only when its smaller files materially improve browsing. Browsertrix writes screenshots to screenshots.warc.gz; when WACZ generation is enabled, that WARC is included and indexed with the other archive records.
Full-page capture command
Enable the documented full-page mode in your Browsertrix Crawler invocation:
Rank #2
- Easily store and access 5TB of content on the go with the Seagate portable drive, a USB external hard Drive
- Designed to work with Windows or Mac computers, this external hard drive makes backup a snap just drag and drop
- To get set up, connect the portable hard drive to a computer for automatic recognition software required
- This USB drive provides plug and play simplicity with the included 18 inch USB 3.0 cable
- The available storage capacity may vary.
browsertrix-crawler crawl --screenshot fullPage
Add your normal seed, scope and output settings for the version you have installed. Because command-line options and plan limits can change, confirm the current Browsertrix Crawler documentation before copying a production command. If you need both the initial viewport and the complete page, configure both documented modes rather than assuming a full-page image contains a separate viewport record.
4. Capture authenticated or interactive content responsibly
Real-browser crawling is appropriate for JavaScript-heavy pages and content that requires a login profile. Browsertrix documents browser-based capture and login profiles. Use an account that is authorized to access the material, minimize the permissions stored in the profile, and follow the site’s access rules. A login profile does not guarantee that every protected page, challenge or state can be reproduced.
Free tools Windows power users keep installed
One-click scans. No signup required.
Plan interactions explicitly
- Identify the click, scroll, tab, form submission or consent choice that reveals the state you need.
- Capture the page after the interaction, not merely the landing URL.
- Note whether the state depends on a session, locale, viewport or feature flag.
- Keep credentials and private data out of collection descriptions and exported notes.
A screenshot is visual evidence, not a complete substitute for the archived network resources or interaction history. Preserve the crawl records and metadata alongside the image so another person can replay and inspect the underlying capture.
5. Monitor each crawl and control URL expansion
Do not treat a scheduled crawl as set-and-forget. Watch live progress, sampled URLs and failure messages. Browsertrix documents live monitoring and URL exclusions, which let you stop runaway patterns and correct scope while a crawl is still useful.
Signals that the scope is wrong
- The URL count rises rapidly through changing query parameters.
- Many captures are login, search or calendar pages that were not in the objective.
- The same content appears under many tracking or session URLs.
- Expected paths never appear because a redirect or exclusion is too broad.
Pause or stop an unsafe run, add exclusions for the offending patterns, and rerun from a clean collection when necessary. Keep the incomplete run as a quality note only if your retention policy requires an audit trail; do not present it as a complete archive.
Rank #3
- Easily store and access 1TB to content on the go with the Seagate Portable Drive, a USB external hard drive.Specific uses: Personal
- Designed to work with Windows or Mac computers, this external hard drive makes backup a snap just drag and drop. Reformatting may be required for Mac
- To get set up, connect the portable hard drive to a computer for automatic recognition no software required
- This USB drive provides plug and play simplicity with the included 18 inch USB 3.0 cable
- The available storage capacity may vary.
6. Inspect, repair and combine captures
Review representative pages from every run: a landing page, a deep page, a media-heavy page, a redirected URL and any authenticated or interactive target. Check that the screenshot mode produced the expected dimensions and that lazy-loaded images are present before accepting the run.
Patch a missed page or interaction
- Open the target interactively in ArchiveWeb.page.
- Perform the required navigation, click or scroll and capture the resulting state.
- Export the interactive capture in WARC or WACZ.
- Add the archived item to the Browsertrix collection as a patch.
- Export the combined collection as one WACZ and record which item was added manually.
ArchiveWeb.page states that captured data remains local unless shared and that sessions can be exported in WARC and WACZ. Use that local workflow for sensitive material, and share only what your authorization permits. A patch supplements a crawl; it does not erase the fact that the automated run missed an interaction.
7. Export and verify a portable archive
Export the completed collection as WACZ. Browsertrix documents downloadable WACZ exports, and its archived-item documentation describes offline playback in ReplayWeb.page. Before distributing an archive, open the export in a compatible viewer and test several URLs, screenshots and metadata records without network access.
- Confirm that the WACZ opens and the index lists the expected records.
- Replay a viewport capture and a full-page capture.
- Check one patched item and its quality note.
- Compare the exported seed list and crawl date with the collection metadata.
- Store an independent copy according to your organization’s retention and backup policy; the cited documentation does not prescribe a universal retention period or backup count.
8. Keep a quality record with every run
For each scheduled run, retain the collection name, UTC start and end time, crawler version, seed URLs, exclusions, recurrence, screenshot modes, login-profile identifier (not the secret), record count, known failures and reviewer. Mark whether the run is complete, partial or patched. This turns a folder of images into an auditable archive.
9. Troubleshooting common failures
Only the first screen appears
Cause: The crawl used view or thumbnail, both documented as initially visible viewport captures. Fix: Enable --screenshot fullPage and rerun the affected URLs. If the page loads content only after scrolling or interaction, capture that state separately.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Rank #4
- Easily store and access 4TB of content on the go with the Seagate Portable Drive, a USB external hard drive.Specific uses: Personal
- Designed to work with Windows or Mac computers, this external hard drive makes backup a snap just drag and drop
- To get set up, connect the portable hard drive to a computer for automatic recognition no software required
- This USB drive provides plug and play simplicity with the included 18 inch USB 3.0 cable
- The available storage capacity may vary.
The crawl keeps discovering URLs
Cause: Facets, tracking parameters, session links or calendars are expanding the scope. Fix: Add URL exclusions, narrow the seed paths and monitor a new run before enabling its recurrence.
Images or embedded resources are missing
Cause: Lazy loading, blocked requests, timing or an interaction was not reproduced. Fix: Inspect the replay, capture the state interactively with ArchiveWeb.page, and patch the Browsertrix collection. Record the limitation rather than silently treating the page as complete.
A login page replaces the protected content
Cause: The login profile expired, lacks permission or does not reproduce the required session state. Fix: Re-authorize the profile, verify access manually, and recapture only the permitted pages. Do not put credentials in URLs or collection metadata.
The WACZ cannot be replayed
Cause: An incomplete export, incompatible viewer or damaged transfer. Fix: Re-export from the source collection, verify the file before sharing, and test it in a compatible ReplayWeb.page environment. Keep the original export and checksum according to your local policy.
Do these 3 things before closing this tab:
1Repair Windows errors before they cause bigger problems2Fix the driver behind crashes, sound loss and screen glitches3Clear out junk files and repair common Windows errorsOr skip the browser setup
For a single clean image, ScreenshotNeo provides a website screenshot API and MCP server for developers. One GET request returns PNG, JPEG or WebP (or a PDF), while its browser accepts consent banners before capture and removes more than 60 known consent platforms, newsletter popups and chat widgets. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads and cache hits are not billed; response headers identify the page verdict and whether the request was billed. The API also supports full-page capture, lazy-image loading, CSS-selector element shots, dark mode, device presets, custom viewports, retina scale, PDF page ranges, custom CSS and JavaScript, clicks, waits, blocking rules, headers, cookies, user agents, authorization, timezone, geolocation, transparent backgrounds, resizing, chosen cache TTLs, signed links, asynchronous jobs with signed webhooks, bulk capture of up to 100 URLs per call, usage reporting and an OpenAPI specification. Parameter names used by other screenshot APIs also work.
cURL
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
Node.js
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
See the ScreenshotNeo documentation for request options and response headers. An MCP server exposes take_screenshot, get_page_info and capture_pdf to Claude, Cursor and other MCP clients, so an AI agent can capture and inspect pages without your building browser orchestration.
Best Value
- [Upgraded Version] - This external hard drive features a mirrored logo stripe combined with a striped anti-slip design, and the rounded corners of the casing make it easier to grip. The stripes also have a heat dissipation function, ensuring stable and fast data transfer.
- 【Ultra-thin and quiet】 - The motherboard adopts JMicron 578 noise-free solution, giving you a quiet working environment. Lightweight and portable size designed to fit in your pocket for easy portability.
- 【Ultra-Fast Data Transfers】 - Pairing this external hard drive with JMicron 578 solution USB 3.0 and USB 2.0 interfaces enables blazing-fast data transfer. It boasts theoretical read speeds of up to 125MB/s and write speeds of up to 103MB/s.
- 【Plug and Play】 - With no software to install, just plug it in and the drive is ready to use.The hard disk chip is wrapped with an aluminum anti-interference layer to increase heat dissipation and protect data.
- 【What You Get】 - 1 x Portable Hard Drive, 1 x USB 3.0 Cable, 1 x User Manual, Gift-type shell packaging ,Three-year manufacturer's warranty and free technical support services.
| Plan | Included shots | Price |
|---|---|---|
| Free | 1,000/month | $0, no card |
| Starter | 3,000 | $5 |
| Growth | 15,000 | $15 |
| Pro | 60,000 | $39 |
| Scale | 250,000 | $99 |
| Business | 1,000,000 | $249 |
Yearly billing provides two months free, and every feature is available on every plan. Start with ScreenshotNeo’s free sign-up to get 1,000 screenshots a month without a card.
10. A practical operating checklist
- Write the preservation objective and define seeds.
- Exclude generated URL patterns before the first scheduled run.
- Choose
view,fullPage,thumbnailor a documented combination. - Set daily, weekly or monthly recurrence to match observed change.
- Use authorized browser profiles for protected content.
- Monitor URL growth and failures during every crawl.
- Review representative replays and patch misses interactively.
- Export WACZ and verify offline playback.
- Store metadata, quality notes and an independent backup.
Frequently Asked Questions
Can a screenshot archive preserve every website interaction?
No. A screenshot records a rendered state. Interactive captures and the underlying WARC/WACZ records provide more context, but no screenshot alone proves every possible state or resource.
The Tool Desk
Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Should I keep both screenshots and WACZ files?
Yes. Screenshots are convenient visual evidence, while WACZ preserves the indexed archive records needed for portable replay and later inspection.
Is monthly capture always enough?
No. Use an interval based on the target’s observed change rate and the purpose of the archive; Browsertrix documents daily, weekly and monthly schedules.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

