The Tool Desk
Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →To archive an entire website, first define what you need to preserve and how it must replay. Use the Wayback Machine for a quick, public reference; run ArchiveBox when you need a controlled local copy in several formats; or use Archive-It for a managed institutional collection. For preservation, retain the crawler’s WARC files, metadata and checksums—not only a PDF or screenshot—and test representative pages offline. No archive guarantees that login flows, forms or server-side features will continue to work.
How do I archive an entire website? Make three decisions first
1. State the purpose
- Public citation: You need a stable link that others can open, usually for a page or small set of pages.
- Private backup: You control the files and can keep them offline or behind authentication.
- Legal or research evidence: You need a defensible record of what was available, when it was captured and what failed.
- Long-term institutional collection: You need scheduled crawling, collection administration, access controls and export procedures.
2. Define the scope
Write down the seed URLs, allowed domains and paths, maximum crawl depth, exclusions, file types and recrawl schedule. Decide whether subdomains, CDN assets, downloadable files and linked third-party services belong in the collection. A single article, a path such as /docs/, an entire domain and a group of related domains are different projects; do not assume that capturing the home page captures the rest.
3. Decide what must replay
Static HTML, images and stylesheets are comparatively straightforward. JavaScript applications, forms, login-gated pages, personalized responses and server-side searches may depend on the original host. An archived copy can preserve the visible response without preserving the function that produced it. List the interactions that matter and plan a separate test for each.
Method 1: Use the Wayback Machine for a quick public snapshot
The Internet Archive’s Wayback Machine is the simplest choice when you need a shareable historical reference or want to inspect earlier versions of a page. Its “Save Pages in the Wayback Machine” and “Archive whole web sites” guidance explains how to submit pages and broader crawls.
#1 Best Overall
- Easily store and access 2TB to content on the go with the Seagate Portable Drive, a USB external hard drive
- Designed to work with Windows or Mac computers, this external hard drive makes backup a snap just drag and drop
- To get set up, connect the portable hard drive to a computer for automatic recognition no software required
- This USB drive provides plug and play simplicity with the included 18 inch USB 3.0 cable
- The available storage capacity may vary.
- Choose and record your seed URL, scope and capture date.
- Submit the page or site through the Wayback Machine’s save workflow.
- Open the resulting timestamped copy in a private browser window.
- Check representative internal links, images, stylesheets, downloads and redirects instead of trusting the index page alone.
- Record missing assets, blocked requests and pages that still depend on the live host.
The service collects publicly available pages. It does not collect pages that require passwords or user-entered form submissions. The Internet Archive warns: “When a dynamic page contains forms, JavaScript, or other elements that require interaction with the originating host, the archive will not contain the original site’s functionality.” Treat a Wayback capture as a public reference copy, not as a complete backup of an application.
When a Wayback capture is enough
Use it for a citation, a change-history check, or a small public record where link sharing matters more than owning the underlying files. It is also useful as a first pass before committing to a larger crawl.
When to choose another method
Choose a controlled crawler when the material is sensitive, you need several derivatives, you require repeatable schedules, or you must retain raw capture data and logs under your own control.
Method 2: Build a controlled local copy with ArchiveBox
ArchiveBox is open-source, self-hosted software that preserves content from websites in multiple formats. Depending on the capture, it can retain ordinary HTML, a browser-rendered SingleFile page, PDF, PNG screenshot, DOM output, article text, JSON, headers, media and WARC data. That variety lets you keep a machine-readable source alongside a human-readable rendering.
Do these 3 things before closing this tab:
1Fix the driver behind crashes, sound loss and screen glitches2Repair Windows errors before they cause bigger problems3Scan for outdated or missing drivers - takes under a minuteA practical ArchiveBox workflow
- Install in isolation. Use a dedicated host, container or virtual environment and restrict who can read the archive directory.
- Add seeds and boundaries. Enter the starting URLs, permitted domains and paths, crawl depth, exclusions and recrawl policy. Do not allow an unrestricted crawl to follow every external link.
- Enable browser rendering. For JavaScript-heavy pages, use the browser-based capture so the rendered DOM and lazy content have a chance to load. Expect some application behavior still not to replay.
- Keep complementary outputs. Retain WARC and raw files for preservation, plus a PDF or PNG when a reviewer needs a fixed visual record.
- Record provenance. Store capture time, original URL, scope, software version, configuration and errors with the collection.
- Test offline. Open representative pages without network access; inspect assets, links, redirects, downloads and at least one JavaScript-heavy page.
- Make a second backup. Copy the archive and its metadata to separate storage, and periodically verify that the copy can be read.
Publishing an ArchiveBox collection safely
ArchiveBox documents two publication models: its built-in web server and export as static HTML. A private backup or research copy has different legal implications from public rehosting for profit. For any shared instance, configure authentication, disable public indexing and public submission by default, place HTTPS in front of the service, and establish a process for DMCA or GDPR requests. Do not expose an archive merely because the software can serve it.
Rank #2
- Easily store and access 5TB of content on the go with the Seagate portable drive, a USB external hard Drive
- Designed to work with Windows or Mac computers, this external hard drive makes backup a snap just drag and drop
- To get set up, connect the portable hard drive to a computer for automatic recognition software required
- This USB drive provides plug and play simplicity with the included 18 inch USB 3.0 cable
- The available storage capacity may vary.
Method 3: Use Archive-It for an institutional collection
Archive-It is the Internet Archive’s service for organizations that need to harvest, build and preserve digital-content collections. It is aimed at libraries, universities, agencies and regulated teams that need managed crawling, collection administration and organizational access controls rather than a single operator’s local directory.
Before committing, confirm the provider’s current crawl scope, limits, pricing, export rights and partnership terms. Document who owns the captured material, who can access it, how takedown requests are handled and how the collection will be exported if the service relationship changes.
Why WARC matters for preservation
Web crawlers receive a seed URL and gather HTML, images and related resources into a WARC (Web ARChive) file. WARC is the preservation-oriented capture package: it keeps the fetched responses and associated records in a form that replay tools can interpret. A PDF or screenshot is a useful derivative for a person to read, but it is not a substitute for the underlying resources.
Free tools Windows power users keep installed
One-click scans. No signup required.
- Keep the original WARC files unchanged.
- Preserve capture metadata, configuration and error logs beside them.
- Calculate or retain checksums when your workflow supports them.
- Record the replay software or viewer used, because rendering can vary between tools and versions.
- Keep a readable derivative such as PDF or PNG for quick review, clearly labeled as a derivative.
How can I save a website for offline use?
For a site you control, an ArchiveBox collection is the practical route in this guide: crawl within defined boundaries, enable browser rendering where needed, keep raw and rendered outputs, then test with the network disabled. Offline availability does not imply offline functionality. A form may display but have nowhere to submit; a search box may require a server; a script may call an API that was never captured. Mark those dependencies in the collection record rather than presenting the copy as a working replacement.
Single page versus whole site
A single-page capture has a narrow, auditable scope and is easier to verify. A whole-site crawl requires seed management, exclusions, depth limits, scheduling and storage planning. If readers need only a few evidence pages, do not create an apparently complete domain archive that was never intended or tested as one.
Rank #3
- Easily store and access 1TB to content on the go with the Seagate Portable Drive, a USB external hard drive.Specific uses: Personal
- Designed to work with Windows or Mac computers, this external hard drive makes backup a snap just drag and drop. Reformatting may be required for Mac
- To get set up, connect the portable hard drive to a computer for automatic recognition no software required
- This USB drive provides plug and play simplicity with the included 18 inch USB 3.0 cable
- The available storage capacity may vary.
Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server for developers. It is useful when you need a clean visual record of selected pages rather than a crawlable preservation package. Before capture, it accepts the cookie or consent banner as a visitor and removes 60+ known consent platforms, newsletter popups and chat widgets; each cleanup step can be disabled. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads and cache hits cost nothing, and the response identifies the result with X-Page-Verdict and X-Billed headers.
One GET request returns PNG, JPEG, WebP or PDF. The API supports full-page capture with lazy images loaded, a CSS-selector element capture, dark mode, 12 device presets or any viewport, retina scale, PDF paper size, margins, landscape mode and page ranges. You can supply custom CSS or JavaScript, click an element before capture, hide selectors, wait for a selector, delay or network idle, and block ads, trackers, requests or resource types. Custom headers, cookies, user agents, Authorization, timezone and geolocation handle pages that vary by request. Other options include transparent backgrounds, resizing, a chosen cache TTL, signed links for public <img> tags, asynchronous jobs with signed webhooks, bulk capture of up to 100 URLs per call, a usage API and an OpenAPI specification. Parameter names used by other screenshot APIs also work, which can simplify migration.
Use the API for visual evidence or a contact sheet of important pages; keep WARC and raw responses when you need preservation, replay or discovery. The MCP server exposes take_screenshot, get_page_info and capture_pdf to Claude, Cursor and other MCP clients, so an AI agent can collect selected views without your maintaining a browser runner.
cURL
See the ScreenshotNeo documentation for request options.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp
Python
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://example.com"}, timeout=90)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
Node.js
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://example.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`HTTP ${res.status}`);
const fs = await import('node:fs/promises');
await fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer()));
Every feature is included on every plan: Free provides 1,000 shots each month with no card; Starter is $5 for 3,000; Growth $15 for 15,000; Pro $39 for 60,000; Scale $99 for 250,000; and Business $249 for 1,000,000. Yearly billing gives two months free. Create a free ScreenshotNeo account to start with the 1,000-shot allowance and no card.
Verification checklist: how to know what you actually captured
- Open representative pages with the network disabled.
- Inspect images, fonts, stylesheets, scripts, downloads, canonical links and redirects.
- Test at least one JavaScript-heavy page and document any missing interaction.
- Confirm every record has its original URL and capture timestamp.
- Compare the declared crawl scope with the URLs actually fetched.
- Preserve checksums where available and verify them after copying storage.
- Record blocked requests, timeouts, robots exclusions and pages that still load from the live site.
- Check that access controls work from an unauthenticated browser before sharing.
An index page loading successfully is not proof that the collection is complete.
Rank #4
- Easily store and access 4TB of content on the go with the Seagate Portable Drive, a USB external hard drive.Specific uses: Personal
- Designed to work with Windows or Mac computers, this external hard drive makes backup a snap just drag and drop
- To get set up, connect the portable hard drive to a computer for automatic recognition no software required
- This USB drive provides plug and play simplicity with the included 18 inch USB 3.0 cable
- The available storage capacity may vary.
Troubleshooting common archive failures
The archive shows a blank page
Often the capture timed out, a script failed, or the page depended on an API response that was not saved. Retry with browser rendering, increase the wait condition, capture after network idle, and inspect logs. Keep the failed attempt and its timestamp rather than silently replacing it.
Images or styles are missing
Check whether assets came from another host, were lazy-loaded, blocked by policy or rewritten through a CDN. Add permitted asset domains, enable lazy-content loading where available, and verify the result offline. Do not broaden the crawl to every external domain without an explicit boundary.
Links return to the live website
Inspect redirects, canonical tags and absolute URLs. A replay tool may intentionally leave an uncaptured destination external. Record those links as unresolved and either capture the destination within scope or label it as live-dependent.
Login pages or forms do not work
Password-protected pages and user-entered form submissions are not collected by the Wayback Machine, and server-side behavior generally cannot be recreated from a static response. For a private project, document the access state and preserve permitted responses; never place credentials or personal submissions in a public archive.
Recommended Free Tools
The collection is unexpectedly huge
Stop the crawl, review seeds and exclusions, cap depth, and block tracking parameters or calendar paths that generate endless URLs. Resume only after the boundary is explicit and logged.
Best Value
- [Upgraded Version] - This external hard drive features a mirrored logo stripe combined with a striped anti-slip design, and the rounded corners of the casing make it easier to grip. The stripes also have a heat dissipation function, ensuring stable and fast data transfer.
- 【Ultra-thin and quiet】 - The motherboard adopts JMicron 578 noise-free solution, giving you a quiet working environment. Lightweight and portable size designed to fit in your pocket for easy portability.
- 【Ultra-Fast Data Transfers】 - Pairing this external hard drive with JMicron 578 solution USB 3.0 and USB 2.0 interfaces enables blazing-fast data transfer. It boasts theoretical read speeds of up to 125MB/s and write speeds of up to 103MB/s.
- 【Plug and Play】 - With no software to install, just plug it in and the drive is ready to use.The hard disk chip is wrapped with an aluminum anti-interference layer to increase heat dissipation and protect data.
- 【What You Get】 - 1 x Portable Hard Drive, 1 x USB 3.0 Cable, 1 x User Manual, Gift-type shell packaging ,Three-year manufacturer's warranty and free technical support services.
A public archive receives a takedown request
Restrict access while reviewing the request, preserve the relevant metadata, and follow the process appropriate to your jurisdiction and hosting arrangement. Public rehosting can trigger copyright, privacy, terms-of-service, DMCA or GDPR obligations that do not apply in the same way to a private research copy.
Choosing among the methods
| Method | Best fit | Scope and replay | Ownership and outputs | Main caution |
|---|---|---|---|---|
| Wayback Machine | Quick, shareable public reference | Page submissions and broader public crawls; dynamic functions may not replay | Public archive service; timestamped replay | Password pages and form submissions are not collected; verify assets |
| ArchiveBox | Controlled local backup or research collection | Defined seeds and boundaries; browser rendering plus multiple derivatives | Self-hosted files including HTML, SingleFile, PDF, PNG, DOM, text, JSON, headers, media and WARC | You administer security, storage, replay testing and legal access |
| Archive-It | Managed organizational or institutional program | Provider-managed harvesting and collection administration | Organizational access controls and service-managed collection | Confirm current limits, pricing, exports and terms directly |
| ScreenshotNeo | Clean visual evidence for selected URLs | Single or bulk screenshots and PDFs; not a substitute for WARC preservation | PNG, JPEG, WebP or PDF via API; MCP tools for AI clients | Use a crawler for discovery, raw capture and replay requirements |
Legal and ethical handling
Copyright, privacy, terms of service and takedown rules vary by country and by use. Keep private captures access-controlled, minimize personal data, honor applicable removal requests and obtain permission before publishing material you do not own. If an archive contains account details, health information, customer records or submitted form data, treat it as sensitive even when the source page was publicly reachable. Write the intended audience and retention period into the collection plan before crawling.
Frequently Asked Questions
What is a seed URL?
It is a starting address supplied to the crawler. The crawler follows links from that address only within the domains, paths and depth rules you define, so documenting seeds makes the collection reproducible.
How often should an archived site be recrawled?
Set the interval from the site’s change rate and the importance of missed updates: frequent changes and high-value records justify shorter intervals, while stable reference pages need less frequent captures. Record the policy with the collection so gaps are explainable.
Can I combine methods?
Yes. A common pattern is a Wayback snapshot for a public citation, ArchiveBox for owner-controlled WARC and derivatives, and ScreenshotNeo for clean visual records of selected pages. Keep each capture’s timestamp and scope separate so one output is not mistaken for another.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

