Do these 3 things before closing this tab:
1Scan for outdated or missing drivers - takes under a minute2Clear out junk files and repair common Windows errors3Fix the driver behind crashes, sound loss and screen glitchesUse HTTrack to make a linked, local copy of a public website. Start with the site’s final canonical URL, set hostname and path boundaries before crawling, then inspect the logs and test the copy. HTTrack downloads HTML, CSS, images and other files recursively, but it does not execute JavaScript, bypass authentication or reproduce every server-side state. For one page, Internet Archive’s Save Page Now is simpler; for scheduled organizational preservation, Archive-It is the managed option.
Choose what “a whole website” means
The right method depends on the outcome you need. A local mirror is useful for offline reading and internal testing. A preservation capture may need WARC or WACZ output, a recorded capture date and a replay system. A single-page citation does not justify crawling thousands of URLs.
| Need | Suitable approach | Main trade-off |
|---|---|---|
| Offline browsing of a linked public site | HTTrack local mirror | Coverage depends on discoverable links, scope, server access and page-generation methods. |
| Save one page for citation or sharing | Internet Archive Save Page Now | Saves one page and its resources, not an outlink crawl. |
| Recurring institutional captures | Archive-It | Paid managed subscription; availability and terms should be confirmed with Internet Archive. |
Do not treat any mirror as the original site’s database. Search indexes, forms, accounts, personalized pages, streaming media and interactive applications can depend on services or states that a downloader cannot reproduce.
Prepare the crawl before downloading
Start at the final URL
Open the site in a browser and note where HTTP, HTTPS, the apex domain and www redirects end. Begin HTTrack at that final destination. A same-host crawl can stop or miss content when the starting URL redirects to another hostname.
#1 Best Overall
- Easily store and access 2TB to content on the go with the Seagate Portable Drive, a USB external hard drive
- Designed to work with Windows or Mac computers, this external hard drive makes backup a snap just drag and drop
- To get set up, connect the portable hard drive to a computer for automatic recognition no software required
- This USB drive provides plug and play simplicity with the included 18 inch USB 3.0 cable
- The available storage capacity may vary.
Define host and path boundaries
Decide whether the copy includes only example.com, related subdomains such as docs.example.com, or assets on a permitted CDN. Third-party analytics, video hosts and advertising domains usually do not belong in the mirror. Narrow boundaries reduce unrelated downloads and server load.
Estimate storage and risk
- Reserve local storage for HTML, images, stylesheets, scripts, documents and duplicate query-string variants.
- Set depth, size and file-type limits for very broad sites.
- Preserve the source URL and capture date if the copy will be used as evidence.
- Obtain permission where required and follow robots.txt and the site owner’s rules. HTTrack’s documentation places responsibility for copying on the operator.
Install HTTrack and make a basic mirror
HTTrack is free software under GPL version 3 or later. The official site lists Windows, macOS/Linux/Unix/BSD and Android interfaces. Its current product page lists version 3.50-4, dated 09/25/2026, with HTTPS support, files larger than 2 GB, long Windows paths and WARC output among the listed changes and features. Use the interface appropriate to your operating system.
Command-line example
Install HTTrack from the official site, then run:
httrack "https://example.com/" -O "./example-mirror"
The first argument is the final canonical URL; -O selects the local output directory. HTTrack follows links recursively, reconstructs relative paths and stores the downloaded files for offline browsing. When complete, open the generated index.html (or the site’s corresponding entry page) in a browser.
Graphical workflow
- Create a new project and choose a local destination with enough free space.
- Enter the final canonical URL, not an earlier HTTP or apex URL that redirects elsewhere.
- Choose the default download action for a first public-site mirror.
- Before starting, review scan rules, limits and the list of allowed hosts.
- Run the crawl, watch its progress, and keep the project directory so you can resume or update it later.
Control how far the crawl reaches
Filters and related hosts
HTTrack’s default behavior stays on the starting host while following links to any depth. Add explicit include rules for required subdomains or CDNs, and exclude paths such as private dashboards, search-result URLs, calendars or unbounded faceted navigation. A broad “follow everything” setting can pull in unrelated resources and enormous URL spaces.
Depth and size limits
Use a maximum link depth when you need only the homepage and a few levels of navigation. Add total-size or per-file limits when storage or bandwidth is constrained. A limit is a coverage trade-off: record it with the capture so another person knows why deeper pages are absent.
Rank #2
- Easily store and access 1TB to content on the go with the Seagate Portable Drive, a USB external hard drive.Specific uses: Personal
- Designed to work with Windows or Mac computers, this external hard drive makes backup a snap just drag and drop. Reformatting may be required for Mac
- To get set up, connect the portable hard drive to a computer for automatic recognition no software required
- This USB drive provides plug and play simplicity with the included 18 inch USB 3.0 cable
- The available storage capacity may vary.
Sitemap seeding
Pages that are not linked from the starting page can be supplied through a sitemap. HTTrack’s guide says sitemap support is off by default; enable it deliberately. Sitemap URLs still pass through your host, path and filter rules, so a sitemap does not override scope.
Documents and PDFs
PDFs and other downloadable files are captured when they are discoverable and their host is allowed. If a document is linked only after JavaScript runs, or is returned after an authenticated request, HTTrack will not discover it through a normal static crawl. Add known document paths or sitemap entries within your permitted scope, then verify the files locally.
What a crawler cannot reliably capture
JavaScript-generated URLs and application state
HTTrack parses HTML and CSS; it does not run JavaScript. Links, images or API calls created only at runtime can therefore be missing. Client-side search, infinite scroll, route transitions and interactive filters often need a browser automation workflow or a separate export from the application.
Authentication and server restrictions
Login-only pages, IP restrictions, paywalls, signed URLs and content requiring a particular session may not be available to an unauthenticated crawler. Do not attempt to bypass access controls. If you administer the site, create an authorized export or provide a controlled test account and document the permitted scope.
External, personalized and streaming resources
Resources outside allowed hosts, user-specific recommendations, real-time dashboards and streaming media may not replay locally. A mirror can contain the page shell while the underlying service remains unavailable.
Rank #3
- High capacity in a small enclosure – The small, lightweight design offers up to 6TB* capacity, making WD Elements portable hard drives the ideal companion for consumers on the go.
- Plug-and-play expandability
- Vast capacities up to 6TB[1] to store your photos, videos, music, important documents and more
- SuperSpeed USB 3.2 Gen 1 (5Gbps)
Resume, update and preserve a capture
Interrupted downloads
Keep the project directory and rerun the same project rather than deleting partial files. HTTrack documents resume behavior so completed files do not need to be fetched again unnecessarily.
Refresh an existing mirror
Use HTTrack’s update function against the existing project when you need a later copy. Keep separate capture dates if you need an audit trail; an update is not the same artifact as the original crawl.
WARC and WACZ output
For preservation and replay workflows, HTTrack’s command guide describes WARC output and WACZ packaging. Select the format your archive or replay software accepts, and retain the crawl configuration, logs, source URLs and capture date beside the package.
Verify that the local copy is useful
- Read
hts-log.txtandhts-err.txt. The guide says these identify URLs that were refused, redirected or filtered. - Open the local entry page without network access and click representative internal links.
- Check several page depths, images, stylesheets, JavaScript-dependent areas and downloadable PDFs.
- Compare a sample of local URLs with the live site’s canonical URLs and note redirects or missing query parameters.
- Search for obvious gaps such as an empty directory, broken navigation, missing fonts or links that still point to the live domain.
- Record excluded hosts, depth and size limits, authentication boundaries and any known dynamic sections.
A successful process is not proof of complete coverage. The logs and representative tests tell you what the mirror actually contains.
Common problems and fixes
The crawl stops at the homepage
Cause: the starting URL redirected to another hostname, or filters exclude the destination. Fix: restart at the final canonical host and review allowed-host and path rules.
Rank #4
- Plug-and-play expandability
- SuperSpeed USB 3.2 Gen 1 (5Gbps)
Images or CSS are missing
Cause: assets are served from a CDN outside scope, blocked by a filter, or inserted by JavaScript. Fix: permit the specific asset host and paths you are authorized to copy, then recrawl; static HTTrack cannot discover runtime-only assets.
Quick wins for a faster PC:
Repair Windows errors before they cause bigger problemsFix Now →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Clear out junk files and repair common Windows errorsFree Scan →PDF links are absent
Cause: the files are unlinked, generated after a script action, protected by login or filtered by extension/path. Fix: seed known URLs or a sitemap, check filters and authentication, and verify the server permits the requests.
The mirror is unexpectedly huge
Cause: faceted navigation, calendars, query strings, third-party hosts or unlimited depth created a large URL space. Fix: exclude parameterized paths, tighten hosts and depth, and set size limits before rerunning.
Pages look blank or interactive controls do nothing
Cause: the application depends on JavaScript, API calls, a session or streaming services. Fix: capture a static export or use an authorized browser-based workflow for the specific states you need; do not claim the static mirror represents the full application.
The server responds slowly or blocks requests
Cause: throttling, robots.txt, rate limits or access policy. Fix: respect the site’s restrictions and HTTrack’s politeness defaults, reduce scope and request rate, and ask the owner for an approved window. Never advise bypassing a block.
The Tool Desk
Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Best Value
- 【Upgraded version】 - The mirror logo strip is combined with the striped non-slip design. The rounded corners of the shell are more suitable for holding. The strips play a heat dissipation function to ensure a stable and fast transmission process.
- 【Ultra-thin and quiet】 - The motherboard adopts JMicron 578 noise-free solution, giving you a quiet working environment. Lightweight and portable size designed to fit in your pocket for easy portability.
- 【Ultra-Fast Data Transfers】 - Pairing this external hard drive with JMicron 578 solution USB 3.0 and USB 2.0 interfaces enables blazing-fast data transfer. It boasts theoretical read speeds of up to 125MB/s and write speeds of up to 103MB/s.
- 【Plug and Play】 - With no software to install, just plug it in and the drive is ready to use.The hard disk chip is wrapped with an aluminum anti-interference layer to increase heat dissipation and protect data.
- 【What You Get】 - 1 x Portable Hard Drive, 1 x USB 3.0 Cable, 1 x User Manual, Gift-type shell packaging ,Three-year manufacturer's warranty and free technical support services.
Or skip the browser setup
For a screenshot of a page rather than a browsable, linked archive, ScreenshotNeo provides a single GET request. It accepts the consent banner like a visitor and removes more than 60 known consent platforms, newsletter popups and chat widgets before capture; each step can be disabled. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads and cache hits are not billed, and response headers identify the page verdict and billing status. It also offers an MCP server for AI agents with take_screenshot, get_page_info and capture_pdf.
cURL
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
Node.js
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
Replace the example URL with the page you need. The ScreenshotNeo documentation covers the 63 capture options, including full-page shots, lazy-image loading, CSS-selector elements, dark mode, device and retina settings, PDF paper and page ranges, custom CSS or JavaScript, clicks, waits, blocked resources, cookies and headers, timezone and geolocation, transparent backgrounds, resizing, caching, signed links, asynchronous webhooks, bulk capture and usage reporting.
| Plan | Included shots | Price |
|---|---|---|
| Free | 1,000/month | No card |
| Starter | 3,000 | $5 |
| Growth | 15,000 | $15 |
| Pro | 60,000 | $39 |
| Scale | 250,000 | $99 |
| Business | 1,000,000 | $249 |
Every feature is on every plan, and yearly billing gives two months free. If you need clean screenshots rather than an offline link graph, start with the free ScreenshotNeo account: 1,000 screenshots per month, no card required.
When to use each approach
- HTTrack: choose it for a local, navigable copy of discoverable public pages, with control over hosts, filters, depth and archival formats.
- Save Page Now: choose it when one public page and its resources are enough.
- Archive-It: choose it when an organization needs recurring, managed collections and support.
- ScreenshotNeo: choose it when the deliverable is a clean image or PDF, an API result, an AI-agent capture or a repeatable page-level workflow rather than a crawlable offline site.
Frequently Asked Questions
Can a mirror replace the original website for legal or evidentiary purposes?
Not automatically. Preserve the source URL, capture date, configuration and logs, and have the intended court, regulator or records policy determine whether the copy is acceptable.
Free tools Windows power users keep installed
One-click scans. No signup required.
How do I capture pages that are not linked anywhere?
Provide known URLs or a sitemap as crawl seeds, subject to HTTrack’s host and filter rules. If the pages require a session or are generated only after JavaScript actions, use an authorized export or browser workflow instead.
Should I store the mirror on an external drive?
Use external storage when the mirror exceeds available internal space. The appropriate capacity depends on the site and limits you set; there is no universal drive size.
The Bottom Line
For a linked public-site copy, HTTrack is the practical starting point: begin at the final host, constrain scope, respect access rules, and verify the logs and local navigation. Use an archival package or managed service when preservation requirements exceed a static mirror, and use ScreenshotNeo when you need clean page screenshots or PDFs instead of a whole-site link graph.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

