Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

JavaScript alone does not make a complete website downloader. For a conventional offline copy, use a recursive crawler to retrieve pages and assets linked in the site’s HTML, XHTML, or CSS. If essential pages or links appear only after browser-side JavaScript runs, use browser automation to render and inspect pages, then build a bounded crawler around what the browser sees. Neither method guarantees a perfect copy of every site. A browser-triggered file download is also not the same thing as mirroring a whole website.

What “download an entire website” means

A website is not necessarily a folder of files that can be copied in one request. A page may link to other pages and assets in its markup, or it may use scripts to fetch content and create links after it loads. Some sites also depend on logins, server-side state, APIs, forms, or interactions that do not translate into a usable offline copy.

Before writing code, decide what “entire” means for your task. Set a starting URL, allowed hosts and paths, exclusions, a maximum number of pages or crawl depth, and which assets matter. A finite boundary helps prevent a crawler from wandering through search results, calendars, or query combinations that can generate effectively endless URLs.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Markup-based site: linked pages and assets are discoverable in ordinary HTML, XHTML, or CSS. A recursive retriever such as GNU Wget is a practical starting point.
  • JavaScript-rendered site: important content or links appear only after client-side code executes. A parser such as Cheerio cannot reveal those by itself; use a real browser through automation such as Playwright.
  • Single file download: a page triggers a downloadable file. Playwright can observe and save that event, but this does not recursively mirror the site.

Judge an approach by whether it executes JavaScript, traverses linked pages and assets, lets you limit host and path scope, persists downloads, and produces files that work when opened locally. The available official documentation does not establish a universal speed or completeness ranking.

#1 Best Overall
Sale
Seagate 2TB Portable Hard Drive | USB 3.0 (STGX2000400)
  • Easily store and access 2TB to content on the go with the Seagate Portable Drive, a USB external hard drive
  • Designed to work with Windows or Mac computers, this external hard drive makes backup a snap just drag and drop
  • To get set up, connect the portable hard drive to a computer for automatic recognition no software required
  • This USB drive provides plug and play simplicity with the included 18 inch USB 3.0 cable
  • The available storage capacity may vary.

Can JavaScript download all pages from a website?

Not reliably with a short browser script. JavaScript running inside a page is constrained by browser security rules and is not automatically a crawler with access to every page, host, or resource. In particular, a script that fetches the current page’s HTML does not execute that HTML’s scripts or discover everything those scripts might request later.

Cheerio is useful for parsing already-fetched HTML or XML, but it does not execute JavaScript, render CSS, or load external resources. It can extract links present in markup; it cannot stand in for a browser on pages whose essential content is inserted after scripts run. See Cheerio’s documentation for its stated role. A browser automation tool can execute page code, but you still have to define a crawl boundary, discover pages, save responses or rendered results, and check the output.

Choose the method that matches the site

Situation Starting point What it does and does not establish
Links and content appear in ordinary markup; you want a conventional offline mirror. GNU Wget recursive retrieval. Wget documents recursive retrieval, reconstruction of a remote directory structure, and conversion of links for local browsing. A bounded crawl and local checks are still needed.
Essential content or links appear after JavaScript executes. Playwright with a real browser, plus your own bounded discovery and saving logic. A browser can run client-side code. Playwright’s cited download API is for page-initiated file downloads, not a turnkey full-site export.
You already have HTML and need to inspect its markup. Cheerio. It parses markup; it does not execute scripts, render pages, or fetch dependent resources.

These are complementary tools rather than interchangeable “download the internet” commands. The simplest suitable method is usually the one that matches where the site exposes its links and content.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Prepare a responsible, bounded crawl

  1. Check the intended use. Review the site’s terms and any permission requirements before crawling or redistributing material. General tool documentation cannot determine whether a particular site permits your use.
  2. Inspect robots.txt. GNU Wget says it respects the Robot Exclusion Standard. Google explains that robots.txt is mainly intended to manage crawler access and request load, not to keep a page out of Google or protect private files. It is crawler guidance, not a security boundary or permission grant.
  3. Choose a finite scope. Decide the starting page, allowed host or hosts, path rules, exclusions, and a practical page or depth cap. Avoid unbounded search, calendar, and faceted-navigation URLs.
  4. Choose based on rendering. If links are in markup, start with recursive retrieval. If critical content appears only after scripts run, use browser automation for discovery and rendering; parsing alone is insufficient.
  5. Plan persistence and verification. Save needed downloads before closing the browser context, then inspect representative local pages and assets.

Method 1: mirror markup-based pages with Wget

For a site whose pages and asset links are visible in ordinary markup, Wget’s recursive mode is the direct route to a conventional offline mirror. A typical command is:

Rank #2
Seagate Portable 1TB External Hard Drive HDD – USB 3.0 for PC, Mac, PlayStation, & Xbox, 1-Year Rescue Service (STGX1000400) , Black
  • Easily store and access 1TB to content on the go with the Seagate Portable Drive, a USB external hard drive.Specific uses: Personal
  • Designed to work with Windows or Mac computers, this external hard drive makes backup a snap just drag and drop. Reformatting may be required for Mac
  • To get set up, connect the portable hard drive to a computer for automatic recognition no software required
  • This USB drive provides plug and play simplicity with the included 18 inch USB 3.0 cable
  • The available storage capacity may vary.
wget --mirror --convert-links --adjust-extension --page-requisites --no-parent https://example.com/

Replace https://example.com/ with the intended starting URL. The command requests a recursive mirror, converts links for local browsing, adjusts file extensions where appropriate, retrieves page requisites such as linked assets, and avoids ascending to the parent directory. Review Wget’s official overview of recursive retrieval and link conversion and its option documentation for the behavior relevant to your use.

This is a starting point, not a guarantee that the result is complete or that every URL on the host is in scope. Set boundaries suitable for the site, inspect what was retrieved, and do not assume that a mirror captures authenticated content, application state, or resources that are not discoverable through the crawl.

Method 2: use Playwright when JavaScript creates the page

Playwright launches a browser that can execute client-side JavaScript. The following Node.js example shows the browser side of the job: load a page, wait for it to settle, and list links visible in the rendered DOM. It is not a full-site downloader; it is a starting point for a crawler that would need to apply host/path rules, visit discovered pages, persist the data you need, and avoid revisiting URLs.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
import { chromium } from 'playwright';

const startUrl = 'https://example.com/';
const allowedHost = new URL(startUrl).host;
const browser = await chromium.launch();
const page = await browser.newPage();

try {
  await page.goto(startUrl, { waitUntil: 'networkidle', timeout: 60000 });
  const links = await page.locator('a[href]').evaluateAll(anchors =>
    anchors.map(anchor => anchor.href)
  );

  const internalLinks = [...new Set(links)].filter(href => {
    try {
      return new URL(href).host === allowedHost;
    } catch {
      return false;
    }
  });

  console.log(internalLinks);
} finally {
  await browser.close();
}

Install Playwright and a supported browser in your project before running the example; the Playwright browsers documentation explains browser installation and hosting considerations. The script prints discovered same-host links only. To turn it into a crawler, add a queue and visited set, apply path exclusions and a page cap, and decide how to save each page’s HTML and required assets. Sites that load content on scroll or after a user action may require additional page-specific steps. Treat these as implementation requirements, not evidence that any generic script can reproduce every site.

Rank #3
Sale
WD 2TB Elements Portable External Hard Drive for Windows, USB 3.2 Gen 1/USB 3.0 for PC & Mac, Plug and Play Ready - WDBU6Y0020BBK-WESN
  • High capacity in a small enclosure – The small, lightweight design offers up to 6TB* capacity, making WD Elements portable hard drives the ideal companion for consumers on the go.
  • Plug-and-play expandability
  • Vast capacities up to 6TB[1] to store your photos, videos, music, important documents and more
  • SuperSpeed USB 3.2 Gen 1 (5Gbps)

Saving a page-triggered file is a different task

If clicking a link or button initiates a file download, Playwright provides a download event and a way to save that file to a chosen path. Its documentation notes that downloads belong to a browser context and are removed when that context closes unless they are saved. This is useful for preserving a file offered by a page, but it does not traverse the site or make an offline mirror. Consult the official Playwright download documentation for the event and persistence API.

Or skip the browser setup

If your task is to capture a page as an image or PDF rather than build a navigable offline mirror, ScreenshotNeo offers a one-request screenshot API and an MCP server. It does not replace a recursive website download: it returns a capture of a requested page. For example, this cURL request saves a WebP capture of the target URL:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp

See the ScreenshotNeo API documentation for request options. Cookie banners, popups, and chat widgets are removed before the shot; bot checks, blank pages, and failed loads are never billed. An MCP server lets AI agents take screenshots, and 1,000 screenshots a month are free with no card; paid plans start at $5 for 3,000. Sign up for 1,000 free screenshots a month with no card.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Verify the offline copy

Open representative saved pages locally rather than judging success from the number of files alone. Check internal navigation, images, stylesheets, scripts, and documents that matter to your use. Try pages from different parts of the site and follow links in the saved copy. A local file may still point to the live site, depend on a remote API, or omit content that required a script, interaction, or login during capture.

Record the boundary you used—starting URL, included host and paths, exclusions, and any page cap—so others understand what the copy includes. Describe it as a bounded offline copy, not a perfect clone, unless you have independently verified the requirements that matter.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Troubleshooting common gaps

The saved page is empty or missing key text

If the live page populates content only after JavaScript runs, a markup parser or basic recursive fetch may see only the initial HTML. Use a browser-rendering approach such as Playwright to inspect the rendered page, and account for any page-specific wait or interaction needed before capture.

Links to deeper pages are missing

Check whether the links are present in the initial markup or created only after scripts execute. If they are markup links, check the crawler’s recursion and scope settings. If they appear after rendering or interaction, a non-rendering parser will not discover them.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Images or styles do not work offline

Confirm that the relevant resources were retrieved and that saved page references point to local files. A successful page download does not by itself prove that dependent resources are present or locally addressable.

A Playwright download disappears after the script exits

Save the download to a chosen persistent path before closing its browser context. Playwright documents that context-owned downloads are removed when the context closes unless saved.

Best Value
Sale
UnionSine 1TB Ultra Slim Portable External Hard Drive HDD-USB 3.0
  • 【Upgraded version】 - The mirror logo strip is combined with the striped non-slip design. The rounded corners of the shell are more suitable for holding. The strips play a heat dissipation function to ensure a stable and fast transmission process.
  • 【Ultra-thin and quiet】 - The motherboard adopts JMicron 578 noise-free solution, giving you a quiet working environment. Lightweight and portable size designed to fit in your pocket for easy portability.
  • 【Ultra-Fast Data Transfers】 - Pairing this external hard drive with JMicron 578 solution USB 3.0 and USB 2.0 interfaces enables blazing-fast data transfer. It boasts theoretical read speeds of up to 125MB/s and write speeds of up to 103MB/s.
  • 【Plug and Play】 - With no software to install, just plug it in and the drive is ready to use.The hard disk chip is wrapped with an aluminum anti-interference layer to increase heat dissipation and protect data.
  • 【What You Get】 - 1 x Portable Hard Drive, 1 x USB 3.0 Cable, 1 x User Manual, Gift-type shell packaging ,Three-year manufacturer's warranty and free technical support services.

The crawl keeps finding URLs

Inspect the URL patterns being discovered. Search, calendar, and faceted paths can generate large or unbounded URL spaces; constrain allowed paths and use a finite crawl boundary rather than following every link indiscriminately.

The local page still depends on the live site

Inspect its links, scripts, and network-dependent behavior. A page that calls a remote API or expects a server-side session may not function offline just because its HTML was saved. Decide whether your goal is readable saved content or a working offline application; the latter may require site-specific handling beyond a generic crawler.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Frequently Asked Questions

Does robots.txt give permission to copy a website?

No. It communicates crawler guidance; it is neither a security boundary nor a grant of permission to crawl or republish content. Check the site’s terms and applicable permissions separately.

Can Cheerio download a JavaScript website?

Cheerio parses HTML and XML markup but does not execute JavaScript, render CSS, or load dependent resources. It can help inspect markup you have already fetched.

Does Playwright’s download API create an offline website?

No. Its download API handles a file initiated by a page and lets you save it. A site mirror requires separate bounded page discovery, retrieval, persistence, and verification.

Quick Recap

SaleBestseller No. 1
Seagate 2TB Portable Hard Drive | USB 3.0 (STGX2000400)
Seagate 2TB Portable Hard Drive | USB 3.0 (STGX2000400)
This USB drive provides plug and play simplicity with the included 18 inch USB 3.0 cable; The available storage capacity may vary.
$129.99
Bestseller No. 2
Seagate Portable 1TB External Hard Drive HDD – USB 3.0 for PC, Mac, PlayStation, & Xbox, 1-Year Rescue Service (STGX1000400) , Black
Seagate Portable 1TB External Hard Drive HDD – USB 3.0 for PC, Mac, PlayStation, & Xbox, 1-Year Rescue Service (STGX1000400) , Black
This USB drive provides plug and play simplicity with the included 18 inch USB 3.0 cable; The available storage capacity may vary.
$119.80
SaleBestseller No. 3

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.