Driver FixRecommendedSound, Wi-Fi or graphics acting up? Check drivers firstFind missing or outdated drivers fast.Check DriversOctober DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsClean PCRecommendedOne scan can reveal what keeps slowing WindowsLook for cleanup and repair opportunities.Run Scan×
Skip to content
Apify

6 Best Node.js Web Scrapers in 2026: Choose by Page Type and Workflow

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The best Node.js scraper depends on what the site returns and how much automation the job needs. For markup already in the HTTP response, start with Cheerio or Node’s built-in fetch. For content rendered by JavaScript, use Playwright or Puppeteer. For a recurring crawl with queues and structured output, use Crawlee. If you want hosted execution and operational tooling rather than running the scraper yourself, consider the Apify platform.

This is a fit-based comparison, not a benchmark ranking: a parser, browser-automation library, crawler framework and hosted platform solve different layers of the problem.

How to choose a Node.js scraper

First find out whether the information you need exists in the HTML returned by an ordinary HTTP request. If it does, parsing that response is usually simpler than opening a browser. If the page fills in its content with JavaScript, you need browser execution or a suitable data endpoint. If the work involves many pages, link discovery, queues and managed output, a crawler framework or hosted platform may be a better fit than a single-page script.

  • HTML is already there: use Cheerio to parse it, often alongside fetch.
  • The browser creates the content: use Playwright or Puppeteer.
  • You need a repeatable crawl: use Crawlee’s crawler classes and shared interface.
  • You want hosted runs and operations: evaluate Apify as a platform, not just as another local library.

Respect the target site’s access rules and applicable requirements. The sources cited here do not establish legal advice or a universal permission rule for scraping.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

At-a-glance comparison

Option What it does Best fit Key limitation or requirement
Cheerio Parses supplied HTML and XML Extracting data from static or server-rendered markup Does not render pages or execute JavaScript; current docs list Node.js 22.19 or later. Cheerio documentation
Playwright Automates real browsers JavaScript-rendered pages and browser interactions Downloads browser binaries; current installation docs list Node.js 22.x, 24.x or 26.x. Playwright documentation
Puppeteer Controls Chrome or Firefox Browser automation that fits a Puppeteer/Chrome-oriented stack The puppeteer package downloads compatible Chrome; puppeteer-core does not. Puppeteer installation
Crawlee Provides crawler classes for HTTP and browser crawling Multi-page crawling, link queues and dataset output More structure than a one-off request may need; quick start says Node.js 16 or later. Crawlee quick start
Node.js fetch + Undici Makes HTTP requests Simple retrieval or calling a suitable endpoint Not a crawler framework or HTML parser. Node.js fetch guide
Apify platform / JavaScript SDK Runs scraper Actors on a hosted platform Hosted execution, scheduling, monitoring or ready-made scrapers A service/platform choice rather than a like-for-like local library. Apify SDK documentation

1. Cheerio: parse HTML without launching a browser

Cheerio offers a jQuery-like API for working with HTML and XML that you already have. It is a strong starting point when a page’s useful information is present in the response body. The important boundary is that Cheerio parses markup; it does not render a page or run its JavaScript. Its documentation puts it plainly: “Cheerio is not a web browser.” Cheerio documentation

When to choose it

  • The server returns the elements and text you need in the initial HTML.
  • You want familiar selector-based extraction without browser startup and browser binaries.
  • You can retrieve pages yourself and want a separate parser for the response.

Runnable example

The following ES module fetches a page and extracts link text and URLs. Use a page you are permitted to retrieve; selectors are site-specific and may need adjustment.

import * as cheerio from 'cheerio';

const url = 'https://example.com/';
const response = await fetch(url);
if (!response.ok) {
  throw new Error(`Request failed: ${response.status} ${response.statusText}`);
}

const html = await response.text();
const $ = cheerio.load(html);
const links = $('a').map((_, element) => ({
  text: $(element).text().trim(),
  href: $(element).attr('href') ?? null,
})).get();

console.log(links);

Install Cheerio with npm install cheerio. Current Cheerio docs specify Node.js 22.19 or later and document both import and require usage; check its current documentation if your project’s runtime differs. Cheerio introduction

2. Playwright: run the page in a browser

Choose Playwright when a page needs JavaScript execution or user-like browser actions before the data appears. It supports Chromium, WebKit and Firefox, and is also a good fit when your project already uses Playwright for testing or browser automation. Playwright documentation

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Install and run

Install the package and its browser binaries. The Playwright installation documentation lists Node.js 22.x, 24.x or 26.x. Browser downloads and the resources needed to run them are part of the operational cost of this approach. Playwright installation

npm init -y
npm install playwright
npx playwright install

Example using Chromium:

import { chromium } from 'playwright';

const browser = await chromium.launch({ headless: true });
try {
  const page = await browser.newPage();
  await page.goto('https://example.com/', { waitUntil: 'domcontentloaded' });
  await page.locator('h1').waitFor();
  const heading = await page.locator('h1').first().textContent();
  console.log(heading?.trim() ?? null);
} finally {
  await browser.close();
}

Practical trade-offs

Browser automation can reach content that is absent from the original HTML, but it requires browser startup and page execution. Select the narrowest reliable wait condition for the page rather than assuming every site becomes ready at the same time. A selector wait can be more targeted than waiting an arbitrary fixed delay; pages with ongoing network activity may not become fully idle.

Crawlee’s documentation describes PlaywrightCrawler as supporting Chromium, Chrome, Firefox, WebKit and other browsers. Playwright’s own installation page documents Chromium, WebKit and Firefox. Use the current version-specific documentation for browser availability and installation details. Crawlee quick start Playwright documentation

3. Puppeteer: browser control for a Puppeteer-oriented stack

Puppeteer automates browsers and is a reasonable choice when its API and Chrome ecosystem fit your application. Do not treat it as Chrome-only: its current documentation describes controlling Chrome or Firefox through DevTools Protocol or WebDriver BiDi. Puppeteer documentation

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Choose the package deliberately

  • puppeteer downloads a compatible Chrome as part of installation.
  • puppeteer-core does not download a browser; use it when you manage the browser separately.

That difference affects installation size, deployment setup and which browser executable your code will launch. Puppeteer installation

Runnable example

import puppeteer from 'puppeteer';

const browser = await puppeteer.launch({ headless: true });
try {
  const page = await browser.newPage();
  await page.goto('https://example.com/', { waitUntil: 'domcontentloaded' });
  const heading = await page.$eval('h1', element => element.textContent?.trim() ?? '');
  console.log(heading);
} finally {
  await browser.close();
}

Use this when browser execution is necessary and Puppeteer fits the existing stack. If you need a broader choice of browser engines or are already standardizing on Playwright, compare the current browser support and setup documentation before choosing. Neither is a universal speed winner based on the sources cited here.

4. Crawlee: organize a crawl, not just a page request

Crawlee is aimed at crawling workflows. Its shared interface includes CheerioCrawler, PuppeteerCrawler and PlaywrightCrawler, letting you choose HTTP-based parsing or browser execution within the same framework. The quick start demonstrates queuing links and writing records to a local JSON dataset. It lists Node.js 16 or later. Crawlee quick start

Pick the crawler class by page behavior

  • CheerioCrawler: plain HTTP retrieval and HTML parsing; it cannot handle JavaScript rendering.
  • PuppeteerCrawler: browser control with Chromium or Chrome.
  • PlaywrightCrawler: browser control with the broader browser set described in Crawlee’s documentation.

Example: queue pages and save records

The quick start’s pattern is to add initial URLs, extract data in a request handler, enqueue discovered links, then export the dataset. Here is a compact PlaywrightCrawler example:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
import { PlaywrightCrawler } from 'crawlee';

const crawler = new PlaywrightCrawler({
  async requestHandler({ request, page, enqueueLinks, pushData }) {
    const title = await page.title();
    await pushData({ url: request.url, title });
    await enqueueLinks({ selector: 'a', strategy: 'same-domain' });
  },
});

await crawler.run(['https://example.com/']);

Crawlee’s quick start documents exporting local dataset records to JSON. Use its queue and storage structure when those capabilities solve a real operational need; for one isolated URL, the framework can add more setup than a direct request and parser. Crawlee quick start

5. Node.js fetch and Undici: the minimal HTTP baseline

Node’s built-in fetch is powered by Undici, according to the Node.js learning documentation. It is useful when you need to retrieve a response or call an endpoint and the returned data is already usable. It does not, by itself, parse HTML into convenient selectors, execute page JavaScript or manage a crawl. Node.js fetch guide

Runnable retrieval example

const response = await fetch('https://example.com/');
if (!response.ok) {
  throw new Error(`HTTP ${response.status}: ${response.statusText}`);
}

const html = await response.text();
console.log(html.slice(0, 500));

Pair the response with Cheerio if you need to select elements from HTML. Prefer a documented data endpoint when one is available and appropriate; a browser is not automatically necessary just because the content is on a website.

6. Apify: hosted execution and managed scraper operations

Apify is a platform route for running scrapers as hosted Actors, rather than simply a local Node.js parsing or browser library. Its official JavaScript/TypeScript SDK creates Actors, while the platform supports running them at scale with monitoring and scheduling. Apify also offers ready-made scrapers, including browser-based options and HTTP-plus-Cheerio approaches. Apify SDK Apify scraping documentation

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

When the platform route fits

  • You want to run jobs in a hosted environment instead of provisioning and operating the runtime yourself.
  • Scheduling and monitoring are part of the problem, not afterthoughts.
  • A ready-made Actor matches the site and data you need closely enough to avoid building the extraction workflow from scratch.

Apify’s SDK documentation identifies it as the official JavaScript/TypeScript Actor library and showed version 3.7 when checked on 2026-09-30 UTC. Hosted operations change the deployment and service decision; compare current platform documentation and terms for your specific workload. Apify SDK documentation

Decision guide: which one should you start with?

  1. Inspect the response first. If the data is in the returned HTML, try fetch plus Cheerio.
  2. Check whether browser execution is essential. If JavaScript produces the content or the workflow needs browser interaction, use Playwright or Puppeteer.
  3. Assess the crawl shape. If you need to discover and queue many pages and structure results, look at Crawlee rather than hand-building orchestration around one-page scripts.
  4. Decide where it should run. If you need hosted scheduling and monitoring, assess Apify as a platform.
  5. Check runtime compatibility. The current documentation requirements differ: Cheerio lists Node.js 22.19 or later; Playwright lists Node.js 22.x, 24.x or 26.x; Crawlee’s quick start says Node.js 16 or later. Confirm requirements for the versions you intend to install. These details were checked on 2026-09-30 UTC. Cheerio Playwright Crawlee

Performance, reliability and cost considerations

There is no supported universal speed ranking among these choices. An HTTP request plus parser avoids browser execution when the needed content is in the response; browser automation performs more work because it runs a page. But the right comparison depends on the target, page behavior and workload, and the sources here do not provide an independent benchmark across all six options.

  • Keep the work proportional: use a parser for static markup rather than launching a browser without a need.
  • Make waits specific: wait for the selector or condition that signals the data is ready; avoid assuming a fixed delay works for every page.
  • Handle failures explicitly: check HTTP status for direct requests and ensure browsers close in a finally block.
  • Account for operations: browser binaries and browser runtime are part of local deployment; hosted platforms shift some execution and operational concerns to a service.
  • Do not generalize vendor comparisons: Apify’s documentation makes a specific claim that its Cheerio Scraper can be as much as 20 times faster than its full-browser Puppeteer solution for the static-content use case it describes. This is Apify’s own claim, not an independent benchmark across the products in this guide. Apify documentation
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Troubleshooting common scraping problems

The extracted HTML has no target content

Likely cause: the browser generates the content after the initial response, so a plain HTTP fetch or Cheerio parse cannot see it. Fix: inspect the page’s behavior and use Playwright or Puppeteer if browser execution is needed, or use an appropriate endpoint if one is available.

A selector returns no matches

Likely cause: the selector does not match the actual response or rendered page, or the page has not reached the state you expect. Fix: inspect the markup you are parsing, verify the selector, and for browser automation wait for a relevant locator before extracting.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Playwright cannot find a browser

Likely cause: the required browser binaries were not installed in the environment. Fix: install the package’s required browsers with npx playwright install and make sure the deployment environment includes them. Playwright installation

Puppeteer launches unsuccessfully after using puppeteer-core

Likely cause: puppeteer-core does not download a browser. Fix: provide a compatible browser installation and configure the executable for your setup, or choose the full puppeteer package if its bundled compatible Chrome download suits your deployment. Puppeteer installation

The request fails or returns an unexpected status

Likely cause: the server response is not successful, the URL is wrong, or the endpoint does not return the expected page. Fix: log the status and response details, check the requested URL, and handle non-2xx responses before parsing. Do not assume that a successful network connection means the expected data was returned.

The crawl grows beyond the pages you intended

Likely cause: link discovery is enqueueing more URLs than the intended scope. Fix: constrain discovered links by domain and the selectors or URL patterns that define the task, and verify the queue behavior on a small scope before a larger run.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Or skip the browser setup

If your task is to capture a visual screenshot rather than extract structured fields, ScreenshotNeo is a website screenshot API and MCP server. A single GET request can return a PNG, JPEG, WebP or PDF. It is an alternative to try first when you need a screenshot without managing a browser locally: it accepts cookie or consent banners before capture and removes more than 60 known consent platforms, newsletter popups and chat widgets; bot checks, blank pages and failed loads are not billed; and its MCP server lets AI agents use take_screenshot, get_page_info and capture_pdf. Those cleanup steps can each be turned off.

cURL example, using the documented API call pattern and a target URL:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

See the ScreenshotNeo documentation for request options and formats. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. Sign up for ScreenshotNeo’s free plan.

Frequently asked questions

Can I scrape websites in Node.js without a headless browser?

Yes. Use fetch for retrieval and a parser such as Cheerio when the response already contains the content you need. A headless browser is needed when the task depends on browser execution or interaction, not for every website request.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Are Cheerio and Crawlee alternatives to one another?

Not exactly. Cheerio parses supplied markup. Crawlee is a crawling framework that offers a Cheerio-based crawler as well as browser-based crawler classes, so it can provide orchestration around different retrieval approaches.

Should I use Playwright or Puppeteer?

Choose based on the browser support, APIs, installation model and existing automation stack your team needs. Both offer browser control; consult their current documentation for version-specific compatibility rather than assuming one is always better.

Is Apify a Node.js scraper library?

Its JavaScript SDK is a library for creating Actors, but Apify is best understood here as a hosted platform and operations option, not as a direct equivalent to a local HTML parser.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Read next

Recommended PC Tool
Recommended PC Tool
Outdated Drivers Are Slowing You DownFree scan - exact matches
PC Slower Than It Used to Be?Free scan - under a minute

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.