The best Node.js scraper depends on what the site returns and how much automation the job needs. For markup already in the HTTP response, start with Cheerio or Node’s built-in fetch. For content rendered by JavaScript, use Playwright or Puppeteer. For a recurring crawl with queues and structured output, use Crawlee. If you want hosted execution and operational tooling rather than running the scraper yourself, consider the Apify platform.
This is a fit-based comparison, not a benchmark ranking: a parser, browser-automation library, crawler framework and hosted platform solve different layers of the problem.
How to choose a Node.js scraper
First find out whether the information you need exists in the HTML returned by an ordinary HTTP request. If it does, parsing that response is usually simpler than opening a browser. If the page fills in its content with JavaScript, you need browser execution or a suitable data endpoint. If the work involves many pages, link discovery, queues and managed output, a crawler framework or hosted platform may be a better fit than a single-page script.
- HTML is already there: use Cheerio to parse it, often alongside
fetch. - The browser creates the content: use Playwright or Puppeteer.
- You need a repeatable crawl: use Crawlee’s crawler classes and shared interface.
- You want hosted runs and operations: evaluate Apify as a platform, not just as another local library.
Respect the target site’s access rules and applicable requirements. The sources cited here do not establish legal advice or a universal permission rule for scraping.
Do these 3 things before closing this tab:
1Fix the driver behind crashes, sound loss and screen glitches2Clear out junk files and repair common Windows errors3Scan for outdated or missing drivers - takes under a minute#1 Best Overall
At-a-glance comparison
| Option | What it does | Best fit | Key limitation or requirement |
|---|---|---|---|
| Cheerio | Parses supplied HTML and XML | Extracting data from static or server-rendered markup | Does not render pages or execute JavaScript; current docs list Node.js 22.19 or later. Cheerio documentation |
| Playwright | Automates real browsers | JavaScript-rendered pages and browser interactions | Downloads browser binaries; current installation docs list Node.js 22.x, 24.x or 26.x. Playwright documentation |
| Puppeteer | Controls Chrome or Firefox | Browser automation that fits a Puppeteer/Chrome-oriented stack | The puppeteer package downloads compatible Chrome; puppeteer-core does not. Puppeteer installation |
| Crawlee | Provides crawler classes for HTTP and browser crawling | Multi-page crawling, link queues and dataset output | More structure than a one-off request may need; quick start says Node.js 16 or later. Crawlee quick start |
Node.js fetch + Undici |
Makes HTTP requests | Simple retrieval or calling a suitable endpoint | Not a crawler framework or HTML parser. Node.js fetch guide |
| Apify platform / JavaScript SDK | Runs scraper Actors on a hosted platform | Hosted execution, scheduling, monitoring or ready-made scrapers | A service/platform choice rather than a like-for-like local library. Apify SDK documentation |
1. Cheerio: parse HTML without launching a browser
Cheerio offers a jQuery-like API for working with HTML and XML that you already have. It is a strong starting point when a page’s useful information is present in the response body. The important boundary is that Cheerio parses markup; it does not render a page or run its JavaScript. Its documentation puts it plainly: “Cheerio is not a web browser.” Cheerio documentation
When to choose it
- The server returns the elements and text you need in the initial HTML.
- You want familiar selector-based extraction without browser startup and browser binaries.
- You can retrieve pages yourself and want a separate parser for the response.
Runnable example
The following ES module fetches a page and extracts link text and URLs. Use a page you are permitted to retrieve; selectors are site-specific and may need adjustment.
import * as cheerio from 'cheerio';
const url = 'https://example.com/';
const response = await fetch(url);
if (!response.ok) {
throw new Error(`Request failed: ${response.status} ${response.statusText}`);
}
const html = await response.text();
const $ = cheerio.load(html);
const links = $('a').map((_, element) => ({
text: $(element).text().trim(),
href: $(element).attr('href') ?? null,
})).get();
console.log(links);
Install Cheerio with npm install cheerio. Current Cheerio docs specify Node.js 22.19 or later and document both import and require usage; check its current documentation if your project’s runtime differs. Cheerio introduction
2. Playwright: run the page in a browser
Choose Playwright when a page needs JavaScript execution or user-like browser actions before the data appears. It supports Chromium, WebKit and Firefox, and is also a good fit when your project already uses Playwright for testing or browser automation. Playwright documentation
Free tools Windows power users keep installed
One-click scans. No signup required.
Install and run
Install the package and its browser binaries. The Playwright installation documentation lists Node.js 22.x, 24.x or 26.x. Browser downloads and the resources needed to run them are part of the operational cost of this approach. Playwright installation
npm init -y
npm install playwright
npx playwright install
Example using Chromium:
import { chromium } from 'playwright';
const browser = await chromium.launch({ headless: true });
try {
const page = await browser.newPage();
await page.goto('https://example.com/', { waitUntil: 'domcontentloaded' });
await page.locator('h1').waitFor();
const heading = await page.locator('h1').first().textContent();
console.log(heading?.trim() ?? null);
} finally {
await browser.close();
}
Practical trade-offs
Browser automation can reach content that is absent from the original HTML, but it requires browser startup and page execution. Select the narrowest reliable wait condition for the page rather than assuming every site becomes ready at the same time. A selector wait can be more targeted than waiting an arbitrary fixed delay; pages with ongoing network activity may not become fully idle.
Rank #2
Crawlee’s documentation describes PlaywrightCrawler as supporting Chromium, Chrome, Firefox, WebKit and other browsers. Playwright’s own installation page documents Chromium, WebKit and Firefox. Use the current version-specific documentation for browser availability and installation details. Crawlee quick start Playwright documentation
3. Puppeteer: browser control for a Puppeteer-oriented stack
Puppeteer automates browsers and is a reasonable choice when its API and Chrome ecosystem fit your application. Do not treat it as Chrome-only: its current documentation describes controlling Chrome or Firefox through DevTools Protocol or WebDriver BiDi. Puppeteer documentation
Choose the package deliberately
puppeteerdownloads a compatible Chrome as part of installation.puppeteer-coredoes not download a browser; use it when you manage the browser separately.
That difference affects installation size, deployment setup and which browser executable your code will launch. Puppeteer installation
Runnable example
import puppeteer from 'puppeteer';
const browser = await puppeteer.launch({ headless: true });
try {
const page = await browser.newPage();
await page.goto('https://example.com/', { waitUntil: 'domcontentloaded' });
const heading = await page.$eval('h1', element => element.textContent?.trim() ?? '');
console.log(heading);
} finally {
await browser.close();
}
Use this when browser execution is necessary and Puppeteer fits the existing stack. If you need a broader choice of browser engines or are already standardizing on Playwright, compare the current browser support and setup documentation before choosing. Neither is a universal speed winner based on the sources cited here.
4. Crawlee: organize a crawl, not just a page request
Crawlee is aimed at crawling workflows. Its shared interface includes CheerioCrawler, PuppeteerCrawler and PlaywrightCrawler, letting you choose HTTP-based parsing or browser execution within the same framework. The quick start demonstrates queuing links and writing records to a local JSON dataset. It lists Node.js 16 or later. Crawlee quick start
Pick the crawler class by page behavior
CheerioCrawler: plain HTTP retrieval and HTML parsing; it cannot handle JavaScript rendering.PuppeteerCrawler: browser control with Chromium or Chrome.PlaywrightCrawler: browser control with the broader browser set described in Crawlee’s documentation.
Example: queue pages and save records
The quick start’s pattern is to add initial URLs, extract data in a request handler, enqueue discovered links, then export the dataset. Here is a compact PlaywrightCrawler example:
import { PlaywrightCrawler } from 'crawlee';
const crawler = new PlaywrightCrawler({
async requestHandler({ request, page, enqueueLinks, pushData }) {
const title = await page.title();
await pushData({ url: request.url, title });
await enqueueLinks({ selector: 'a', strategy: 'same-domain' });
},
});
await crawler.run(['https://example.com/']);
Crawlee’s quick start documents exporting local dataset records to JSON. Use its queue and storage structure when those capabilities solve a real operational need; for one isolated URL, the framework can add more setup than a direct request and parser. Crawlee quick start
5. Node.js fetch and Undici: the minimal HTTP baseline
Node’s built-in fetch is powered by Undici, according to the Node.js learning documentation. It is useful when you need to retrieve a response or call an endpoint and the returned data is already usable. It does not, by itself, parse HTML into convenient selectors, execute page JavaScript or manage a crawl. Node.js fetch guide
Runnable retrieval example
const response = await fetch('https://example.com/');
if (!response.ok) {
throw new Error(`HTTP ${response.status}: ${response.statusText}`);
}
const html = await response.text();
console.log(html.slice(0, 500));
Pair the response with Cheerio if you need to select elements from HTML. Prefer a documented data endpoint when one is available and appropriate; a browser is not automatically necessary just because the content is on a website.
6. Apify: hosted execution and managed scraper operations
Apify is a platform route for running scrapers as hosted Actors, rather than simply a local Node.js parsing or browser library. Its official JavaScript/TypeScript SDK creates Actors, while the platform supports running them at scale with monitoring and scheduling. Apify also offers ready-made scrapers, including browser-based options and HTTP-plus-Cheerio approaches. Apify SDK Apify scraping documentation
Quick wins for a faster PC:
Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Clear out junk files and repair common Windows errorsFree Scan →When the platform route fits
- You want to run jobs in a hosted environment instead of provisioning and operating the runtime yourself.
- Scheduling and monitoring are part of the problem, not afterthoughts.
- A ready-made Actor matches the site and data you need closely enough to avoid building the extraction workflow from scratch.
Apify’s SDK documentation identifies it as the official JavaScript/TypeScript Actor library and showed version 3.7 when checked on 2026-09-30 UTC. Hosted operations change the deployment and service decision; compare current platform documentation and terms for your specific workload. Apify SDK documentation
Decision guide: which one should you start with?
- Inspect the response first. If the data is in the returned HTML, try
fetchplus Cheerio. - Check whether browser execution is essential. If JavaScript produces the content or the workflow needs browser interaction, use Playwright or Puppeteer.
- Assess the crawl shape. If you need to discover and queue many pages and structure results, look at Crawlee rather than hand-building orchestration around one-page scripts.
- Decide where it should run. If you need hosted scheduling and monitoring, assess Apify as a platform.
- Check runtime compatibility. The current documentation requirements differ: Cheerio lists Node.js 22.19 or later; Playwright lists Node.js 22.x, 24.x or 26.x; Crawlee’s quick start says Node.js 16 or later. Confirm requirements for the versions you intend to install. These details were checked on 2026-09-30 UTC. Cheerio Playwright Crawlee
Performance, reliability and cost considerations
There is no supported universal speed ranking among these choices. An HTTP request plus parser avoids browser execution when the needed content is in the response; browser automation performs more work because it runs a page. But the right comparison depends on the target, page behavior and workload, and the sources here do not provide an independent benchmark across all six options.
Rank #4
- Keep the work proportional: use a parser for static markup rather than launching a browser without a need.
- Make waits specific: wait for the selector or condition that signals the data is ready; avoid assuming a fixed delay works for every page.
- Handle failures explicitly: check HTTP status for direct requests and ensure browsers close in a
finallyblock. - Account for operations: browser binaries and browser runtime are part of local deployment; hosted platforms shift some execution and operational concerns to a service.
- Do not generalize vendor comparisons: Apify’s documentation makes a specific claim that its Cheerio Scraper can be as much as 20 times faster than its full-browser Puppeteer solution for the static-content use case it describes. This is Apify’s own claim, not an independent benchmark across the products in this guide. Apify documentation
Troubleshooting common scraping problems
The extracted HTML has no target content
Likely cause: the browser generates the content after the initial response, so a plain HTTP fetch or Cheerio parse cannot see it. Fix: inspect the page’s behavior and use Playwright or Puppeteer if browser execution is needed, or use an appropriate endpoint if one is available.
A selector returns no matches
Likely cause: the selector does not match the actual response or rendered page, or the page has not reached the state you expect. Fix: inspect the markup you are parsing, verify the selector, and for browser automation wait for a relevant locator before extracting.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Playwright cannot find a browser
Likely cause: the required browser binaries were not installed in the environment. Fix: install the package’s required browsers with npx playwright install and make sure the deployment environment includes them. Playwright installation
Puppeteer launches unsuccessfully after using puppeteer-core
Likely cause: puppeteer-core does not download a browser. Fix: provide a compatible browser installation and configure the executable for your setup, or choose the full puppeteer package if its bundled compatible Chrome download suits your deployment. Puppeteer installation
The request fails or returns an unexpected status
Likely cause: the server response is not successful, the URL is wrong, or the endpoint does not return the expected page. Fix: log the status and response details, check the requested URL, and handle non-2xx responses before parsing. Do not assume that a successful network connection means the expected data was returned.
The crawl grows beyond the pages you intended
Likely cause: link discovery is enqueueing more URLs than the intended scope. Fix: constrain discovered links by domain and the selectors or URL patterns that define the task, and verify the queue behavior on a small scope before a larger run.
Or skip the browser setup
If your task is to capture a visual screenshot rather than extract structured fields, ScreenshotNeo is a website screenshot API and MCP server. A single GET request can return a PNG, JPEG, WebP or PDF. It is an alternative to try first when you need a screenshot without managing a browser locally: it accepts cookie or consent banners before capture and removes more than 60 known consent platforms, newsletter popups and chat widgets; bot checks, blank pages and failed loads are not billed; and its MCP server lets AI agents use take_screenshot, get_page_info and capture_pdf. Those cleanup steps can each be turned off.
cURL example, using the documented API call pattern and a target URL:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
See the ScreenshotNeo documentation for request options and formats. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. Sign up for ScreenshotNeo’s free plan.
Frequently asked questions
Can I scrape websites in Node.js without a headless browser?
Yes. Use fetch for retrieval and a parser such as Cheerio when the response already contains the content you need. A headless browser is needed when the task depends on browser execution or interaction, not for every website request.
Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchPC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Are Cheerio and Crawlee alternatives to one another?
Not exactly. Cheerio parses supplied markup. Crawlee is a crawling framework that offers a Cheerio-based crawler as well as browser-based crawler classes, so it can provide orchestration around different retrieval approaches.
Should I use Playwright or Puppeteer?
Choose based on the browser support, APIs, installation model and existing automation stack your team needs. Both offer browser control; consult their current documentation for version-specific compatibility rather than assuming one is always better.
Is Apify a Node.js scraper library?
Its JavaScript SDK is a library for creating Actors, but Apify is best understood here as a hosted platform and operations option, not as a direct equivalent to a local HTML parser.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.
The Tool Desk
Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →




