Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

The best JavaScript scraping library in 2026 depends on how the target page delivers data. Start with Node.js fetch and Cheerio when the required markup is already in the HTML response. Use Playwright or Puppeteer when JavaScript execution, browser APIs, or interaction is required. Choose Crawlee when you need one crawler interface that can switch between HTTP and browser-based work while handling queues, retries, and concurrency.

This guide gives you a practical decision path, current runtime requirements, working Node.js examples, and recovery steps for common failures. There is no single winner for every site.

Quick decision: which library should you use?

Your page or project Best starting point Why
The data is present in the initial HTML response Node.js fetch plus Cheerio Lowest overhead; parses HTML without launching a browser.
The page renders data with JavaScript or needs clicks, scrolling, login, or other interaction Playwright Full browser automation with Chromium, Firefox, and WebKit support documented by the project.
You already have a Chrome/Chromium Puppeteer codebase Puppeteer Familiar API and a sensible choice when WebKit is not required.
You are crawling many URLs and want shared queues, retries, and HTTP/browser modes Crawlee Provides CheerioCrawler, PlaywrightCrawler, and PuppeteerCrawler behind a common framework.

Make the first decision by inspecting the response, not by guessing from how a page looks in a browser. A page that appears dynamic may still contain all useful data in server-rendered HTML; conversely, a simple-looking page may fetch its content only after scripts run.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

First test: is the data in the initial HTML?

Fetch the response and inspect it

Use a normal HTTP request before installing browser automation:

const response = await fetch('https://example.com/products');
if (!response.ok) throw new Error(`HTTP ${response.status}`);
const html = await response.text();
console.log(html.includes('product-card'));

Save the response and search for a distinctive text string, link, or element that you can see in the page. Also check whether the server returned a login page, consent wall, rate-limit message, or an empty shell.

Parse static markup with Cheerio

Cheerio creates a queryable HTML or XML structure with a jQuery-like API. It is not a browser: the project documentation explicitly says it provides “no visual rendering, no CSS, no loading of external resources, and no JavaScript execution.” If the required elements are absent from the response, changing selectors will not make them appear.

import * as cheerio from 'cheerio';

const response = await fetch('https://example.com/products');
if (!response.ok) throw new Error(`HTTP ${response.status}`);
const $ = cheerio.load(await response.text());

const products = $('.product-card').map((_, el) => ({
  name: $(el).find('.name').text().trim(),
  price: $(el).find('.price').text().trim(),
  url: $(el).find('a').attr('href')
})).get();
console.log(products);

Use absolute URLs when following links, normalize whitespace, and treat missing attributes as normal input rather than assuming every card has identical markup.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

When a real browser is required

Use Playwright for cross-browser coverage

Playwright is the strongest general default when your scraper must execute page JavaScript or interact with a site and you care about more than Chromium. Its documented browser engines include Chromium, Firefox, and WebKit. Cross-engine runs can reveal browser-specific behavior that a Chromium-only test misses.

import { chromium } from 'playwright';

const browser = await chromium.launch({ headless: true });
const page = await browser.newPage();
await page.goto('https://example.com/products', { waitUntil: 'domcontentloaded' });
await page.waitForSelector('.product-card');

const products = await page.locator('.product-card').evaluateAll(cards =>
  cards.map(card => ({
    name: card.querySelector('.name')?.textContent?.trim() ?? null,
    price: card.querySelector('.price')?.textContent?.trim() ?? null,
    url: card.querySelector('a')?.href ?? null
  }))
);
console.log(products);
await browser.close();

Install the package and the browser binaries in the environment where the job runs. Pin versions in your project and verify the required browser installation in CI or a container image; a package install alone does not guarantee that an executable is available.

Use Puppeteer for Chromium-focused projects

Puppeteer remains reasonable when your existing code, team, and deployment are already built around it or when Chrome/Chromium is the only target. Do not choose it expecting WebKit coverage: the cited comparison material identifies WebKit as unsupported by Puppeteer.

import puppeteer from 'puppeteer';

const browser = await puppeteer.launch({ headless: true });
const page = await browser.newPage();
await page.goto('https://example.com/products', { waitUntil: 'domcontentloaded' });
await page.waitForSelector('.product-card');
const products = await page.$$eval('.product-card', cards =>
  cards.map(card => ({
    name: card.querySelector('.name')?.textContent?.trim() ?? null,
    price: card.querySelector('.price')?.textContent?.trim() ?? null
  }))
);
console.log(products);
await browser.close();

When Crawlee is the better choice

Crawlee is a crawling framework rather than only a parser or browser driver. Its documented crawler classes include CheerioCrawler for plain HTTP, PlaywrightCrawler for Playwright-powered pages, and PuppeteerCrawler for Puppeteer-powered pages. You can keep a shared request queue and handler style while moving a route from HTTP to a browser when its behavior demands it.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
import { CheerioCrawler } from 'crawlee';

const crawler = new CheerioCrawler({
  async requestHandler({ request, $, enqueueLinks, log }) {
    const title = $('h1').first().text().trim();
    log.info(`${request.url}: ${title}`);
    await enqueueLinks({ selector: 'a.next', label: 'LIST' });
  }
});

await crawler.run(['https://example.com/catalog']);

For a JavaScript-rendered route, use the corresponding Playwright or Puppeteer crawler and install that browser package separately. Crawlee is useful when crawl management is a first-class concern; it is not evidence of a universal size threshold at which every project must adopt a framework.

Current requirements and installation checks

Runtime requirements differ across these packages and can change. Cheerio’s current introduction states Node.js 22.19 or later. Crawlee’s version 3.18 quick start states a minimum Node.js version of 16 and says Playwright and Puppeteer are not bundled when their crawler classes are used. These statements are package-specific, not a single ecosystem requirement.

  1. Check the package’s current documentation and release notes at implementation time.
  2. Run node --version in the same image or host that will execute the scraper.
  3. Install the crawler’s optional browser dependency explicitly when using PlaywrightCrawler or PuppeteerCrawler.
  4. In CI, perform a smoke test that launches the browser and loads a known page before running a large crawl.
  5. Lock dependency versions and review browser binary changes as part of upgrades.

Comparison by operating characteristic

Characteristic HTTP plus Cheerio Playwright or Puppeteer Crawlee
JavaScript execution No Yes, in a real browser Yes through its browser crawler classes
External resources and visual rendering Neither is provided by Cheerio Browser loads resources according to your settings Depends on selected crawler
Browser engines None Playwright documents Chromium, Firefox, and WebKit; Puppeteer is Chromium-focused and does not support WebKit Select Playwright or Puppeteer crawler
Queues and crawl orchestration Usually your code Usually your code or an added framework Built around crawler classes and request management
Typical setup cost Node.js and a parser Package plus browser binaries Framework plus the selected parser or browser package

A July 2026 secondary comparison recommends fetch plus Cheerio for static HTML, Playwright when JavaScript is needed, and Crawlee when queues, retries, and concurrency matter. Treat that as a practical heuristic, not a controlled performance benchmark.

Build a reliable scraper

Wait for the condition that proves readiness

Prefer a specific selector, response, or application state over a fixed sleep. A delay can be too short on a slow run and wasteful on a fast one. In Playwright, combine waitForSelector or locator assertions with a navigation timeout appropriate for your site.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Control concurrency and load

Start conservatively, observe response times and error rates, and increase concurrency only when the target and your own infrastructure remain healthy. Respect the site’s terms, robots guidance where applicable, authentication rules, and rate limits. Retries should be bounded and should distinguish transient network failures from deterministic HTTP errors.

Make extraction tolerant

Selectors break when classes are renamed or markup is A/B tested. Use stable attributes where available, validate required fields, record the source URL, and keep malformed records in a quarantine stream instead of silently dropping them.

Separate navigation from parsing

Store raw HTML or a controlled snapshot for debugging when policy permits. This lets you fix a selector without repeatedly loading the target and helps distinguish a changed page from a failed request.

Troubleshooting

Cheerio returns an empty list

Cause: the content is injected after load, the selector is wrong, or the response is a block/login page. Fix: save and inspect the raw response; if the data is absent, move that route to Playwright or Puppeteer and wait for the rendered element.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Browser launch fails in CI

Cause: browser binaries or operating-system libraries are missing, or the installed browser does not match the package. Fix: install the documented browser dependencies in the build image, run a launch smoke test, and keep package and browser versions aligned.

The page never reaches a network-idle state

Cause: analytics, WebSockets, polling, or advertisements keep connections open. Fix: wait for a meaningful selector or a specific response instead of global network idle, and block nonessential resources only when doing so does not remove data you need.

Selectors work locally but fail in production

Cause: different locale, viewport, authentication state, feature flag, or bot challenge. Fix: log URL, status, viewport, and a short HTML diagnostic; reproduce the production context locally and add explicit handling for challenge or login pages.

Requests are slow or intermittently fail

Cause: target throttling, overloaded browsers, DNS/TLS issues, or unbounded concurrency. Fix: cap concurrency, use bounded exponential backoff for transient errors, set realistic timeouts, and record which stage failed.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup

If your goal is a clean image or PDF rather than extracting fields into your own process, ScreenshotNeo provides a hosted screenshot API and MCP server. It accepts consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed; each response identifies the page verdict and billing status in X-Page-Verdict and X-Billed headers.

A single call can return PNG, JPEG, WebP, or PDF:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

See the complete parameter reference in the ScreenshotNeo documentation. Options include full-page and element capture, 12 device presets or custom viewports, dark mode, retina scale, PDF paper and page controls, custom CSS and JavaScript, clicks, selector or delay waits, request blocking, headers, cookies, user agents, authorization, timezone, geolocation, transparent backgrounds, resizing, chosen cache TTLs, signed image links, asynchronous jobs with signed webhooks, bulk capture of up to 100 URLs per call, usage reporting, and an OpenAPI specification. Its parameter names are compatible with those used by other screenshot APIs, which can simplify migration. An MCP server exposes take_screenshot, get_page_info, and capture_pdf to Claude, Cursor, and other MCP clients.

Every feature is available on every plan. The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 shots. Yearly billing provides two months free. Create a free ScreenshotNeo account to try it without a card.

Cost, performance, and maintenance trade-offs

Cheerio generally avoids browser startup and rendering work, so it is the economical first path when the response already contains the data. Browser automation consumes more CPU, memory, and wall-clock time because it launches a browser, executes scripts, and may load many resources. The evidence available here does not establish a controlled head-to-head benchmark, so choose based on page behavior and operational requirements rather than a claimed universal speed ratio.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Crawlee adds framework overhead but can reduce custom orchestration code when queues, retries, and multiple crawler modes are important. Whichever stack you choose, measure your own pages: capture success rate, median and tail latency, browser memory, HTTP status distribution, retry count, and extracted-record validation failures.

What the language figures do—and do not—tell you

An Apify 2026 report excerpt attributes 71.7% of respondents to Python use and 17% to JavaScript preference. The excerpt does not provide sampling methods or respondent counts, so these are respondent figures, not market-wide language shares. They do not determine which library is correct for your site.

Frequently Asked Questions

Can Cheerio scrape a React or Vue site?

Only if the required data is present in the server response. Cheerio will not execute the React or Vue JavaScript that fills an empty HTML shell; use a browser crawler or find an authorized data endpoint instead.

Should I use Playwright or Puppeteer for a new project?

Use Playwright when Chromium, Firefox, and WebKit coverage or cross-browser testing matters. Puppeteer is a practical fit for an existing Puppeteer/Chromium codebase.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Does Crawlee replace Cheerio and Playwright?

No. Crawlee coordinates crawler workflows and offers classes that use Cheerio, Playwright, or Puppeteer. You still select and install the underlying mode your route requires.

How do I avoid scraping a login or bot-challenge page as if it were data?

Validate status, title, URL, and one or more expected content markers before extraction; route challenge and authentication responses to explicit error handling.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.