Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
The best JavaScript scraping library in 2026 depends on how the target page delivers data. Start with Node.js fetch and Cheerio when the required markup is already in the HTML response. Use Playwright or Puppeteer when JavaScript execution, browser APIs, or interaction is required. Choose Crawlee when you need one crawler interface that can switch between HTTP and browser-based work while handling queues, retries, and concurrency.
This guide gives you a practical decision path, current runtime requirements, working Node.js examples, and recovery steps for common failures. There is no single winner for every site.
Quick decision: which library should you use?
| Your page or project | Best starting point | Why |
|---|---|---|
| The data is present in the initial HTML response | Node.js fetch plus Cheerio |
Lowest overhead; parses HTML without launching a browser. |
| The page renders data with JavaScript or needs clicks, scrolling, login, or other interaction | Playwright | Full browser automation with Chromium, Firefox, and WebKit support documented by the project. |
| You already have a Chrome/Chromium Puppeteer codebase | Puppeteer | Familiar API and a sensible choice when WebKit is not required. |
| You are crawling many URLs and want shared queues, retries, and HTTP/browser modes | Crawlee | Provides CheerioCrawler, PlaywrightCrawler, and PuppeteerCrawler behind a common framework. |
Make the first decision by inspecting the response, not by guessing from how a page looks in a browser. A page that appears dynamic may still contain all useful data in server-rendered HTML; conversely, a simple-looking page may fetch its content only after scripts run.
Recommended Free Tools
First test: is the data in the initial HTML?
Fetch the response and inspect it
Use a normal HTTP request before installing browser automation:
#1 Best Overall
const response = await fetch('https://example.com/products');
if (!response.ok) throw new Error(`HTTP ${response.status}`);
const html = await response.text();
console.log(html.includes('product-card'));
Save the response and search for a distinctive text string, link, or element that you can see in the page. Also check whether the server returned a login page, consent wall, rate-limit message, or an empty shell.
Parse static markup with Cheerio
Cheerio creates a queryable HTML or XML structure with a jQuery-like API. It is not a browser: the project documentation explicitly says it provides “no visual rendering, no CSS, no loading of external resources, and no JavaScript execution.” If the required elements are absent from the response, changing selectors will not make them appear.
import * as cheerio from 'cheerio';
const response = await fetch('https://example.com/products');
if (!response.ok) throw new Error(`HTTP ${response.status}`);
const $ = cheerio.load(await response.text());
const products = $('.product-card').map((_, el) => ({
name: $(el).find('.name').text().trim(),
price: $(el).find('.price').text().trim(),
url: $(el).find('a').attr('href')
})).get();
console.log(products);
Use absolute URLs when following links, normalize whitespace, and treat missing attributes as normal input rather than assuming every card has identical markup.
Free tools Windows power users keep installed
One-click scans. No signup required.
When a real browser is required
Use Playwright for cross-browser coverage
Playwright is the strongest general default when your scraper must execute page JavaScript or interact with a site and you care about more than Chromium. Its documented browser engines include Chromium, Firefox, and WebKit. Cross-engine runs can reveal browser-specific behavior that a Chromium-only test misses.
import { chromium } from 'playwright';
const browser = await chromium.launch({ headless: true });
const page = await browser.newPage();
await page.goto('https://example.com/products', { waitUntil: 'domcontentloaded' });
await page.waitForSelector('.product-card');
const products = await page.locator('.product-card').evaluateAll(cards =>
cards.map(card => ({
name: card.querySelector('.name')?.textContent?.trim() ?? null,
price: card.querySelector('.price')?.textContent?.trim() ?? null,
url: card.querySelector('a')?.href ?? null
}))
);
console.log(products);
await browser.close();
Install the package and the browser binaries in the environment where the job runs. Pin versions in your project and verify the required browser installation in CI or a container image; a package install alone does not guarantee that an executable is available.
Use Puppeteer for Chromium-focused projects
Puppeteer remains reasonable when your existing code, team, and deployment are already built around it or when Chrome/Chromium is the only target. Do not choose it expecting WebKit coverage: the cited comparison material identifies WebKit as unsupported by Puppeteer.
import puppeteer from 'puppeteer';
const browser = await puppeteer.launch({ headless: true });
const page = await browser.newPage();
await page.goto('https://example.com/products', { waitUntil: 'domcontentloaded' });
await page.waitForSelector('.product-card');
const products = await page.$$eval('.product-card', cards =>
cards.map(card => ({
name: card.querySelector('.name')?.textContent?.trim() ?? null,
price: card.querySelector('.price')?.textContent?.trim() ?? null
}))
);
console.log(products);
await browser.close();
When Crawlee is the better choice
Crawlee is a crawling framework rather than only a parser or browser driver. Its documented crawler classes include CheerioCrawler for plain HTTP, PlaywrightCrawler for Playwright-powered pages, and PuppeteerCrawler for Puppeteer-powered pages. You can keep a shared request queue and handler style while moving a route from HTTP to a browser when its behavior demands it.
import { CheerioCrawler } from 'crawlee';
const crawler = new CheerioCrawler({
async requestHandler({ request, $, enqueueLinks, log }) {
const title = $('h1').first().text().trim();
log.info(`${request.url}: ${title}`);
await enqueueLinks({ selector: 'a.next', label: 'LIST' });
}
});
await crawler.run(['https://example.com/catalog']);
For a JavaScript-rendered route, use the corresponding Playwright or Puppeteer crawler and install that browser package separately. Crawlee is useful when crawl management is a first-class concern; it is not evidence of a universal size threshold at which every project must adopt a framework.
Current requirements and installation checks
Runtime requirements differ across these packages and can change. Cheerio’s current introduction states Node.js 22.19 or later. Crawlee’s version 3.18 quick start states a minimum Node.js version of 16 and says Playwright and Puppeteer are not bundled when their crawler classes are used. These statements are package-specific, not a single ecosystem requirement.
- Check the package’s current documentation and release notes at implementation time.
- Run
node --versionin the same image or host that will execute the scraper. - Install the crawler’s optional browser dependency explicitly when using PlaywrightCrawler or PuppeteerCrawler.
- In CI, perform a smoke test that launches the browser and loads a known page before running a large crawl.
- Lock dependency versions and review browser binary changes as part of upgrades.
Comparison by operating characteristic
| Characteristic | HTTP plus Cheerio | Playwright or Puppeteer | Crawlee |
|---|---|---|---|
| JavaScript execution | No | Yes, in a real browser | Yes through its browser crawler classes |
| External resources and visual rendering | Neither is provided by Cheerio | Browser loads resources according to your settings | Depends on selected crawler |
| Browser engines | None | Playwright documents Chromium, Firefox, and WebKit; Puppeteer is Chromium-focused and does not support WebKit | Select Playwright or Puppeteer crawler |
| Queues and crawl orchestration | Usually your code | Usually your code or an added framework | Built around crawler classes and request management |
| Typical setup cost | Node.js and a parser | Package plus browser binaries | Framework plus the selected parser or browser package |
A July 2026 secondary comparison recommends fetch plus Cheerio for static HTML, Playwright when JavaScript is needed, and Crawlee when queues, retries, and concurrency matter. Treat that as a practical heuristic, not a controlled performance benchmark.
Build a reliable scraper
Wait for the condition that proves readiness
Prefer a specific selector, response, or application state over a fixed sleep. A delay can be too short on a slow run and wasteful on a fast one. In Playwright, combine waitForSelector or locator assertions with a navigation timeout appropriate for your site.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Control concurrency and load
Start conservatively, observe response times and error rates, and increase concurrency only when the target and your own infrastructure remain healthy. Respect the site’s terms, robots guidance where applicable, authentication rules, and rate limits. Retries should be bounded and should distinguish transient network failures from deterministic HTTP errors.
Make extraction tolerant
Selectors break when classes are renamed or markup is A/B tested. Use stable attributes where available, validate required fields, record the source URL, and keep malformed records in a quarantine stream instead of silently dropping them.
Separate navigation from parsing
Store raw HTML or a controlled snapshot for debugging when policy permits. This lets you fix a selector without repeatedly loading the target and helps distinguish a changed page from a failed request.
Troubleshooting
Cheerio returns an empty list
Cause: the content is injected after load, the selector is wrong, or the response is a block/login page. Fix: save and inspect the raw response; if the data is absent, move that route to Playwright or Puppeteer and wait for the rendered element.
Browser launch fails in CI
Cause: browser binaries or operating-system libraries are missing, or the installed browser does not match the package. Fix: install the documented browser dependencies in the build image, run a launch smoke test, and keep package and browser versions aligned.
The page never reaches a network-idle state
Cause: analytics, WebSockets, polling, or advertisements keep connections open. Fix: wait for a meaningful selector or a specific response instead of global network idle, and block nonessential resources only when doing so does not remove data you need.
Selectors work locally but fail in production
Cause: different locale, viewport, authentication state, feature flag, or bot challenge. Fix: log URL, status, viewport, and a short HTML diagnostic; reproduce the production context locally and add explicit handling for challenge or login pages.
Requests are slow or intermittently fail
Cause: target throttling, overloaded browsers, DNS/TLS issues, or unbounded concurrency. Fix: cap concurrency, use bounded exponential backoff for transient errors, set realistic timeouts, and record which stage failed.
Or skip the browser setup
If your goal is a clean image or PDF rather than extracting fields into your own process, ScreenshotNeo provides a hosted screenshot API and MCP server. It accepts consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed; each response identifies the page verdict and billing status in X-Page-Verdict and X-Billed headers.
A single call can return PNG, JPEG, WebP, or PDF:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python:
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
See the complete parameter reference in the ScreenshotNeo documentation. Options include full-page and element capture, 12 device presets or custom viewports, dark mode, retina scale, PDF paper and page controls, custom CSS and JavaScript, clicks, selector or delay waits, request blocking, headers, cookies, user agents, authorization, timezone, geolocation, transparent backgrounds, resizing, chosen cache TTLs, signed image links, asynchronous jobs with signed webhooks, bulk capture of up to 100 URLs per call, usage reporting, and an OpenAPI specification. Its parameter names are compatible with those used by other screenshot APIs, which can simplify migration. An MCP server exposes take_screenshot, get_page_info, and capture_pdf to Claude, Cursor, and other MCP clients.
Every feature is available on every plan. The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 shots. Yearly billing provides two months free. Create a free ScreenshotNeo account to try it without a card.
Cost, performance, and maintenance trade-offs
Cheerio generally avoids browser startup and rendering work, so it is the economical first path when the response already contains the data. Browser automation consumes more CPU, memory, and wall-clock time because it launches a browser, executes scripts, and may load many resources. The evidence available here does not establish a controlled head-to-head benchmark, so choose based on page behavior and operational requirements rather than a claimed universal speed ratio.
Quick wins for a faster PC:
Scan for outdated or missing drivers - takes under a minuteDriver Scan →Clear out junk files and repair common Windows errorsFree Scan →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Crawlee adds framework overhead but can reduce custom orchestration code when queues, retries, and multiple crawler modes are important. Whichever stack you choose, measure your own pages: capture success rate, median and tail latency, browser memory, HTTP status distribution, retry count, and extracted-record validation failures.
What the language figures do—and do not—tell you
An Apify 2026 report excerpt attributes 71.7% of respondents to Python use and 17% to JavaScript preference. The excerpt does not provide sampling methods or respondent counts, so these are respondent figures, not market-wide language shares. They do not determine which library is correct for your site.
Frequently Asked Questions
Can Cheerio scrape a React or Vue site?
Only if the required data is present in the server response. Cheerio will not execute the React or Vue JavaScript that fills an empty HTML shell; use a browser crawler or find an authorized data endpoint instead.
Should I use Playwright or Puppeteer for a new project?
Use Playwright when Chromium, Firefox, and WebKit coverage or cross-browser testing matters. Puppeteer is a practical fit for an existing Puppeteer/Chromium codebase.
Do these 3 things before closing this tab:
1Fix the driver behind crashes, sound loss and screen glitches2Repair Windows errors before they cause bigger problems3Scan for outdated or missing drivers - takes under a minuteDoes Crawlee replace Cheerio and Playwright?
No. Crawlee coordinates crawler workflows and offers classes that use Cheerio, Playwright, or Puppeteer. You still select and install the underlying mode your route requires.
How do I avoid scraping a login or bot-challenge page as if it were data?
Validate status, title, URL, and one or more expected content markers before extraction; route challenge and authentication responses to explicit error handling.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

