Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

Use Cheerio for data already present in the HTTP response, Playwright when a real browser must execute JavaScript, Puppeteer when Chrome/Firefox automation is enough, and Crawlee when you need queues, retries, proxies, sessions, storage, and scaling. A tiered crawler that tries HTTP parsing first and opens a browser only for JavaScript-dependent pages is usually the most efficient design.

The decision in one table

Library Best fit Strengths Main limitations
Cheerio Static HTML/XML and pages whose fields are in the initial response Very low overhead; jQuery-like selectors and traversal No visual rendering, external-resource loading, or JavaScript execution; SPA content can be absent
Puppeteer Chrome/Firefox automation, screenshots, PDFs, UI interaction, and browser-state workflows High-level JavaScript API over CDP/WebDriver BiDi; headless by default Browser installation can fail when package-manager scripts are blocked; heavier runtime than HTTP parsing
Playwright Cross-browser scraping and interactions that need robust waiting Chromium, Firefox, WebKit, Chrome, and Edge; locators, auto-waiting, contexts, frames, tabs, and web-first assertions Matching browser binaries are required; updates can require another browser install; higher resource cost than Cheerio
Crawlee Production crawlers needing scheduling and operational controls CheerioCrawler, PuppeteerCrawler, and PlaywrightCrawler with queues, storage, scaling, proxies, sessions, retries, routing, Docker, and TypeScript support More framework complexity; browser crawlers require separate Playwright or Puppeteer installation

Choose by how the page produces data

Start with Cheerio for server-delivered markup

Cheerio parses HTML or XML. It is not a browser: it does not render visually, load external resources, or execute JavaScript. That is why it is fast and inexpensive, and also why a client-rendered single-page application may yield an empty result even though a human sees products, prices, or comments in a browser.

Inspect the raw response first. If the required text appears in “view source” or in the response body returned by an HTTP client, Cheerio is the simplest choice. Use CSS selectors, traversal, and normalization without paying for a browser process.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Escalate to a browser for client-rendered sites

Use Playwright or Puppeteer when JavaScript builds the DOM, when a click or form submission is required, when content is inside frames or tabs, when the site varies by browser engine, or when you need screenshots or PDFs. A browser also preserves cookies, local storage, and other state that ordinary HTTP parsing does not create.

Use Crawlee when operations become the problem

Crawlee is a framework rather than a competing parser. Its CheerioCrawler handles HTTP pages, while PuppeteerCrawler and PlaywrightCrawler handle browser pages through a common architecture. Persistent request queues, storage, retries, routing, sessions, proxy rotation, resource-based scaling, Docker support, and deployment conventions matter once a script becomes a continuing crawl.

Cheerio: the low-overhead path

Install the HTTP client and parser:

npm install axios cheerio

This complete Node.js example extracts article titles from the initial response and reports missing selectors instead of silently returning bad data:

const axios = require('axios');
const cheerio = require('cheerio');

async function scrapeStatic(url) {
  const response = await axios.get(url, {
    headers: { 'User-Agent': 'Mozilla/5.0 (compatible; ResearchBot/1.0)' },
    timeout: 30000,
    validateStatus: status => status >= 200 && status < 400
  });
  const $ = cheerio.load(response.data);
  const titles = $('article h2, article h3').map((_, el) => $(el).text().trim()).get();
  if (titles.length === 0) {
    throw new Error('No matching titles in the initial HTML; test a browser crawler.');
  }
  return titles;
}

scrapeStatic('https://example.com/news')
  .then(items => console.log(items))
  .catch(error => console.error(error.message));

Do not “fix” an empty result by adding arbitrary delays to Cheerio; delays cannot make a non-browser parser execute JavaScript. Check the response, identify the endpoint or rendered route, and move only that workload to a browser.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Playwright and Puppeteer: real browser automation

Playwright setup and a robust extraction

Playwright supports Chromium, Firefox, WebKit, Chrome, and Edge. Its locators wait for elements to be actionable, reducing hand-written sleep calls and race conditions.

npm install playwright
npx playwright install chromium
const { chromium } = require('playwright');

(async () => {
  const browser = await chromium.launch({ headless: true });
  const context = await browser.newContext({
    viewport: { width: 1440, height: 900 },
    locale: 'en-US',
    timezoneId: 'UTC'
  });
  const page = await context.newPage();
  await page.goto('https://example.com/catalog', {
    waitUntil: 'domcontentloaded',
    timeout: 60000
  });
  const cards = page.locator('[data-product-card]');
  await cards.first().waitFor({ state: 'visible', timeout: 30000 });
  const products = await cards.evaluateAll(nodes => nodes.map(node => ({
    name: node.querySelector('.name')?.textContent.trim() ?? null,
    price: node.querySelector('.price')?.textContent.trim() ?? null
  })));
  console.log(products);
  await browser.close();
})().catch(error => { console.error(error); process.exitCode = 1; });

Prefer a locator tied to a stable role, label, or data attribute. Use a specific readiness condition such as a selector, a network-idle rule, or a response event rather than a large fixed delay. Create a fresh browser context per logical session when cookies or authentication must not leak between accounts.

Rank #2
Sale
HTML and CSS: Design and Build Websites
  • HTML CSS Design and Build Web Sites
  • Comes with secure packaging
  • It can be a gift option

Puppeteer setup and when it is enough

Puppeteer controls Chrome and Firefox and is headless by default. It is a good fit when your deployment standardizes on those engines and its API ecosystem already matches your automation. The equivalent minimal flow is:

npm install puppeteer
const puppeteer = require('puppeteer');

(async () => {
  const browser = await puppeteer.launch({ headless: true });
  const page = await browser.newPage();
  await page.goto('https://example.com/catalog', {
    waitUntil: 'networkidle2',
    timeout: 60000
  });
  const products = await page.$$eval('[data-product-card]', nodes => nodes.map(node => ({
    name: node.querySelector('.name')?.textContent.trim() ?? null,
    price: node.querySelector('.price')?.textContent.trim() ?? null
  })));
  console.log(products);
  await browser.close();
})().catch(error => { console.error(error); process.exitCode = 1; });

Choose Playwright instead when WebKit coverage, multiple browser engines, or its locator and auto-waiting model is a central requirement. Both browser approaches consume substantially more CPU, memory, startup time, and maintenance effort than Cheerio.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Crawlee for a production crawler

Crawlee version 3.18 provides a common foundation for HTTP and browser crawlers. Install the core package and the crawler engines you intend to use:

npm install crawlee playwright
npx playwright install chromium
const { PlaywrightCrawler } = require('crawlee');

const crawler = new PlaywrightCrawler({
  maxRequestsPerCrawl: 100,
  maxConcurrency: 5,
  requestHandler: async ({ page, request, log }) => {
    await page.locator('[data-product-card]').first().waitFor({ state: 'visible' });
    const products = await page.$$eval('[data-product-card]', nodes => nodes.map(node => ({
      name: node.querySelector('.name')?.textContent.trim() ?? null,
      price: node.querySelector('.price')?.textContent.trim() ?? null
    })));
    log.info(`Collected ${products.length} products`, { url: request.url });
  },
  failedRequestHandler: async ({ request, log }) => {
    log.error(`Failed after retries: ${request.url}`);
  }
});

crawler.run(['https://example.com/catalog']).catch(console.error);

Use CheerioCrawler for the majority of URLs and route only JavaScript-dependent requests to PlaywrightCrawler or PuppeteerCrawler. That design preserves Crawlee’s queues, retries, sessions, storage, and proxy controls without paying browser costs for every page.

A practical selection workflow

  1. Fetch one representative URL. Save the response body and inspect whether the required fields exist before scripts run.
  2. Parse with Cheerio if they do. Add explicit status, timeout, and selector checks so a template change becomes an alert.
  3. Identify the browser requirement if they do not. Determine whether the page needs JavaScript, a click, authentication, a frame, a browser-specific rendering path, or a screenshot.
  4. Use Playwright for cross-engine behavior and reliable waits. Use Puppeteer when Chrome/Firefox control is sufficient and its API is already established in your codebase.
  5. Add Crawlee when reliability is a system concern. Queues, persistence, retries, sessions, proxies, and scaling are reasons to adopt the framework, not merely the presence of one dynamic page.
  6. Measure and cap concurrency. Increase workers only while CPU, memory, target response times, and error rates remain acceptable.

Installation and runtime failures

Browser executable is missing

Playwright versions require matching browser binaries. After upgrading the package, rerun the appropriate browser installation command. In CI, install browsers in the image build and cache that layer when possible.

Puppeteer starts but cannot launch

If a package manager blocks install scripts, Puppeteer may not download its browser and will fail at runtime. Allow the install script in the controlled build, or configure an explicitly managed browser executable and verify that the binary is present before starting workers.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Selectors time out

A timeout can mean the selector is wrong, the page is still gated, the content is in an iframe, or the request failed. Capture the URL, status, console messages, screenshot, and HTML on failure. Replace broad sleeps with a selector or response condition that represents the data you actually need.

Results are empty or duplicated

Empty results usually indicate that Cheerio received a shell page or that the browser waited for the wrong event. Duplicates often come from pagination, infinite scroll, retries, or multiple tabs. Record the canonical URL and a stable item identifier, and make writes idempotent.

Performance, reliability, and cost trade-offs

  • Latency and resource use: Cheerio avoids browser startup and external-resource work. Browser crawlers pay for processes, pages, JavaScript execution, fonts, images, and cleanup.
  • Concurrency: HTTP parsing can usually run at much higher concurrency; browsers need lower limits sized to available memory. Watch resident memory and file descriptors, not only request counts.
  • Reliability: Set navigation and selector timeouts, retry transient failures with a limit, preserve failed URLs, and record the reason for each retry. A retry cannot repair a permanent consent wall or a changed selector.
  • Freshness: Cache immutable or slowly changing responses, but avoid serving stale authenticated data. Browser contexts should be closed even when extraction throws.
  • Observability: Log URL, status, elapsed time, crawler type, browser engine, attempt number, and a failure artifact. This lets you see whether escalation to a browser is actually justified.

Robots.txt, permissions, and legal boundaries

RFC 9309 (IETF, September 2022) defines robots.txt processing as a requested protocol and states that its rules are not access authorization. A parseable robots.txt should be followed after successful retrieval; unavailable and unreachable files have different handling, and cached rules generally should not be used for more than 24 hours unless the file is unreachable.

Robots.txt is only one input. Review the target’s terms, authentication boundaries, privacy and copyright obligations, rate limits, and applicable law. Do not bypass a login, CAPTCHA, paywall, or technical access control merely because a URL is discoverable. Obtain permission for authenticated or high-volume collection and provide a contact path for removal requests where appropriate.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #4
Sale
Web Design with HTML, CSS, JavaScript and jQuery Set
  • Brand: Wiley
  • Set of 2 Volumes
  • A handy two-book set that uniquely combines related technologies Highly visual format and accessible language makes these books highly effective learning tools Perfect for beginning web designers and front-end developers
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

When the deliverable is a screenshot, not extracted data

If your browser workflow exists mainly to produce a clean page image or PDF, ScreenshotNeo can replace the browser setup with one API request. It accepts consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each cleanup step can be disabled. Only clean shots are billed: bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are free, and response headers identify the page verdict and billing status.

Or skip the browser setup

See the ScreenshotNeo API documentation for the full parameter list. This cURL request returns a WebP image:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

The same request in Python:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

And in Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

ScreenshotNeo has 63 options, including full-page capture with lazy images loaded, CSS-selector element capture, dark mode, 12 device presets plus custom viewports, retina scale, PDF paper size/margins/landscape/page ranges, HTML/CSS-to-image, custom CSS and JavaScript, pre-capture clicks, hidden selectors, selector/delay/network-idle waits, ad/tracker/request/resource blocking, headers/cookies/user agent/Authorization, timezone and geolocation, transparent backgrounds, resizing, configurable-TTL caching, signed image links, asynchronous jobs with signed webhooks, bulk capture of up to 100 URLs per call, a usage API, and an OpenAPI specification. Common parameter names used by other screenshot APIs also work, which eases migration.

An MCP server exposes take_screenshot, get_page_info, and capture_pdf to Claude, Cursor, and other MCP clients, so AI agents can perform captures directly. Plans include 1,000 shots per month free with no card; paid tiers are Starter $5 for 3,000, Growth $15 for 15,000, Pro $39 for 60,000, Scale $99 for 250,000, and Business $249 for 1,000,000. Yearly billing gives two months free, and every feature is included on every plan. Sign up free for 1,000 screenshots a month with no card.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

FAQ

Can I use more than one library in the same crawl?

Yes. A router can send ordinary responses to Cheerio and promote only JavaScript-dependent requests to Playwright or Puppeteer; Crawlee supplies the shared queue and retry layer when needed.

Why did a browser upgrade break a previously working deployment?

The automation package and its browser binary must remain compatible. Reinstall the binaries after package updates and test the image in CI before releasing it.

Is robots.txt permission to scrape?

No. It expresses crawler preferences under RFC 9309, not authorization. Terms, authentication rules, privacy, copyright, rate limits, and local law still apply.

Frequently Asked Questions

Can I use more than one library in the same crawl?

Yes. Route ordinary responses to Cheerio and promote only JavaScript-dependent requests to Playwright or Puppeteer; Crawlee can provide a shared queue and retry layer.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Why did a browser upgrade break a previously working deployment?

The automation package and its browser binary must remain compatible. Reinstall the binaries after package updates and test the image in CI before releasing it.

Is robots.txt permission to scrape?

No. Under RFC 9309 it expresses crawler preferences, not authorization. Terms, authentication rules, privacy, copyright, rate limits, and applicable law still apply.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.