October DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsWindows FixRecommendedWindows errors stealing your time? Find the fix fastScan stability, cleanup and performance issues.Fix NowOctober DealsAmazon USDeal season is back - check today's better picksAmazon US: current deals, useful picks and tech finds.See Picks×
Skip to content
Cheerio

Web Scraping with Cheerio in 2026: A Practical Node.js Guide

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Cheerio is the right tool when the data you need is already in a page’s HTML response. It parses HTML or XML with a fast, jQuery-like API, but it does not open a browser, execute JavaScript, click controls, or wait for client-rendered content. The key decision is therefore simple: inspect the response first. If the desired nodes are present, Cheerio can extract them efficiently; if the response is only an application shell, use browser automation such as Puppeteer or Playwright instead.

This guide covers installation, loading choices, selectors, encoding, parser configuration, URL fetching, failure diagnosis, security, and a browser-free alternative for producing page screenshots.

What Cheerio can—and cannot—scrape

Cheerio parses markup supplied by your application. Its API resembles jQuery, so you can select elements, traverse the tree, read text and attributes, and manipulate the parsed document. It does not render a page or execute its scripts; as Cheerio’s documentation puts it, “Cheerio is not a web browser.”

Use Cheerio when

  • The HTTP response contains the titles, links, prices, tables or metadata you need.
  • You want low overhead and deterministic parsing without launching Chromium.
  • You can fetch pages with Node.js and then process their markup.

Use a browser when

  • The initial response is an empty app shell and JavaScript inserts the data later.
  • You must execute scripts, interact with controls, handle scrolling, or reproduce a logged-in browser session.
  • The site requires browser-only APIs or visual state before the data appears.

When in doubt, save the raw response and search it for a distinctive value visible in the browser. If it is absent, changing CSS selectors will not solve the problem.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Install Cheerio and verify your runtime

Install the package in a Node.js project:

npm install cheerio

The package listing showed Cheerio 1.2.0 on 2026-09-29, while the introductory documentation stated Node.js 22.19 or later. Both values can change, so check the current package metadata and documentation before pinning a production environment.

How do I scrape a website with Cheerio?

The basic workflow is: obtain HTML, load it, select nodes, and extract fields. This complete example fetches a page with Node’s built-in fetch and collects article titles and links.

import * as cheerio from 'cheerio';

const target = 'https://example.com/news';
const response = await fetch(target, {
  headers: { 'user-agent': 'MyResearchBot/1.0' }
});

if (!response.ok) {
  throw new Error(`HTTP ${response.status} for ${target}`);
}

const html = await response.text();
const $ = cheerio.load(html);
const articles = [];

$('article').each((_, element) => {
  const title = $(element).find('h2, h3').first().text().trim();
  const href = $(element).find('a').first().attr('href');
  if (title) articles.push({ title, href });
});

console.log(articles);

text() returns the combined text of a selection. attr('href') reads an attribute; for an empty selection it normally returns undefined rather than throwing. Check the selection before assuming a match:

const cards = $('.product-card');
if (cards.length === 0) {
  console.error('No product cards. Inspect the response HTML and selector.');
}

Choose the loader that matches your input

Cheerio provides five practical entry points. Selecting the right one avoids encoding and memory surprises.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Method Input Best use
load String Already-decoded HTML or XML text
loadBuffer Buffer Bytes when the character encoding is uncertain
stringStream Decoded text stream Incremental parsing when your stream is already decoded
decodeStream Raw byte stream Streaming bytes while Cheerio sniffs encoding
fromURL URL Let Cheerio fetch and parse a page

Stream and URL loaders depend on Node.js APIs and are not included in the browser build. Prefer loadBuffer or decodeStream when decoding is not certain; decoding bytes as UTF-8 before parsing can corrupt non-ASCII text.

How do I load a URL with Cheerio?

fromURL combines fetching and parsing:

import * as cheerio from 'cheerio';

const $ = await cheerio.fromURL('https://example.com/catalog');
const names = $('h2.product-name').map((_, el) => $(el).text().trim()).get();
console.log(names);

Documented behavior matters in production:

  • Redirects are followed up to five times.
  • Non-2xx responses reject with an Undici response error.
  • Responses whose content type is not HTML or XML are rejected.
  • XML mode is selected from the response content type.
  • Encoding comes from a declared content-type charset when present; otherwise bytes are sniffed.
  • baseURI reflects the final URL after redirects.

Custom request options

When supplying requestOptions, explicitly include method; omitting it causes the call to fail. A supplied headers object replaces the default Accept header, so include the headers you need rather than assuming they are merged.

const $ = await cheerio.fromURL('https://example.com/data', {
  requestOptions: {
    method: 'GET',
    headers: {
      accept: 'text/html,application/xhtml+xml',
      'user-agent': 'CatalogBot/1.0'
    }
  }
});

If you need specialized retries, authentication, proxy handling, or response logging, fetch with your own HTTP client and pass the resulting string or buffer to load or loadBuffer.

Selectors, attributes and reliable extraction

Cheerio supports familiar CSS and jQuery-style traversal: descendant selectors, classes, IDs, attribute selectors, .find(), .first(), .last(), .eq(), .map() and .each(). Build selectors around stable semantics such as a data attribute or meaningful element rather than generated class names.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
const rows = $('table[data-testid="orders"] tbody tr').map((_, row) => {
  const cells = $(row).find('td');
  return {
    id: $(cells[0]).text().trim(),
    status: $(cells[1]).text().trim(),
    url: $(row).find('a').attr('href') ?? null
  };
}).get();

Normalize whitespace and handle optional fields explicitly. For relative links, resolve against the page URL with the standard URL class:

const absolute = href ? new URL(href, 'https://example.com/catalog').href : null;

When extraction unexpectedly returns an empty string or undefined, log a small slice of the response, count the selection, and verify that the server’s markup—not the browser’s post-rendered DOM—contains the target.

parse5 or htmlparser2?

Cheerio uses parse5 by default for HTML and htmlparser2 by default for XML. Parser choice affects standards fidelity, tolerance of malformed markup, speed and memory use.

Choice Characteristics Consider it when
parse5 Browser-oriented HTML parsing and standards fidelity You need results closest to how a browser interprets HTML
htmlparser2 Faster, lower-memory and more forgiving of malformed markup Throughput or imperfect source markup matters more than browser equivalence

Do not switch parsers merely because a selector failed. First confirm that the response contains the node; parser changes cannot create client-rendered content.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Why does my Cheerio selector return nothing?

  1. Client rendering: the server returned an app shell. Inspect the raw response and use Puppeteer or Playwright if JavaScript is required.
  2. Wrong selector: check the actual tag, class, nesting and attribute names in the response.
  3. Wrong document: redirects, consent pages, bot checks or an error page may have replaced the expected content.
  4. Encoding damage: use loadBuffer or decodeStream when bytes are not known to be UTF-8.
  5. Namespace or XML differences: confirm the content type and whether XML mode was selected.
  6. Timing assumptions: Cheerio parses once; it does not wait for network requests or timers.

Capture status, final URL, content type and response length in logs. Those four values often reveal the failure faster than changing selectors.

Security, limits and responsible collection

Cheerio parses markup and does not execute scripts, but it is not a sanitizer. Limit the size of untrusted input, validate sources and fields, and sanitize markup before rendering any extracted or user-supplied HTML in a browser. Treat extracted text as data and encode it at the output boundary.

There is no universal legal answer for scraping. Permission depends on the target site, terms, access controls, jurisdiction, the data involved and your intended use. Review applicable site policies and obtain qualified advice for consequential projects; do not treat a parser library as authorization.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Performance and reliability practices

  • Fetch only the pages and fields you need; avoid loading browser automation for static documents.
  • Set network timeouts and bound response sizes before parsing.
  • Reuse a controlled HTTP client when you need retries, rate limits or observability.
  • Prefer streaming loaders for large responses, while recognizing that selectors requiring the complete tree still need sufficient memory.
  • Cache responsibly and respect robots directives, terms and server capacity.
  • Store the source URL, final URL, status, content type and parser mode alongside results for debugging.

Or skip the browser setup

If your immediate need is a clean visual capture rather than structured DOM fields, ScreenshotNeo provides a single-call website screenshot API and MCP server. It accepts cookie and consent banners as a visitor, then removes more than 60 known consent platforms, newsletter popups and chat widgets before capture; each step can be disabled. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads and cache hits are not billed, and response headers report the page verdict and billing status.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

One request returns PNG, JPEG, WebP or PDF:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

See the ScreenshotNeo documentation for all options, including full-page and element captures, device presets, custom CSS and JavaScript, waits, blocking rules, cookies and headers, PDF settings, caching, signed links, asynchronous webhooks, bulk capture and the usage API. Its MCP server exposes take_screenshot, get_page_info and capture_pdf to Claude, Cursor and other MCP clients.

There is a free allowance of 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.

Frequently Asked Questions

Can Cheerio scrape a JavaScript-rendered page?

Not when the desired content exists only after scripts run. Use Puppeteer or Playwright for rendering and interaction, or locate an underlying server/API response that already contains the data.

Which Cheerio loader should I use for unknown character encoding?

Use loadBuffer for a complete byte buffer or decodeStream for a raw byte stream. They let Cheerio inspect encoding instead of forcing an early UTF-8 decode.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Does Cheerio sanitize HTML?

No. Limit input size and sanitize untrusted markup before rendering it.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Read next

Recommended PC Tool
Recommended PC Tool
Windows Errors? Fix Them Before They SpreadFree repair scan
Outdated Drivers Are Slowing You DownFree scan - exact matches

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.