October DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsClean PCRecommendedOne scan can reveal what keeps slowing WindowsLook for cleanup and repair opportunities.Run ScanOctober DealsAmazon USDeal season is back - check today's better picksAmazon US: current deals, useful picks and tech finds.See Picks×
Skip to content
Cheerio

What Is Cheerio in JavaScript? A Practical Guide to Parsing HTML

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Cheerio is a JavaScript library for parsing HTML or XML that you already have, then selecting, reading, and changing elements with a jQuery-like API. It is useful for tasks such as extracting links or text from markup and transforming a document. It is not a browser: it does not visually render a page or execute its JavaScript, so it cannot by itself reveal content that a site inserts only after the page loads.

What Cheerio does

Cheerio takes markup as input and exposes methods for working with the resulting document structure. A typical workflow is to obtain HTML or XML, load it into Cheerio, select the elements you need, and read or modify their contents. If you change the document, you can serialize it back into markup.

Its API will feel familiar to developers who have used jQuery: select elements with CSS-style selectors, traverse the document, inspect text or attributes, and make changes. The important difference is where the input comes from. In a browser, jQuery works with a live page; Cheerio starts with markup supplied to your Node.js program.

For example, given the string <h2 class="title">Hello world</h2>, Cheerio can find the heading and return its text. It is working with the parsed markup, not displaying the page to a user.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What Cheerio cannot do

Cheerio does not run a web page as a browser would. It does not execute page JavaScript, apply CSS for visual rendering, or load external resources such as images and scripts. A page may initially return a nearly empty HTML shell and then use client-side JavaScript to fetch and display its content. If the data is absent from the markup you give Cheerio, Cheerio alone cannot produce it.

This is the key decision point: if the data is already in the HTML or XML, Cheerio may be a good fit; if you need the page to execute scripts or need browser automation, use a browser-oriented tool instead. The Cheerio introduction names Puppeteer and Playwright for browser automation, and jsdom as a DOM-emulation option. Those are different approaches, not features built into Cheerio.

Install and use Cheerio

The basic example below assumes a Node.js project with npm available. It uses an ES module import, loads an HTML string, selects a heading, and serializes the document.

  1. In your project directory, install the package: npm install cheerio.

    Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  2. Save the following as example.mjs:

    import * as cheerio from 'cheerio';
    
    const markup = '<main><h2 class="title">Hello world</h2></main>';
    const $ = cheerio.load(markup);
    
    const heading = $('h2.title').text();
    console.log(heading); // Hello world
    
    console.log($.html());
  3. Run it with node example.mjs. The first output is the selected heading’s text; the second is serialized HTML for the loaded document.

The project documentation also shows CommonJS usage with require. Choose the import style that matches your project’s module configuration. The example uses a string already in memory; it does not fetch a website.

Choose the right way to load input

Cheerio offers different loading methods depending on what you have. The documentation describes these routes:

Input you have Loading route When it fits
Markup as a string load Use when another part of your program has already obtained or constructed the text.
Raw bytes, encoding not known loadBuffer Use the byte-oriented loader when you want Cheerio to perform encoding sniffing.
A stream of decoded text stringStream Use when your input is already arriving as text through a stream.
A stream of raw bytes decodeStream Use when input arrives as bytes and encoding sniffing is needed.
A URL to load fromURL Use when you want Cheerio to load a URL; it refuses responses whose content type is neither HTML nor XML.

These methods address how markup reaches Cheerio. They do not turn the library into a browser or cause client-side scripts to run. In particular, fromURL is a URL-loading convenience, not a solution for pages whose useful data appears only after browser-side JavaScript executes.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Understand the parser choice

Cheerio’s documented defaults differ by markup type. For HTML, it uses parse5 by default, following HTML parsing rules. For XML, htmlparser2 is the default. The project describes htmlparser2 as faster, lower-memory, and more forgiving of malformed markup; it can also be selected for HTML when those properties are desirable or parse5’s browser-oriented parsing is not suitable.

Those are qualitative descriptions from the project documentation, not a guarantee about speed or memory use for a particular input. Parser choice can affect how imperfect markup is interpreted. If your input is malformed, or if exact tree structure matters, test the parser behavior against representative documents rather than assuming different parsers will produce identical results.

Common Cheerio use cases

For each use case, the practical prerequisite is the same: your program must have the relevant markup. If a website requires JavaScript execution, a login flow, or browser interaction to expose the content, plan for a browser-based step before parsing or choose a tool designed to automate a browser.

Cheerio versus browser tools and DOM emulation

Need Cheerio Other approach named by Cheerio
Parse and select from markup already available Designed for this task, with a jQuery-like API. A browser tool may be unnecessary if no browser behavior is needed.
Execute client-side page JavaScript or automate a browser Does not do this. Puppeteer or Playwright are the browser-automation options named in Cheerio’s introduction.
Use a DOM-emulation approach Cheerio provides its own markup-processing API rather than a full browser environment. jsdom is the DOM-emulation option named in Cheerio’s introduction.
Choose parsing behavior for HTML Uses parse5 by default; htmlparser2 can be selected when its documented characteristics better fit the input. Consider the parser behavior required by your document, not a presumed universal speed advantage.

Cheerio and browser automation can also be used in sequence: a browser-oriented tool can obtain a rendered page state, while a markup parser can be useful for subsequent extraction or transformation when you have the resulting markup. Whether that division is worthwhile depends on where the data appears and whether browser behavior is required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Troubleshooting common problems

A selector returns no result

First inspect the exact markup passed to Cheerio. The element may not be present in the original response, its class or tag may differ, or the desired content may be added later by page JavaScript. Cheerio does not execute that JavaScript. If the element is present, check that the selector matches the actual structure and spelling in the input.

The page looks different from the browser

Cheerio does not render CSS, load external resources, or execute scripts. A browser may show content or layout that cannot be inferred from the initial HTML alone. Use a browser-oriented approach when your task depends on the rendered or interactive page, rather than expecting Cheerio to reproduce it.

Loading a URL is rejected

Cheerio’s documentation says fromURL refuses a response whose content type is neither HTML nor XML. Confirm that the URL returns a supported content type and that you are actually trying to load markup. A URL that points to another kind of resource is not suitable input for that method.

Malformed markup produces an unexpected structure

Parser behavior matters. Cheerio defaults to parse5 for HTML and htmlparser2 for XML; htmlparser2 can also be chosen for HTML. Review the parser configuration and try the parser whose handling best matches your markup. Do not treat a parser’s tolerance of malformed input as proof that the resulting tree matches your intended structure.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Input encoding is unclear

If you have raw bytes rather than a decoded string, use the documented byte-oriented loading path, such as loadBuffer or decodeStream. The loading guide says these byte-oriented methods perform encoding sniffing. If you already converted bytes to text incorrectly upstream, loading that string cannot restore lost or corrupted characters.

Or skip the browser setup

Cheerio is for parsing markup; it is not a screenshot tool or a browser renderer. If your actual goal is to capture a clean screenshot or PDF of a website, ScreenshotNeo is a separate option: a single GET request can return PNG, JPEG, WebP, or PDF. For example, this cURL request saves a WebP screenshot of Stripe; replace the target URL as needed. See the ScreenshotNeo API documentation for parameters and response details.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Sign up for ScreenshotNeo’s free plan to try 1,000 screenshots a month without a card.

Practical decision rule

Use Cheerio when you have markup and need a convenient way to inspect or transform it. Reach for browser automation when page scripts or browser interactions must happen first, and use a screenshot service when the deliverable is an image or PDF rather than extracted markup. Making that distinction before writing selectors prevents the most common category error: asking a parser to create content that only a running page can supply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Read next

Recommended PC Tool
Recommended PC Tool
Windows Errors? Fix Them Before They SpreadFree repair scan
Outdated Drivers Are Slowing You DownFree scan - exact matches

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.