Load your markup with cheerio.load(), then query the returned $ function with a selector such as p:contains("Hello"). Cheerio’s :contains() pseudo-class performs substring matching. If you need exact whole-text equality, select the candidate elements first and compare their extracted text in JavaScript.
Find text with :contains()
Install Cheerio in your Node.js project:
npm install cheerio
With ECMAScript modules, import Cheerio and load an HTML string:
import * as cheerio from 'cheerio';
const html = `
<ul>
<li>Apple</li>
<li>Green apple</li>
<li>Banana</li>
</ul>
`;
const $ = cheerio.load(html);
const matches = $('li:contains("Apple")');
console.log(matches.length); // 2
console.log(matches.map((_, element) => $(element).text()).get());
// [ 'Apple', 'Green apple' ]
The selector first limits the search to li elements, then keeps elements whose text contains the substring Apple. It is case-sensitive in this example, so matching and normalization should be chosen deliberately.
Substring matching versus exact text
Use :contains() for intentional substring matches
li:contains("an"), for example, can match any list item containing those characters. This is useful when a label may include additional words, such as finding both “Apple” and “Green apple”. Cheerio also supports most standard pseudo-classes and selector extensions such as :first, :last, and :eq(n) through its selector engine. Those positional extensions are Cheerio features, not selectors you should expect to work in a browser’s native CSS selector API.
Free tools Windows power users keep installed
One-click scans. No signup required.
#1 Best Overall
Compare extracted text for exact equality
The documented containment selector is not an exact-equality operator. Select a stable candidate set, extract each element’s text, and compare it:
const exact = $('li').filter((_, element) => {
return $(element).text().trim() === 'Apple';
});
console.log(exact.length); // 1
console.log(exact.first().text()); // Apple
This approach also lets you define your own policy. trim() ignores leading and trailing whitespace; you can normalize repeated whitespace or case when your application requires it. Do not silently normalize values when capitalization or spacing carries meaning.
Normalize safely when the page is inconsistent
const wanted = 'apple';
const normalized = $('li').filter((_, element) => {
const value = $(element).text().replace(/s+/g, ' ').trim().toLowerCase();
return value === wanted;
});
Keep the selector fixed and treat the text as data. Building a selector directly from untrusted input can produce parsing surprises when the value contains selector-special characters.
Load the right kind of input
HTML strings with load
cheerio.load(markup) is the normal choice when your program already has HTML as a string. Cheerio parses it as a document by default and may add html, head, and body wrappers.
The Tool Desk
Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →const $ = cheerio.load('<li>One</li>');
console.log($('body > li').length);
Fragments without document wrappers
If the input is only a fragment and those wrappers would be inconvenient, pass false as the third argument:
Rank #2
- HTML CSS Design and Build Web Sites
- Comes with secure packaging
- It can be a gift option
const $ = cheerio.load('<li>One</li>', null, false);
console.log($('li').length); // 1
Buffers and streams
Use loadBuffer for raw bytes when the encoding is unknown. Use stringStream for a stream that already contains decoded text, and decodeStream for a raw-byte stream whose encoding must be detected. The byte-oriented methods can sniff the encoding.
Fetching a URL with fromURL
fromURL is asynchronous and asks Cheerio to retrieve a URL before parsing it. Use it only when having Cheerio perform that fetch is appropriate for your application; otherwise fetch the response yourself so you can control headers, retries, authentication, limits, and error handling.
const $ = await cheerio.fromURL('https://example.com');
const title = $('title').text().trim();
console.log(title);
Regardless of the loader, Cheerio can only inspect the markup it receives.
Extract text correctly
.text() returns text content
.text() reads the selected nodes’ raw textContent. If a selected element contains inline script or style source, that source can appear in the result.
const $ = cheerio.load(`
<div class="card">
<span>Ready</span>
<script>const internal = true;</script>
</div>
`);
console.log($('.card').text());
Use innerText semantics when script and style text should be skipped
const visibleTextLike = $('.card').prop('innerText');
Cheerio documents innerText as skipping script and style text, but it still works from the parsed tree. Cheerio does not apply CSS, so text hidden by display: none or a hidden attribute can still be included. Treat either result as a structural extraction, not as a pixel-accurate view of what a browser displays.
Rank #3
Build selectors that survive markup changes
Text is often the right fallback, but a stable attribute or structural anchor is usually less fragile than a long chain of classes. Prefer a selector such as [data-testid="price"] or a known container, then apply text comparison inside that scope.
const price = $('[data-testid="product-card"]')
.find('.price')
.filter((_, element) => $(element).text().trim() === '$19.99');
Scope matters: the same label may occur in navigation, a footer, and the main content. Narrowing first avoids a false positive and makes the intention clear.
Recommended Free Tools
When a value is supplied by a user or another untrusted system, do not interpolate it into a selector string. Select a known element and compare its extracted value instead:
const wantedLabel = getLabelFromRequest();
const hit = $('button').filter((_, element) => {
return $(element).text().trim() === wantedLabel;
});
Why an apparently correct text selector returns nothing
Inspect the selection length first
const selection = $('p:contains("Hello")');
console.log('matches:', selection.length);
console.log('loaded HTML:', $.html());
Cheerio returns an empty selection when no nodes match. Calling .text() on that selection simply produces an empty string, which can hide the real problem.
The element is created by client-side JavaScript
Cheerio is not a web browser. It parses the HTML supplied to it; it does not execute scripts, render a page, load external resources, or run a client-side framework. If the browser creates the element after JavaScript runs, it will be absent from the source Cheerio receives. Obtain the rendered markup with a browser automation tool such as Puppeteer or Playwright, then pass that HTML to Cheerio.
Rank #4
- Brand: Wiley
- Set of 2 Volumes
- A handy two-book set that uniquely combines related technologies Highly visual format and accessible language makes these books highly effective learning tools Perfect for beginning web designers and front-end developers
The class or ID is dynamic
Generated class names and IDs can change between requests or builds. Anchor to a stable data- attribute, element relationship, heading, or carefully scoped text instead. Inspect the actual markup for the request that failed rather than relying on a copied browser snapshot.
The search is scoped incorrectly
A selector run against the wrong container cannot see the desired node. Start with a broad, known anchor, verify its length, and then add narrower conditions one at a time.
Whitespace or case differs
Exact comparisons fail when the source includes line breaks, non-breaking spaces, or different capitalization. Log a representative value, then choose explicit trimming, whitespace folding, or case handling that matches your data contract.
Security and output safety
Cheerio is a parser and DOM manipulation library, not an HTML sanitizer. Scripts and event-handler attributes in input can survive parsing and serialization. If you will render extracted or serialized markup in a browser, sanitize it with a dedicated sanitizer.
Text output can contain characters such as <, >, and quotation marks. Send extracted values to a text context or escape them for the output context you are targeting. Parsing untrusted HTML does not make that HTML safe to publish.
Quick wins for a faster PC:
Scan for outdated or missing drivers - takes under a minuteDriver Scan →Repair Windows errors before they cause bigger problemsFix Now →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Best Value
Complete Node.js example
This script demonstrates loading, substring matching, exact comparison, scoped selection, and a useful empty-result diagnostic:
import * as cheerio from 'cheerio';
const html = `
<main>
<h1>Products</h1>
<ul data-testid="product-list">
<li>Apple</li>
<li>Green apple</li>
<li>Banana</li>
</ul>
</main>
`;
const $ = cheerio.load(html);
const list = $('[data-testid="product-list"]');
if (list.length === 0) {
throw new Error('Product list was not found in the supplied HTML');
}
const containsApple = list.find('li:contains("Apple")');
console.log('Contains Apple:', containsApple.map((_, el) => $(el).text().trim()).get());
const exactlyApple = list.find('li').filter((_, el) => {
return $(el).text().trim() === 'Apple';
});
console.log('Exactly Apple:', exactlyApple.map((_, el) => $(el).text().trim()).get());
Run it with a project configured for ES modules (for example, a package.json containing "type": "module"), or translate the import to the CommonJS form documented by Cheerio:
const cheerio = require('cheerio');
Performance and reliability choices
- Parse once: call
loadonce for a document and reuse the returned$function for all queries. - Reduce the search space: select a stable container before filtering by text.
- Prefer deterministic comparisons: exact JavaScript comparisons make the normalization policy visible in code.
- Control network behavior: when fetching yourself, set timeouts, response-size limits, and error handling before passing HTML to Cheerio.
- Test real variants: include missing nodes, duplicate labels, extra whitespace, script/style descendants, and empty documents in automated tests.
Cheerio is deterministic for a given input string, buffer, or stream. Reliability problems usually come from the fetch, changing source markup, or client-side rendering that happened after the HTML was obtained—not from a browser viewport or CSS layout.
Or skip the browser setup
If your actual goal is a screenshot of a rendered page rather than text extraction from supplied HTML, ScreenshotNeo provides a single HTTP request. It accepts cookie and consent banners as a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each cleanup step can be disabled. Bot checks and CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and the response identifies the page verdict and billing status in X-Page-Verdict and X-Billed headers. Its MCP server exposes take_screenshot, get_page_info, and capture_pdf to Claude, Cursor, and other MCP clients.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
cURL
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
Node.js
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
See the ScreenshotNeo documentation for parameters. The free plan includes 1,000 screenshots each month with no card; paid plans start at $5 for 3,000 screenshots. Create a free ScreenshotNeo account.
Quick decision guide
| Need | Use | Reason |
|---|---|---|
| Any element containing a phrase | :contains("text") |
Concise substring matching |
| One element whose complete text equals a value | Candidate selector plus .filter() |
Explicit exact comparison and normalization |
| Raw bytes with unknown encoding | loadBuffer or decodeStream |
Encoding detection for byte input |
| Browser-created content | Browser automation, then Cheerio | Cheerio does not execute client JavaScript |
| Visual capture of a live page | ScreenshotNeo | Rendered capture with cleanup and usage-aware billing |
Frequently Asked Questions
Does Cheerio’s :contains() match exact text?
No. It is documented as a substring match. Select candidate elements and compare their extracted text in JavaScript for exact equality.
Can Cheerio find text added by React or another client-side framework?
Not from the original HTML alone. Cheerio does not run scripts or render a browser page; obtain rendered markup with browser automation first.
Why does .text() include unexpected JavaScript?
.text() reads raw textContent, which can include script and style source inside the selected element. Use .prop('innerText') when that behavior better fits your extraction.
PC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minuteQuick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

