Load your markup with cheerio.load(), then query the returned $ function with a selector such as p:contains("Hello"). Cheerio’s :contains() pseudo-class performs substring matching. If you need exact whole-text equality, select the candidate elements first and compare their extracted text in JavaScript.
Find text with :contains()
Install Cheerio in your Node.js project:
npm install cheerio
With ECMAScript modules, import Cheerio and load an HTML string:
import * as cheerio from 'cheerio';
const html = `
<ul>
<li>Apple</li>
<li>Green apple</li>
<li>Banana</li>
</ul>
`;
const $ = cheerio.load(html);
const matches = $('li:contains("Apple")');
console.log(matches.length); // 2
console.log(matches.map((_, element) => $(element).text()).get());
// [ 'Apple', 'Green apple' ]
The selector first limits the search to li elements, then keeps elements whose text contains the substring Apple. It is case-sensitive in this example, so matching and normalization should be chosen deliberately.
Substring matching versus exact text
Use :contains() for intentional substring matches
li:contains("an"), for example, can match any list item containing those characters. This is useful when a label may include additional words, such as finding both “Apple” and “Green apple”. Cheerio also supports most standard pseudo-classes and selector extensions such as :first, :last, and :eq(n) through its selector engine. Those positional extensions are Cheerio features, not selectors you should expect to work in a browser’s native CSS selector API.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
#1 Best Overall
Compare extracted text for exact equality
The documented containment selector is not an exact-equality operator. Select a stable candidate set, extract each element’s text, and compare it:
const exact = $('li').filter((_, element) => {
return $(element).text().trim() === 'Apple';
});
console.log(exact.length); // 1
console.log(exact.first().text()); // Apple
This approach also lets you define your own policy. trim() ignores leading and trailing whitespace; you can normalize repeated whitespace or case when your application requires it. Do not silently normalize values when capitalization or spacing carries meaning.
Normalize safely when the page is inconsistent
const wanted = 'apple';
const normalized = $('li').filter((_, element) => {
const value = $(element).text().replace(/s+/g, ' ').trim().toLowerCase();
return value === wanted;
});
Keep the selector fixed and treat the text as data. Building a selector directly from untrusted input can produce parsing surprises when the value contains selector-special characters.
Load the right kind of input
HTML strings with load
cheerio.load(markup) is the normal choice when your program already has HTML as a string. Cheerio parses it as a document by default and may add html, head, and body wrappers.
Recommended Free Tools
const $ = cheerio.load('<li>One</li>');
console.log($('body > li').length);
Fragments without document wrappers
If the input is only a fragment and those wrappers would be inconvenient, pass false as the third argument:
Rank #2
- HTML CSS Design and Build Web Sites
- Comes with secure packaging
- It can be a gift option
const $ = cheerio.load('<li>One</li>', null, false);
console.log($('li').length); // 1
Buffers and streams
Use loadBuffer for raw bytes when the encoding is unknown. Use stringStream for a stream that already contains decoded text, and decodeStream for a raw-byte stream whose encoding must be detected. The byte-oriented methods can sniff the encoding.
Fetching a URL with fromURL
fromURL is asynchronous and asks Cheerio to retrieve a URL before parsing it. Use it only when having Cheerio perform that fetch is appropriate for your application; otherwise fetch the response yourself so you can control headers, retries, authentication, limits, and error handling.
const $ = await cheerio.fromURL('https://example.com');
const title = $('title').text().trim();
console.log(title);
Regardless of the loader, Cheerio can only inspect the markup it receives.
Extract text correctly
.text() returns text content
.text() reads the selected nodes’ raw textContent. If a selected element contains inline script or style source, that source can appear in the result.
const $ = cheerio.load(`
<div class="card">
<span>Ready</span>
<script>const internal = true;</script>
</div>
`);
console.log($('.card').text());
Use innerText semantics when script and style text should be skipped
const visibleTextLike = $('.card').prop('innerText');
Cheerio documents innerText as skipping script and style text, but it still works from the parsed tree. Cheerio does not apply CSS, so text hidden by display: none or a hidden attribute can still be included. Treat either result as a structural extraction, not as a pixel-accurate view of what a browser displays.
Rank #3
Build selectors that survive markup changes
Text is often the right fallback, but a stable attribute or structural anchor is usually less fragile than a long chain of classes. Prefer a selector such as [data-testid="price"] or a known container, then apply text comparison inside that scope.
const price = $('[data-testid="product-card"]')
.find('.price')
.filter((_, element) => $(element).text().trim() === '$19.99');
Scope matters: the same label may occur in navigation, a footer, and the main content. Narrowing first avoids a false positive and makes the intention clear.
Free tools Windows power users keep installed
One-click scans. No signup required.
When a value is supplied by a user or another untrusted system, do not interpolate it into a selector string. Select a known element and compare its extracted value instead:
const wantedLabel = getLabelFromRequest();
const hit = $('button').filter((_, element) => {
return $(element).text().trim() === wantedLabel;
});
Why an apparently correct text selector returns nothing
Inspect the selection length first
const selection = $('p:contains("Hello")');
console.log('matches:', selection.length);
console.log('loaded HTML:', $.html());
Cheerio returns an empty selection when no nodes match. Calling .text() on that selection simply produces an empty string, which can hide the real problem.
The element is created by client-side JavaScript
Cheerio is not a web browser. It parses the HTML supplied to it; it does not execute scripts, render a page, load external resources, or run a client-side framework. If the browser creates the element after JavaScript runs, it will be absent from the source Cheerio receives. Obtain the rendered markup with a browser automation tool such as Puppeteer or Playwright, then pass that HTML to Cheerio.
Rank #4
- Brand: Wiley
- Set of 2 Volumes
- A handy two-book set that uniquely combines related technologies Highly visual format and accessible language makes these books highly effective learning tools Perfect for beginning web designers and front-end developers
The class or ID is dynamic
Generated class names and IDs can change between requests or builds. Anchor to a stable data- attribute, element relationship, heading, or carefully scoped text instead. Inspect the actual markup for the request that failed rather than relying on a copied browser snapshot.
Do these 3 things before closing this tab:
1Repair Windows errors before they cause bigger problems2Scan for outdated or missing drivers - takes under a minute3Clear out junk files and repair common Windows errorsThe search is scoped incorrectly
A selector run against the wrong container cannot see the desired node. Start with a broad, known anchor, verify its length, and then add narrower conditions one at a time.
Whitespace or case differs
Exact comparisons fail when the source includes line breaks, non-breaking spaces, or different capitalization. Log a representative value, then choose explicit trimming, whitespace folding, or case handling that matches your data contract.
Security and output safety
Cheerio is a parser and DOM manipulation library, not an HTML sanitizer. Scripts and event-handler attributes in input can survive parsing and serialization. If you will render extracted or serialized markup in a browser, sanitize it with a dedicated sanitizer.
Text output can contain characters such as <, >, and quotation marks. Send extracted values to a text context or escape them for the output context you are targeting. Parsing untrusted HTML does not make that HTML safe to publish.
Best Value
Complete Node.js example
This script demonstrates loading, substring matching, exact comparison, scoped selection, and a useful empty-result diagnostic:
import * as cheerio from 'cheerio';
const html = `
<main>
<h1>Products</h1>
<ul data-testid="product-list">
<li>Apple</li>
<li>Green apple</li>
<li>Banana</li>
</ul>
</main>
`;
const $ = cheerio.load(html);
const list = $('[data-testid="product-list"]');
if (list.length === 0) {
throw new Error('Product list was not found in the supplied HTML');
}
const containsApple = list.find('li:contains("Apple")');
console.log('Contains Apple:', containsApple.map((_, el) => $(el).text().trim()).get());
const exactlyApple = list.find('li').filter((_, el) => {
return $(el).text().trim() === 'Apple';
});
console.log('Exactly Apple:', exactlyApple.map((_, el) => $(el).text().trim()).get());
Run it with a project configured for ES modules (for example, a package.json containing "type": "module"), or translate the import to the CommonJS form documented by Cheerio:
const cheerio = require('cheerio');
Performance and reliability choices
- Parse once: call
loadonce for a document and reuse the returned$function for all queries. - Reduce the search space: select a stable container before filtering by text.
- Prefer deterministic comparisons: exact JavaScript comparisons make the normalization policy visible in code.
- Control network behavior: when fetching yourself, set timeouts, response-size limits, and error handling before passing HTML to Cheerio.
- Test real variants: include missing nodes, duplicate labels, extra whitespace, script/style descendants, and empty documents in automated tests.
Cheerio is deterministic for a given input string, buffer, or stream. Reliability problems usually come from the fetch, changing source markup, or client-side rendering that happened after the HTML was obtained—not from a browser viewport or CSS layout.
Or skip the browser setup
If your actual goal is a screenshot of a rendered page rather than text extraction from supplied HTML, ScreenshotNeo provides a single HTTP request. It accepts cookie and consent banners as a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each cleanup step can be disabled. Bot checks and CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and the response identifies the page verdict and billing status in X-Page-Verdict and X-Billed headers. Its MCP server exposes take_screenshot, get_page_info, and capture_pdf to Claude, Cursor, and other MCP clients.
Windows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallOutdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchcURL
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
Node.js
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
See the ScreenshotNeo documentation for parameters. The free plan includes 1,000 screenshots each month with no card; paid plans start at $5 for 3,000 screenshots. Create a free ScreenshotNeo account.
Quick decision guide
| Need | Use | Reason |
|---|---|---|
| Any element containing a phrase | :contains("text") |
Concise substring matching |
| One element whose complete text equals a value | Candidate selector plus .filter() |
Explicit exact comparison and normalization |
| Raw bytes with unknown encoding | loadBuffer or decodeStream |
Encoding detection for byte input |
| Browser-created content | Browser automation, then Cheerio | Cheerio does not execute client JavaScript |
| Visual capture of a live page | ScreenshotNeo | Rendered capture with cleanup and usage-aware billing |
Frequently Asked Questions
Does Cheerio’s :contains() match exact text?
No. It is documented as a substring match. Select candidate elements and compare their extracted text in JavaScript for exact equality.
Can Cheerio find text added by React or another client-side framework?
Not from the original HTML alone. Cheerio does not run scripts or render a browser page; obtain rendered markup with browser automation first.
Why does .text() include unexpected JavaScript?
.text() reads raw textContent, which can include script and style source inside the selected element. Use .prop('innerText') when that behavior better fits your extraction.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




