The Tool Desk
Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
Use Cheerio when the fields are already in the HTML response, jsdom when your code needs DOM semantics without a full browser, and Playwright when JavaScript execution or browser network behavior creates the data. For large responses, fetch and transform as streams instead of buffering everything. The reliable workflow is to define the source contract, validate status and content type, choose the lightest execution model that contains the data, normalize and validate records, and make failures observable.
Start with the source, not the library
Before writing a selector, describe the source contract in concrete terms:
- Location: a page URL, an API endpoint, or a paginated sequence.
- Response: expected status codes and content type (HTML, XML, JSON, or a file).
- Fields: required and optional values, their types, and how missing values should be represented.
- Access: authentication headers or cookies, rate limits, redirect limits, and a timeout.
- Provenance: source URL, retrieval time, and any page or request identifier that lets you audit a record.
This contract prevents a common error: a script that returns rows successfully while silently extracting an error page, a login form, or a client-rendered shell with none of the desired data.
Choose the execution model
| Need | Best starting point | What it does | What it does not do |
|---|---|---|---|
| Delivered HTML or XML | Cheerio | Parses markup and provides fast, jQuery-like traversal. | It does not execute JavaScript, render a page, or load external resources. |
| DOM-shaped application logic | jsdom | Emulates many WHATWG DOM and HTML standards in pure JavaScript. | It is not a complete browser and cannot reproduce every browser API or behavior. |
| Client rendering, authenticated browser state, or network interception | Playwright | Runs a real browser, observes requests and responses, and can alter traffic. | It costs more memory and startup time than a parser. |
| Very large bodies or incremental records | Node HTTP/Web Streams plus an incremental parser | Streams data with backpressure instead of retaining the whole response. | A stream alone does not understand HTML structure; you still need a parser or record framing. |
Node’s HTTP API is intentionally low-level and does not buffer an entire request or response, making it suitable for streaming. The Web Streams API follows WHATWG semantics and can interoperate with Node streams through Readable.toWeb() and Readable.fromWeb(); see Node HTTP documentation and Node Web Streams documentation.
#1 Best Overall
Fetch and validate before parsing
A robust fetch checks the timeout, status, content type, and redirect behavior before handing bytes to an extractor. This Node.js example uses the built-in fetch available in current Node releases and keeps the response bounded in memory for a normal page.
const { setTimeout: delay } = require('node:timers/promises');
async function fetchHtml(url, { timeoutMs = 30_000, userAgent = 'example-extractor/1.0' } = {}) {
const controller = new AbortController();
const timer = setTimeout(() => controller.abort(), timeoutMs);
try {
const response = await fetch(url, {
signal: controller.signal,
redirect: 'follow',
headers: {
'user-agent': userAgent,
'accept': 'text/html,application/xhtml+xml'
}
});
if (!response.ok) {
throw new Error(`HTTP ${response.status} for ${url}`);
}
const type = response.headers.get('content-type') || '';
if (!type.includes('text/html') && !type.includes('application/xhtml+xml')) {
throw new Error(`Unexpected content type: ${type || 'missing'}`);
}
return { response, finalUrl: response.url };
} finally {
clearTimeout(timer);
}
}
(async () => {
const { response, finalUrl } = await fetchHtml('https://example.com/articles');
const html = await response.text();
console.log({ finalUrl, bytes: Buffer.byteLength(html) });
})();
For production, add an explicit maximum body size, a bounded redirect policy, and structured logs. A 404 or 503 is still an HTTP response; treat it as a failure rather than parsing its error document as if it were a page.
Extract static markup with Cheerio
Cheerio is the lightest choice when the desired values are present in the delivered response. Its load() method accepts a string. For unknown encodings, loadBuffer() accepts bytes and performs encoding detection. stringStream() and decodeStream() support stream-oriented input, while fromURL() performs the fetch for you. The loader details and options are documented at Cheerio loading methods.
A complete static-page extractor
const cheerio = require('cheerio');
async function extractArticles(url) {
const response = await fetch(url, {
headers: { 'user-agent': 'article-extractor/1.0', 'accept': 'text/html' }
});
if (!response.ok) throw new Error(`HTTP ${response.status}`);
const type = response.headers.get('content-type') || '';
if (!type.includes('text/html')) throw new Error(`Not HTML: ${type}`);
const html = await response.text();
const $ = cheerio.load(html);
const records = [];
$('article').each((index, element) => {
const title = $(element).find('h2, h3').first().text().replace(/s+/g, ' ').trim();
const href = $(element).find('a[href]').first().attr('href');
if (!title || !href) return;
records.push({
title,
url: new URL(href, response.url).href,
position: index,
sourceUrl: response.url,
retrievedAt: new Date().toISOString()
});
});
if (records.length === 0) throw new Error('No article records found; check rendering or selectors');
return records;
}
extractArticles('https://example.com/news').then(console.log).catch(console.error);
Use the loader that matches the bytes
fromURL() follows up to five redirects, rejects non-2xx responses, refuses non-markup content types, and uses the final URL as the base URI. When supplying request options, provide the HTTP method; custom headers replace the defaults, so include a user agent and an Accept header yourself. For malformed input or XML, Cheerio’s parser configuration matters: standards-oriented parse5 is the default HTML parser, while htmlparser2 is faster, uses less memory, and is more forgiving of malformed markup. See Cheerio parser configuration.
Cheerio never runs page JavaScript or loads external resources. If the initial response contains only an application shell and scripts, the target fields will be absent. Cheerio’s own introduction points to Puppeteer, Playwright, or a DOM-emulation project such as jsdom for that case: Cheerio introduction.
Use jsdom for DOM-shaped extraction logic
Choose jsdom when your extraction code expects document, CSS selectors, element properties, or other DOM semantics, but does not require a full browser. jsdom implements many WHATWG DOM and HTML standards in pure JavaScript and is intended to emulate enough of a browser for testing and scraping web applications. Its scope and options are described in the jsdom README.
Rank #2
const { JSDOM } = require('jsdom');
async function extractWithDom(url) {
const response = await fetch(url, { headers: { 'user-agent': 'dom-extractor/1.0' } });
if (!response.ok) throw new Error(`HTTP ${response.status}`);
const html = await response.text();
const dom = new JSDOM(html, { url });
const { document } = dom.window;
return [...document.querySelectorAll('[data-product]')].map((node) => ({
name: node.querySelector('.name')?.textContent.trim() || null,
price: node.querySelector('.price')?.textContent.trim() || null,
sourceUrl: document.URL
}));
}
extractWithDom('https://example.com/catalog').then(console.log).catch(console.error);
Do not assume jsdom will execute a site’s scripts exactly as Chrome does. If data appears only after framework code, service workers, browser APIs, or complex network activity run, move to Playwright.
Quick wins for a faster PC:
Repair Windows errors before they cause bigger problemsFix Now →Scan for outdated or missing drivers - takes under a minuteDriver Scan →Use Playwright when the browser is part of the data source
Playwright is appropriate for client-rendered pages, login sessions, interactions, and network-dependent extraction. It can intercept requests, fetch a response for inspection or modification, change headers, set a maximum redirect count, and expose request lifecycle events. The relevant APIs are documented at Playwright route and Playwright request.
const { chromium } = require('playwright');
(async () => {
const browser = await chromium.launch();
const page = await browser.newPage({ userAgent: 'browser-extractor/1.0' });
page.on('response', response => {
if (response.status() >= 400) {
console.warn('HTTP error', response.status(), response.url());
}
});
page.on('requestfailed', request => {
console.warn('Request failed', request.url(), request.failure()?.errorText);
});
await page.goto('https://example.com/app', { waitUntil: 'networkidle', timeout: 60_000 });
await page.locator('[data-product]').first().waitFor({ state: 'visible', timeout: 15_000 });
const products = await page.locator('[data-product]').evaluateAll(nodes => nodes.map(node => ({
name: node.querySelector('.name')?.textContent.trim() || null,
price: node.querySelector('.price')?.textContent.trim() || null
})));
console.log(products);
await browser.close();
})();
Inspect or modify a response with route.fetch()
await page.route('**/api/products**', async route => {
const upstream = await route.fetch({
maxRedirects: 5,
headers: { ...route.request().headers(), 'x-extractor': '1' }
});
const json = await upstream.json();
const filtered = json.filter(item => item.active);
await route.fulfill({ response: upstream, json: filtered });
});
Listen to request, response, requestfinished, and requestfailed when diagnosing a data flow. A 404 or 503 normally appears as a response event, not a failed request, so inspect the status explicitly.
Stream large responses instead of buffering them
For feeds or exports that can exceed available memory, consume the body incrementally and apply backpressure. A line-delimited JSON endpoint is a straightforward example because each newline delimits one record:
const { Transform } = require('node:stream');
const { pipeline } = require('node:stream/promises');
async function streamNdjson(url, onRecord) {
const response = await fetch(url, { headers: { accept: 'application/x-ndjson' } });
if (!response.ok || !(response.headers.get('content-type') || '').includes('ndjson')) {
throw new Error(`Unexpected response: ${response.status}`);
}
let carry = '';
const splitter = new Transform({
writableObjectMode: false,
readableObjectMode: true,
transform(chunk, encoding, callback) {
carry += chunk.toString('utf8');
const lines = carry.split('n');
carry = lines.pop();
for (const line of lines) {
if (line.trim()) this.push(JSON.parse(line));
}
callback();
},
flush(callback) {
if (carry.trim()) this.push(JSON.parse(carry));
callback();
}
});
splitter.on('data', onRecord);
await pipeline(response.body, splitter);
}
streamNdjson('https://example.com/export.ndjson', record => {
// Persist or process one record, then release it.
console.log(record.id);
}).catch(console.error);
Node’s Web Streams and Node streams are different interfaces; use the documented conversion helpers when a library expects one type. For HTML, arbitrary tags can span chunks, so do not split on byte boundaries and call a regex a parser. Use a streaming-capable HTML parser, or choose Cheerio’s stringStream()/decodeStream() when their input and output model fits your workload.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Normalize, validate, and preserve provenance
- Collapse incidental whitespace while preserving meaningful text and non-breaking spaces where required.
- Resolve relative links against the final response URL, not the originally requested URL after redirects.
- Parse numbers and dates with locale and timezone rules stated in the source contract; retain the original text if conversion fails.
- Require key fields and reject or quarantine incomplete records instead of silently emitting partial rows.
- Store the source URL and retrieval timestamp with each record so a later correction can be traced.
Selectors are part of your code’s interface. Keep representative HTML fixtures and rerun them in tests whenever the source layout changes. A sudden zero-record result should be an alert, not a successful empty export.
Rank #3
Retries, limits, and responsible operation
Retry only transient failures, with a small maximum and exponential backoff. Do not retry deterministic 4xx responses or parser validation failures. Make jobs idempotent: checkpoint the last page or record key, and write output so a restart cannot duplicate committed records. Set concurrency and request delays to respect the site’s published limits, terms, access controls, and robots guidance applicable to your source.
Record timing, status, content type, redirect destination, bytes received, and extracted-record counts. Those measurements distinguish a slow origin from a selector break and make partial failures recoverable.
Performance and cost trade-offs
| Approach | Startup and memory | When it wins | Typical limitation |
|---|---|---|---|
| Node HTTP/fetch plus streaming | Lowest overhead; memory can remain bounded. | Large APIs, exports, and line-delimited feeds. | You must handle framing and parsing. |
| Cheerio | Lightweight for a complete document. | Many static pages with predictable markup. | Requires the data to be in delivered bytes. |
| jsdom | More memory and CPU than a direct parser. | Code that needs DOM methods and browser-like document behavior. | Not full browser execution. |
| Playwright | Highest startup, memory, and operational cost. | JavaScript-rendered apps, interactions, sessions, and network interception. | Browser lifecycle, timeouts, and request failures add complexity. |
Benchmark your actual pages and fields rather than assuming a library is universally faster. Reuse browser contexts when safe, close pages promptly, cap concurrency, and avoid loading images or other unneeded resources in browser jobs.
Or skip the browser setup
If your goal is a clean screenshot or PDF of a page before extracting or reviewing it, ScreenshotNeo provides a single HTTP call. It accepts cookie and consent banners like a visitor, removes more than 60 known consent platforms plus newsletter popups and chat widgets, and lets each cleanup step be turned off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed; response headers identify the page verdict and billing result. Its MCP server exposes take_screenshot, get_page_info, and capture_pdf to Claude, Cursor, and other MCP clients.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
See the ScreenshotNeo API documentation for all options, including full-page lazy-image loading, CSS-selector element capture, dark mode, device presets and custom viewports, retina scale, PDF paper and page settings, custom CSS or JavaScript, clicks, waits, blocked resources, headers, cookies, user agents, authorization, timezone, geolocation, transparent backgrounds, resizing, chosen cache TTLs, signed image links, asynchronous jobs with signed webhooks, bulk capture of up to 100 URLs per call, usage reporting, and the OpenAPI specification.
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
The Free plan includes 1,000 shots per month with no card. Paid plans start at $5 for 3,000 shots; yearly billing gives two months free. Create a free ScreenshotNeo account.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Troubleshooting extraction failures
Selectors return zero records
Save the raw response and inspect it. If the values are absent, the page is likely client-rendered; switch from Cheerio to jsdom or Playwright. If values are present, check selector scope, case, whitespace, and whether a redirect changed the final document.
Rank #4
HTML is actually an error or login page
Log status, final URL, and content type before parsing. Supply required cookies or authorization, and stop on non-2xx responses instead of exporting the error document.
Characters are corrupted
Use a byte-aware loader such as Cheerio’s loadBuffer() when encoding is uncertain. A string created with the wrong encoding cannot be repaired reliably later.
The process runs out of memory
Do not concatenate unbounded chunks. Stream record-framed data, lower concurrency, release page objects, and use a streaming parser for large markup. Browser jobs should process one page at a time unless measurements justify more.
Playwright reports a timeout
Distinguish navigation completion from data readiness. Wait for a specific selector or response that proves the field exists, set a bounded timeout, and log failed requests. Avoid treating networkidle as universal proof that an application is finished.
Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minuteWindows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallIntermittent 429, 503, or connection failures
Reduce concurrency, honor retry-after when supplied, add bounded exponential backoff, and checkpoint progress. Do not retry forever or convert repeated access failures into duplicate records.
FAQ
Can I use regular expressions to extract HTML?
Use an HTML parser for nested or malformed markup. Regular expressions can help with a small, already-isolated text fragment, but they do not model HTML structure safely.
How do I preserve the exact page version I extracted?
Store the retrieved timestamp, final URL, response metadata, and (where permitted) a fixture or content hash alongside normalized records.
What should an empty result mean in a data pipeline?
Treat it as a state requiring validation unless the source contract explicitly permits an empty collection. Emit a metric or alert so a layout change cannot masquerade as a valid zero.
Free tools Windows power users keep installed
One-click scans. No signup required.
Is a browser always more accurate than a parser?
Only when browser execution is part of the source behavior. For data already delivered as HTML, a parser avoids browser variability and is usually simpler to operate.
Frequently Asked Questions
Which Node.js version should I target for this workflow?
Use a currently supported Node.js release with built-in Web Streams and fetch, and pin your parser and browser-library versions in the project lockfile so upgrades are deliberate.
How can I test an extractor without repeatedly calling the live site?
Save representative, permissioned responses as fixtures and run selector, normalization, and validation tests against those files; reserve live checks for a small monitoring job.
Should extracted data be stored as strings or typed values?
Keep the original text for auditability, then add validated typed fields for dates, numbers, and identifiers. That preserves what the source said while making downstream queries reliable.
Do these 3 things before closing this tab:
1Fix the driver behind crashes, sound loss and screen glitches2Clear out junk files and repair common Windows errors3Scan for outdated or missing drivers - takes under a minuteQuick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

