You can scrape a static HTML page with node-fetch by requesting its URL, checking the HTTP response, and parsing the returned HTML with a library such as Cheerio. node-fetch fetches the response; it does not provide HTML selectors or run the page’s JavaScript. The distinction matters: if the data appears only after browser-side rendering, a plain fetch will not see it.
What node-fetch does—and what it does not
node-fetch implements the Fetch API for Node.js. It lets a script make HTTP requests and read response bodies using methods such as text() and json(). It is not an HTML parser, browser, or permission to collect any particular website’s data.
A basic scraper therefore has three jobs: fetch a page, decide whether the response is usable, and parse its HTML to extract the fields you need. Cheerio is one option for parsing HTML with a jQuery-like selector API. Fetching and parsing are separate stages, which makes errors easier to diagnose.
Check your Node.js version and module system
Version choice affects both installation and imports. The node-fetch maintainers document v3 as ESM-only and requiring Node.js 12.20.0 or later. The current Cheerio documentation states Node.js 22.19 or later; when using a particular Cheerio release, check that release’s runtime requirement and use the stricter requirement of the packages you install.
#1 Best Overall
- Using ESM: Use
node-fetchv3 withimportstatements and a project configured for ES modules. - Using CommonJS:
require('node-fetch')does not work with v3. Use node-fetch v2 if you need the CommonJS interface, or use dynamicimport()from CommonJS code. - Choosing Cheerio: Check its documented Node.js requirement for the exact version you plan to install. The currently documented requirement may be higher than node-fetch’s.
For a new ESM project using current Cheerio, check that the installed Node.js version satisfies both packages before debugging import or installation errors.
Install the packages
For a project using npm, install node-fetch and Cheerio:
npm install node-fetch cheerio
The example below is ESM. To make import syntax work in a Node.js project, either use an .mjs file or set "type": "module" in the project’s package.json. Save the example as scrape.mjs if you want to avoid changing that setting.
Fetch and parse a static page
This runnable example requests an HTML page, follows redirects up to a set limit, caps the response size, rejects unsuccessful HTTP statuses, and extracts its title and description with Cheerio:
Rank #2
import fetch from 'node-fetch';
import * as cheerio from 'cheerio';
const url = 'https://example.com/';
const controller = new AbortController();
const timeout = setTimeout(() => controller.abort(), 15_000);
try {
const response = await fetch(url, {
method: 'GET',
redirect: 'follow',
follow: 10,
size: 2_000_000,
signal: controller.signal,
headers: {
'user-agent': 'ExampleResearchBot/1.0 (contact: [email protected])',
accept: 'text/html,application/xhtml+xml',
},
});
if (!response.ok) {
throw new Error(`HTTP ${response.status} ${response.statusText} for ${url}`);
}
const contentType = response.headers.get('content-type') ?? '';
if (!contentType.includes('text/html')) {
throw new Error(`Expected HTML but received ${contentType || 'no content type'}`);
}
const html = await response.text();
const $ = cheerio.load(html);
const title = $('title').first().text().trim();
const description = $('meta[name="description"]').attr('content')?.trim() ?? '';
console.log({ url, title, description });
} catch (error) {
if (error.name === 'AbortError') {
console.error(`Request timed out: ${url}`);
} else {
console.error(error);
}
process.exitCode = 1;
} finally {
clearTimeout(timeout);
}
Replace the example URL and user-agent contact with values appropriate to your project. The size cap is in bytes; choose one suited to the pages you expect. A response that exceeds it should be treated as a failure rather than read without bounds.
Why the status check is essential
A 404 or 500 response usually does not cause fetch() to reject. The request can resolve normally with a response object whose status indicates an HTTP failure. Check response.ok (true for successful 2xx responses) or explicitly allow only the status codes your application expects before parsing the body. Reserve the catch branch for network, cancellation, parsing, and other thrown errors; it is not a substitute for checking HTTP status.
What the selectors return
cheerio.load(html) creates a parsed document. $('title').first().text().trim() selects the first title element, gets its text, and trims surrounding whitespace. The description lookup reads the content attribute of the description meta tag. Pages may omit either field, so extraction code should tolerate empty or missing values.
Extracting more than one record
After inspecting the page’s HTML and identifying a selector that represents one record, select all matching elements and map them into structured data. For example, if each item is in an element with the class product, and each has a title link and price element:
Rank #3
const products = $('.product').map((_, element) => {
const item = $(element);
return {
title: item.find('.product-title').first().text().trim(),
href: item.find('.product-title').first().attr('href') ?? null,
price: item.find('.price').first().text().trim(),
};
}).get();
console.log(products);
Those selectors are illustrative, not universal: replace them with selectors that exist in the target page’s returned HTML. A relative link such as /products/1 is not an absolute URL; resolve it against the page URL when your downstream code needs a complete address. Validate extracted values rather than assuming every card has every field.
Choose the right response handling
Use the response method that matches the resource. For HTML, read response.text() and pass the string to an HTML parser. For a JSON endpoint, use response.json() and handle parse failures. For binary data, use the response body as a stream or buffer as appropriate rather than decoding it as text.
node-fetch supports Node streams and documents automatic decoding for gzip, deflate, and Brotli responses. It also provides controls for redirects and response size. These capabilities help bound and interpret a request, but they do not guarantee that a site will return the page you expect.
Handle timeouts, redirects, and response size
Cancel slow requests
In node-fetch v3, the old non-standard timeout option is removed. Use an AbortSignal, as in the example, to cancel a request that takes too long. Always clear the timer after the request finishes so it does not remain active unnecessarily. A timeout is a policy decision: set it to fit the target and your application, not as a promise that every slower page is defective.
Free tools Windows power users keep installed
One-click scans. No signup required.
Rank #4
Set a redirect policy
The example uses redirect: 'follow' with a maximum of ten redirects. You can choose 'manual' when the application needs to inspect redirect responses, or 'error' when redirects should cause failure. Set the follow limit intentionally; redirect loops and unexpectedly long chains should not run indefinitely.
Limit body size
The size option limits the response body size. This helps avoid accidentally buffering an unexpectedly large page into memory when calling response.text(). Select a limit based on realistic page sizes for your task and handle a size-limit failure as a failed fetch, rather than silently accepting incomplete content.
Cookies, headers, and sessions
node-fetch does not store cookies between requests by default. If a site explicitly permits access that requires a session, you must manage cookies deliberately: inspect relevant response headers and forward permitted cookie values on a later request, or use a cookie-jar solution. A single request with a Cookie header is not equivalent to a browser maintaining a complete session, and cookies should be protected as credentials.
Request headers can identify your client or state acceptable response types. Use an honest, stable user-agent and contact information where appropriate. Do not use headers to impersonate a browser or bypass access controls. A successful HTTP response is not evidence that automated collection is permitted.
Do these 3 things before closing this tab:
1Repair Windows errors before they cause bigger problems2Fix the driver behind crashes, sound loss and screen glitches3Clear out junk files and repair common Windows errorsJavaScript-rendered pages are a different problem
node-fetch downloads the HTTP response; it does not execute JavaScript in a browser environment. If the HTML response already contains the data, fetch plus Cheerio can be enough. If a page fills its content only after client-side scripts run, the fetched source may contain only a shell and your selectors may find nothing.
When that happens, first determine whether the site offers a documented API or other permitted data source. If browser rendering is genuinely required, use an appropriate browser automation approach and account for its additional runtime, resource use, and site rules. Do not treat a screenshot as extracted text: image capture can show a rendered page, but it does not replace a parser or return structured records.
Be deliberate about request load and URL safety
Scraping is not made safe or permitted simply by using a particular library. Review the site’s terms and applicable robots guidance, make requests at a measured rate, cache results when suitable, and avoid concurrency that creates unnecessary load. These are operational safeguards, not a guarantee of permission for a particular target.
If your application accepts URLs from users, treat them as untrusted input. Restrict allowed schemes to those you need, validate or allow-list hosts, and guard against server-side request forgery. Otherwise a user may cause your server to request internal services or addresses that should not be reachable. Cheerio’s URL-loading documentation also notes security considerations around user-supplied URLs; validate before fetching, not only while parsing.
Quick wins for a faster PC:
Repair Windows errors before they cause bigger problemsFix Now →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Clear out junk files and repair common Windows errorsFree Scan →Common node-fetch scraping failures
| Symptom | Likely cause | What to do |
|---|---|---|
| A 404 or 500 is parsed as if it were a normal page | HTTP error statuses resolve to a response; they do not automatically throw. | Check response.ok or an explicit status allow-list before reading and parsing the body. |
require('node-fetch') fails |
node-fetch v3 is ESM-only. | Use ESM imports, dynamic import(), or node-fetch v2 for a CommonJS project. |
| The request hangs or takes too long | No cancellation policy is configured, or the server is slow. | Pass an AbortSignal with a suitable timeout and handle AbortError. |
| Memory use grows while reading pages | The response body is larger than expected or is being buffered without a bound. | Set a suitable size limit; for larger expected data, consider streaming and process it incrementally. |
| Selectors return empty strings or no records | The selector does not match the returned HTML, the fields are absent, or JavaScript adds them after fetch. | Inspect the fetched HTML, verify selectors against it, tolerate missing fields, and assess whether browser rendering or an allowed API is needed. |
| A page works once but a later request lacks access | Cookies are not retained automatically, or access depends on a session. | Manage permitted cookies explicitly or use an appropriate cookie jar; do not attempt to evade access restrictions. |
| A package installs but fails on the project’s Node version | The selected Cheerio release may require a newer runtime than node-fetch. | Check the exact release requirements and satisfy the stricter package requirement. |
| The request follows an unexpected chain or fails on redirects | The target redirects repeatedly, or the chosen redirect policy does not match the application. | Set the redirect mode and maximum follow count deliberately; inspect redirects when needed. |
Or skip the browser setup
If the task is to capture a rendered page as an image or PDF rather than extract structured HTML, ScreenshotNeo is a website screenshot API and MCP server. It can accept a URL in one GET request and return PNG, JPEG, WebP, or PDF. Its capture flow accepts cookie or consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; those steps can be turned off. Bot checks and CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers report the page verdict and billing status. Its MCP server includes take_screenshot, get_page_info, and capture_pdf for MCP clients including Claude and Cursor.
For a single image capture with cURL:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp
See the ScreenshotNeo API documentation for authentication and request options. The service is for rendered captures; use node-fetch and an HTML parser when you need to collect structured page data. ScreenshotNeo’s free plan includes 1,000 shots per month with no card, and paid plans start at $5 for 3,000 shots. Sign up free for 1,000 screenshots a month, with no card required.
Frequently Asked Questions
Does node-fetch need Cheerio to scrape a page?
No. node-fetch can retrieve HTML on its own; Cheerio or another parser is useful when you need to select and extract elements from that HTML.
Will node-fetch return the same content I see in my browser?
Not necessarily. It retrieves the HTTP response but does not run the page’s browser-side JavaScript or reproduce a browser session automatically.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




