The reliable pattern is layered: request ordinary pages with Axios, parse the returned HTML with Cheerio, and use Playwright only when the required data appears after JavaScript or browser interaction. At larger volumes, add explicit timeouts, retries, bounded concurrency, queues, deduplication, observability, resumability, and carefully controlled proxies. Cheerio is a parser, not a browser, so it cannot execute page scripts or render a client-side application.
Choose the least powerful layer that contains your data
Start by checking the response that the server returns. If the fields you need are present in that HTML, an HTTP request plus Cheerio is simpler to deploy and maintain than a browser. If the response contains only an application shell and JavaScript later fetches or constructs the data, escalate that URL to browser automation.
| Approach | Use it when | JavaScript and interaction | Operational trade-off |
|---|---|---|---|
| Axios plus Cheerio | Required fields are in the initial server response | Not executed; no visual rendering or external-resource loading | Smallest runtime and easiest deployment, but you must implement request controls |
| Playwright | Fields depend on JavaScript, scrolling, clicks, or browser-only behavior | Executes pages in Chromium, Firefox, or WebKit | Browser binaries, operating-system dependencies, updates, and higher worker complexity |
| Managed crawling/rendering API | You want to outsource some fetching, proxy, or rendering operations | Depends on the vendor and selected mode | Less infrastructure work, with vendor dependency and terms to evaluate; comparable cost and reliability figures are not established here |
This split is a design decision, not a performance promise. Measure your own workload, follow each target’s published rules, and reduce traffic when responses indicate overload.
Install Node.js, Axios, Cheerio, and Playwright
- Use a maintained Node.js release. The current Cheerio introduction states that its current release runs on Node.js 22.19 or later; treat that as version-sensitive and verify it when you upgrade.
- Create a project and install the packages:
mkdir layered-scraper cd layered-scraper npm init -y npm install axios cheerio playwright npx playwright install chromium - If your deployment image is minimal Linux, install the operating-system libraries requested by Playwright’s installation process, and include browser updates in routine maintenance.
- Use ES modules by adding
"type": "module"topackage.json, or convert the examples to CommonJS.
Build the HTTP-first scraper
Axios retrieves a response; Cheerio traverses the markup that you give it. Configure the timeout and retry policy yourself instead of relying on an undocumented default. Check both HTTP status and the presence of expected content before accepting a result.
Do these 3 things before closing this tab:
1Repair Windows errors before they cause bigger problems2Scan for outdated or missing drivers - takes under a minute3Clear out junk files and repair common Windows errors#1 Best Overall
import axios from 'axios';
import * as cheerio from 'cheerio';
const retryableStatuses = new Set([408, 425, 429, 500, 502, 503, 504]);
function sleep(ms) {
return new Promise(resolve => setTimeout(resolve, ms));
}
async function fetchHtml(url, { attempts = 3, timeoutMs = 15000 } = {}) {
let lastError;
for (let attempt = 1; attempt <= attempts; attempt += 1) {
try {
const response = await axios.get(url, {
timeout: timeoutMs,
headers: {
'User-Agent': 'ExampleResearchBot/1.0 ([email protected])',
'Accept': 'text/html,application/xhtml+xml'
},
validateStatus: () => true
});
if (response.status >= 200 && response.status < 300) {
return response.data;
}
const error = new Error(`HTTP ${response.status} for ${url}`);
error.status = response.status;
if (!retryableStatuses.has(response.status) || attempt === attempts) {
throw error;
}
lastError = error;
} catch (error) {
lastError = error;
const status = error.status;
const mayRetry = !status || retryableStatuses.has(status);
if (!mayRetry || attempt === attempts) throw error;
}
const delay = Math.min(8000, 500 * 2 ** (attempt - 1));
await sleep(delay);
}
throw lastError;
}
function parseArticle(html, url) {
const $ = cheerio.load(html);
const title = $('h1').first().text().trim();
const description = $('meta[name="description"]').attr('content')?.trim() || null;
const links = $('a[href]').map((_, element) => ({
text: $(element).text().trim(),
href: $(element).attr('href')
})).get();
if (!title) {
throw new Error(`Expected h1 was not found in ${url}`);
}
return { url, title, description, links };
}
const url = process.argv[2];
if (!url) throw new Error('Usage: node static-scraper.js https://example.com/article');
const html = await fetchHtml(url);
const record = parseArticle(html, url);
console.log(JSON.stringify(record, null, 2));
Selectors should describe stable semantics rather than fragile positional paths. Normalize whitespace, convert dates and numbers deliberately, and keep the original URL with every record. Save a small HTML fixture for each parser so selector changes can be tested without repeatedly contacting a live site.
Know when Cheerio is not enough
Cheerio does not run JavaScript, paint a page, load external resources, or provide browser APIs. A client-rendered page may therefore produce an empty list or an application shell in the Axios response even though a human sees records in a browser. Compare the raw response with the rendered DOM and inspect the network requests triggered by the page. Only escalate when the missing fields are actually produced by execution or interaction.
Render the difficult pages with Playwright
Playwright supports Chromium, Firefox, and WebKit. The example below launches Chromium, waits for a meaningful selector, and extracts the resulting DOM. Set a navigation timeout and close the browser in a finally block so failed jobs do not leak workers.
Rank #2
import { chromium } from 'playwright';
export async function renderArticle(url) {
const browser = await chromium.launch({ headless: true });
try {
const page = await browser.newPage({
userAgent: 'ExampleResearchBot/1.0 ([email protected])'
});
const response = await page.goto(url, {
waitUntil: 'domcontentloaded',
timeout: 30000
});
if (!response || !response.ok()) {
throw new Error(`Navigation failed: ${response?.status() ?? 'no response'}`);
}
await page.locator('h1').first().waitFor({ state: 'visible', timeout: 10000 });
const result = await page.evaluate(() => ({
title: document.querySelector('h1')?.textContent?.trim() || null,
description: document.querySelector('meta[name="description"]')?.getAttribute('content') || null,
links: [...document.querySelectorAll('a[href]')].map(a => ({
text: a.textContent.trim(),
href: a.href
}))
}));
if (!result.title) throw new Error('Rendered page has no h1');
return { url, ...result };
} finally {
await browser.close();
}
}
const url = process.argv[2];
console.log(JSON.stringify(await renderArticle(url), null, 2));
Use explicit waits for evidence of readiness: a selector, a known response, a bounded delay, or network-idle behavior where it is appropriate. Avoid an unbounded “sleep and hope” rule. If a click opens the data, perform the click and wait for the relevant response or selector. Playwright’s request and response events can help identify an endpoint, but call that endpoint directly only when the site permits it and the endpoint is intended for that use.
The Tool Desk
Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Combine both paths with a verified fallback
Keep extraction logic independent from transport. First fetch and parse with Axios and Cheerio. If a required field is absent, enqueue the URL for a browser worker. Record which path produced each record so a sudden increase in fallbacks is visible.
import { fetchHtml } from './static-scraper.js';
import { renderArticle } from './render-scraper.js';
import * as cheerio from 'cheerio';
export async function scrape(url) {
try {
const html = await fetchHtml(url);
const $ = cheerio.load(html);
const title = $('h1').first().text().trim();
if (title) return { url, title, mode: 'http' };
} catch (error) {
// Keep the error for logs; a browser fallback may still succeed.
}
const rendered = await renderArticle(url);
return { ...rendered, mode: 'browser' };
}
In production, distinguish “not found in HTML” from “HTTP request failed.” A 404 should normally be recorded as a permanent result, while a timeout may be retried or sent to a browser queue.
Rank #3
Or skip the browser setup
ScreenshotNeo is a hosted website screenshot API and MCP server. It is useful when your output is a clean PNG, JPEG, WebP, or PDF rather than structured fields. It accepts consent banners as a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be disabled. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing result. For AI workflows, its MCP server exposes take_screenshot, get_page_info, and capture_pdf to Claude, Cursor, and other MCP clients.
See the ScreenshotNeo API documentation for all options, including full-page lazy-image loading, CSS-selector element capture, dark mode, device presets, arbitrary viewports, retina scale, PDF paper and page-range controls, custom CSS and JavaScript, clicks, selector or network-idle waits, request blocking, headers, cookies, user agents, authorization, timezone and geolocation, transparent backgrounds, resizing, chosen cache TTLs, signed image links, asynchronous webhooks, bulk capture of up to 100 URLs per call, usage reporting, and the OpenAPI specification.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
The Free plan includes 1,000 shots per month without a card. Paid plans are Starter $5 for 3,000, Growth $15 for 15,000, Pro $39 for 60,000, Scale $99 for 250,000, and Business $249 for 1,000,000; yearly billing provides two months free, and every feature is included on every plan. Create a free ScreenshotNeo account to try it.
Engineer scale as a set of controls
Bound concurrency and use a queue
Do not launch one unbounded promise per URL. A queue lets you cap active HTTP requests and browser contexts separately, pause intake, retry selected failures, and resume after a process restart. Keep browser concurrency lower when its workers consume more CPU or memory in your environment; there is no universal safe number.
Rank #4
export async function mapWithConcurrency(items, limit, fn) {
const results = new Array(items.length);
let next = 0;
async function worker() {
while (true) {
const index = next++;
if (index >= items.length) return;
results[index] = await fn(items[index], index);
}
}
await Promise.all(Array.from({ length: limit }, worker));
return results;
}
Use separate limits for ordinary HTTP jobs and browser jobs. Apply backpressure when the queue grows, and deduplicate canonical URLs before they enter the queue.
Classify failures before retrying
- Permanent: malformed URL, disallowed target, 401/403 requiring a different authorized workflow, or 404. Record and stop.
- Transient: timeout, connection reset, 408, 429, and many 5xx responses. Retry a small, bounded number of times with exponential backoff and jitter.
- Content failure: a successful response that lacks required fields. Send it to the browser path or a review queue rather than retrying the same request forever.
Honor Retry-After when supplied, stop after the configured attempt count, and include status, elapsed time, attempt number, and parser version in logs.
Make jobs observable and restartable
- Persist job state such as queued, running, succeeded, permanent failure, and retry scheduled.
- Store a content hash or source timestamp when deduplicating, and keep the source URL and capture mode with each record.
- Emit counters for HTTP versus browser captures, status classes, timeout rate, selector-missing rate, and queue age.
- Write checkpoints transactionally so a crash can resume without duplicating completed work.
- Keep raw HTML or a redacted diagnostic sample for parser debugging, subject to your data policy.
Timeouts, waiting, and dynamic content
Use separate budgets for connection, response, navigation, selector waits, and total job time. A page can finish navigation while an API call is still populating a table, so wait for the table’s selector or a documented response rather than simply increasing the navigation timeout. For infinite scroll, set a maximum number of scrolls or a total time budget and stop when the item count stops increasing.
Proxy configuration and trust
Playwright accepts HTTP(S) and SOCKSv5 proxies at browser launch or context level, with credentials and bypass hosts. A minimal launch configuration is:
const browser = await chromium.launch({
proxy: {
server: 'http://proxy.example:8080',
username: process.env.PROXY_USER,
password: process.env.PROXY_PASSWORD,
bypass: 'localhost,127.0.0.1'
}
});
Use only infrastructure you are authorized to use. A proxy is not an anonymity guarantee: operators may see connection metadata and, under some configurations, content. Node.js also has version-specific environment-proxy behavior; verify the behavior of the exact runtime you deploy. Rotation is an operations choice, not permission to evade access controls, rate limits, or bot checks.
Self-managed versus a managed crawler
| Choice | Control | Maintenance | When it fits |
|---|---|---|---|
| Axios and Cheerio you run | Highest control over requests, parsing, storage, and pacing | HTTP reliability, queues, proxies, and parser upkeep are yours | Data is in initial HTML and you need a lean service |
| Playwright you run | Browser and network behavior are directly configurable | Browser binaries, OS dependencies, updates, and worker isolation are yours | JavaScript or interaction is unavoidable |
| Managed API | Defined by vendor settings and API limits | Less infrastructure work; vendor availability, pricing, data handling, and terms become dependencies | You prefer outsourcing fetching, proxy management, or rendering |
Crawlbase’s vendor-authored guide describes its service as returning fetched HTML with optional JavaScript rendering and rotating residential IPs. Those are the vendor’s claims, not an independent benchmark or endorsement; verify current terms, pricing, data handling, and limits before adopting any managed service.
Quick wins for a faster PC:
Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Repair Windows errors before they cause bigger problemsFix Now →Troubleshoot common failures
| Symptom | Likely cause | Fix |
|---|---|---|
| Cheerio returns an empty list | The data is inserted after JavaScript runs, or the selector changed | Inspect the raw HTML, validate the selector against a saved fixture, then use the browser fallback only if the field is absent from the response |
| Playwright cannot launch | Browser binary or Linux dependency is missing | Run the matching Playwright install command in the build image and keep package and browser versions current |
| Navigation times out | Slow target, blocked resource, or an overly short budget | Capture timing details, set bounded navigation and selector budgets, block nonessential resources where allowed, and retry only transient failures |
| Repeated 429 or 403 responses | Traffic exceeds the site’s limits, or access requires authorization | Stop or slow the queue, honor published rules and Retry-After, and obtain permission; do not treat proxies as a bypass |
| Records are duplicated after a restart | No durable job state or URL canonicalization | Persist idempotent job keys and completion state before acknowledging work |
| Rendered text is still missing | The page needs a click, scroll, login, geolocation, or a later API response | Reproduce the required browser action explicitly, wait for its response or selector, and confirm that the workflow is authorized |
Collect responsibly
Before running a scraper, assess the target’s terms, access rules, published robots guidance, request rate, data type, and purpose. Robots.txt is useful operational guidance but is not, by itself, a complete legal permission or prohibition. Personal data, authentication boundaries, copyrighted material, and jurisdiction-specific rules can change the analysis; obtain advice for consequential projects. Identify yourself accurately, provide a contact address where appropriate, minimize retained data, and stop when the operator asks you to stop or responses show overload.
Practical checklist
- Confirm the target and purpose are permitted.
- Capture a raw response and prove whether each required field is server-rendered.
- Use Axios with explicit timeout, status handling, and bounded retry/backoff.
- Parse with Cheerio and test selectors against fixtures.
- Escalate only missing client-rendered fields to Playwright.
- Install and update the required browser binaries and OS dependencies.
- Queue work with separate HTTP and browser concurrency limits.
- Persist state, deduplicate URLs, classify failures, and expose metrics.
- Use only trusted, authorized proxies and document their data-access implications.
Frequently Asked Questions
How can I test parser changes without contacting a live site?
Save representative HTML responses as fixtures and run the Cheerio parser against those files in automated tests. Add a new fixture when a permitted markup change is intentional.
What should I do when a site offers a documented API?
Prefer that API when it fits your purpose, then implement its authentication, pagination, quotas, and terms instead of scraping pages unnecessarily.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

