Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The reliable pattern is layered: request ordinary pages with Axios, parse the returned HTML with Cheerio, and use Playwright only when the required data appears after JavaScript or browser interaction. At larger volumes, add explicit timeouts, retries, bounded concurrency, queues, deduplication, observability, resumability, and carefully controlled proxies. Cheerio is a parser, not a browser, so it cannot execute page scripts or render a client-side application.

Choose the least powerful layer that contains your data

Start by checking the response that the server returns. If the fields you need are present in that HTML, an HTTP request plus Cheerio is simpler to deploy and maintain than a browser. If the response contains only an application shell and JavaScript later fetches or constructs the data, escalate that URL to browser automation.

Approach Use it when JavaScript and interaction Operational trade-off
Axios plus Cheerio Required fields are in the initial server response Not executed; no visual rendering or external-resource loading Smallest runtime and easiest deployment, but you must implement request controls
Playwright Fields depend on JavaScript, scrolling, clicks, or browser-only behavior Executes pages in Chromium, Firefox, or WebKit Browser binaries, operating-system dependencies, updates, and higher worker complexity
Managed crawling/rendering API You want to outsource some fetching, proxy, or rendering operations Depends on the vendor and selected mode Less infrastructure work, with vendor dependency and terms to evaluate; comparable cost and reliability figures are not established here

This split is a design decision, not a performance promise. Measure your own workload, follow each target’s published rules, and reduce traffic when responses indicate overload.

Install Node.js, Axios, Cheerio, and Playwright

  1. Use a maintained Node.js release. The current Cheerio introduction states that its current release runs on Node.js 22.19 or later; treat that as version-sensitive and verify it when you upgrade.
  2. Create a project and install the packages:
    mkdir layered-scraper
    cd layered-scraper
    npm init -y
    npm install axios cheerio playwright
    npx playwright install chromium
  3. If your deployment image is minimal Linux, install the operating-system libraries requested by Playwright’s installation process, and include browser updates in routine maintenance.
  4. Use ES modules by adding "type": "module" to package.json, or convert the examples to CommonJS.

Build the HTTP-first scraper

Axios retrieves a response; Cheerio traverses the markup that you give it. Configure the timeout and retry policy yourself instead of relying on an undocumented default. Check both HTTP status and the presence of expected content before accepting a result.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
import axios from 'axios';
import * as cheerio from 'cheerio';

const retryableStatuses = new Set([408, 425, 429, 500, 502, 503, 504]);

function sleep(ms) {
  return new Promise(resolve => setTimeout(resolve, ms));
}

async function fetchHtml(url, { attempts = 3, timeoutMs = 15000 } = {}) {
  let lastError;

  for (let attempt = 1; attempt <= attempts; attempt += 1) {
    try {
      const response = await axios.get(url, {
        timeout: timeoutMs,
        headers: {
          'User-Agent': 'ExampleResearchBot/1.0 ([email protected])',
          'Accept': 'text/html,application/xhtml+xml'
        },
        validateStatus: () => true
      });

      if (response.status >= 200 && response.status < 300) {
        return response.data;
      }

      const error = new Error(`HTTP ${response.status} for ${url}`);
      error.status = response.status;
      if (!retryableStatuses.has(response.status) || attempt === attempts) {
        throw error;
      }
      lastError = error;
    } catch (error) {
      lastError = error;
      const status = error.status;
      const mayRetry = !status || retryableStatuses.has(status);
      if (!mayRetry || attempt === attempts) throw error;
    }

    const delay = Math.min(8000, 500 * 2 ** (attempt - 1));
    await sleep(delay);
  }

  throw lastError;
}

function parseArticle(html, url) {
  const $ = cheerio.load(html);
  const title = $('h1').first().text().trim();
  const description = $('meta[name="description"]').attr('content')?.trim() || null;
  const links = $('a[href]').map((_, element) => ({
    text: $(element).text().trim(),
    href: $(element).attr('href')
  })).get();

  if (!title) {
    throw new Error(`Expected h1 was not found in ${url}`);
  }

  return { url, title, description, links };
}

const url = process.argv[2];
if (!url) throw new Error('Usage: node static-scraper.js https://example.com/article');

const html = await fetchHtml(url);
const record = parseArticle(html, url);
console.log(JSON.stringify(record, null, 2));

Selectors should describe stable semantics rather than fragile positional paths. Normalize whitespace, convert dates and numbers deliberately, and keep the original URL with every record. Save a small HTML fixture for each parser so selector changes can be tested without repeatedly contacting a live site.

Know when Cheerio is not enough

Cheerio does not run JavaScript, paint a page, load external resources, or provide browser APIs. A client-rendered page may therefore produce an empty list or an application shell in the Axios response even though a human sees records in a browser. Compare the raw response with the rendered DOM and inspect the network requests triggered by the page. Only escalate when the missing fields are actually produced by execution or interaction.

Render the difficult pages with Playwright

Playwright supports Chromium, Firefox, and WebKit. The example below launches Chromium, waits for a meaningful selector, and extracts the resulting DOM. Set a navigation timeout and close the browser in a finally block so failed jobs do not leak workers.

import { chromium } from 'playwright';

export async function renderArticle(url) {
  const browser = await chromium.launch({ headless: true });
  try {
    const page = await browser.newPage({
      userAgent: 'ExampleResearchBot/1.0 ([email protected])'
    });

    const response = await page.goto(url, {
      waitUntil: 'domcontentloaded',
      timeout: 30000
    });
    if (!response || !response.ok()) {
      throw new Error(`Navigation failed: ${response?.status() ?? 'no response'}`);
    }

    await page.locator('h1').first().waitFor({ state: 'visible', timeout: 10000 });
    const result = await page.evaluate(() => ({
      title: document.querySelector('h1')?.textContent?.trim() || null,
      description: document.querySelector('meta[name="description"]')?.getAttribute('content') || null,
      links: [...document.querySelectorAll('a[href]')].map(a => ({
        text: a.textContent.trim(),
        href: a.href
      }))
    }));

    if (!result.title) throw new Error('Rendered page has no h1');
    return { url, ...result };
  } finally {
    await browser.close();
  }
}

const url = process.argv[2];
console.log(JSON.stringify(await renderArticle(url), null, 2));

Use explicit waits for evidence of readiness: a selector, a known response, a bounded delay, or network-idle behavior where it is appropriate. Avoid an unbounded “sleep and hope” rule. If a click opens the data, perform the click and wait for the relevant response or selector. Playwright’s request and response events can help identify an endpoint, but call that endpoint directly only when the site permits it and the endpoint is intended for that use.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Combine both paths with a verified fallback

Keep extraction logic independent from transport. First fetch and parse with Axios and Cheerio. If a required field is absent, enqueue the URL for a browser worker. Record which path produced each record so a sudden increase in fallbacks is visible.

import { fetchHtml } from './static-scraper.js';
import { renderArticle } from './render-scraper.js';
import * as cheerio from 'cheerio';

export async function scrape(url) {
  try {
    const html = await fetchHtml(url);
    const $ = cheerio.load(html);
    const title = $('h1').first().text().trim();
    if (title) return { url, title, mode: 'http' };
  } catch (error) {
    // Keep the error for logs; a browser fallback may still succeed.
  }

  const rendered = await renderArticle(url);
  return { ...rendered, mode: 'browser' };
}

In production, distinguish “not found in HTML” from “HTTP request failed.” A 404 should normally be recorded as a permanent result, while a timeout may be retried or sent to a browser queue.

Or skip the browser setup

ScreenshotNeo is a hosted website screenshot API and MCP server. It is useful when your output is a clean PNG, JPEG, WebP, or PDF rather than structured fields. It accepts consent banners as a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be disabled. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing result. For AI workflows, its MCP server exposes take_screenshot, get_page_info, and capture_pdf to Claude, Cursor, and other MCP clients.

See the ScreenshotNeo API documentation for all options, including full-page lazy-image loading, CSS-selector element capture, dark mode, device presets, arbitrary viewports, retina scale, PDF paper and page-range controls, custom CSS and JavaScript, clicks, selector or network-idle waits, request blocking, headers, cookies, user agents, authorization, timezone and geolocation, transparent backgrounds, resizing, chosen cache TTLs, signed image links, asynchronous webhooks, bulk capture of up to 100 URLs per call, usage reporting, and the OpenAPI specification.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

The Free plan includes 1,000 shots per month without a card. Paid plans are Starter $5 for 3,000, Growth $15 for 15,000, Pro $39 for 60,000, Scale $99 for 250,000, and Business $249 for 1,000,000; yearly billing provides two months free, and every feature is included on every plan. Create a free ScreenshotNeo account to try it.

Engineer scale as a set of controls

Bound concurrency and use a queue

Do not launch one unbounded promise per URL. A queue lets you cap active HTTP requests and browser contexts separately, pause intake, retry selected failures, and resume after a process restart. Keep browser concurrency lower when its workers consume more CPU or memory in your environment; there is no universal safe number.

export async function mapWithConcurrency(items, limit, fn) {
  const results = new Array(items.length);
  let next = 0;

  async function worker() {
    while (true) {
      const index = next++;
      if (index >= items.length) return;
      results[index] = await fn(items[index], index);
    }
  }

  await Promise.all(Array.from({ length: limit }, worker));
  return results;
}

Use separate limits for ordinary HTTP jobs and browser jobs. Apply backpressure when the queue grows, and deduplicate canonical URLs before they enter the queue.

Classify failures before retrying

  • Permanent: malformed URL, disallowed target, 401/403 requiring a different authorized workflow, or 404. Record and stop.
  • Transient: timeout, connection reset, 408, 429, and many 5xx responses. Retry a small, bounded number of times with exponential backoff and jitter.
  • Content failure: a successful response that lacks required fields. Send it to the browser path or a review queue rather than retrying the same request forever.

Honor Retry-After when supplied, stop after the configured attempt count, and include status, elapsed time, attempt number, and parser version in logs.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Make jobs observable and restartable

  • Persist job state such as queued, running, succeeded, permanent failure, and retry scheduled.
  • Store a content hash or source timestamp when deduplicating, and keep the source URL and capture mode with each record.
  • Emit counters for HTTP versus browser captures, status classes, timeout rate, selector-missing rate, and queue age.
  • Write checkpoints transactionally so a crash can resume without duplicating completed work.
  • Keep raw HTML or a redacted diagnostic sample for parser debugging, subject to your data policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Timeouts, waiting, and dynamic content

Use separate budgets for connection, response, navigation, selector waits, and total job time. A page can finish navigation while an API call is still populating a table, so wait for the table’s selector or a documented response rather than simply increasing the navigation timeout. For infinite scroll, set a maximum number of scrolls or a total time budget and stop when the item count stops increasing.

Proxy configuration and trust

Playwright accepts HTTP(S) and SOCKSv5 proxies at browser launch or context level, with credentials and bypass hosts. A minimal launch configuration is:

const browser = await chromium.launch({
  proxy: {
    server: 'http://proxy.example:8080',
    username: process.env.PROXY_USER,
    password: process.env.PROXY_PASSWORD,
    bypass: 'localhost,127.0.0.1'
  }
});

Use only infrastructure you are authorized to use. A proxy is not an anonymity guarantee: operators may see connection metadata and, under some configurations, content. Node.js also has version-specific environment-proxy behavior; verify the behavior of the exact runtime you deploy. Rotation is an operations choice, not permission to evade access controls, rate limits, or bot checks.

Self-managed versus a managed crawler

Choice Control Maintenance When it fits
Axios and Cheerio you run Highest control over requests, parsing, storage, and pacing HTTP reliability, queues, proxies, and parser upkeep are yours Data is in initial HTML and you need a lean service
Playwright you run Browser and network behavior are directly configurable Browser binaries, OS dependencies, updates, and worker isolation are yours JavaScript or interaction is unavoidable
Managed API Defined by vendor settings and API limits Less infrastructure work; vendor availability, pricing, data handling, and terms become dependencies You prefer outsourcing fetching, proxy management, or rendering

Crawlbase’s vendor-authored guide describes its service as returning fetched HTML with optional JavaScript rendering and rotating residential IPs. Those are the vendor’s claims, not an independent benchmark or endorsement; verify current terms, pricing, data handling, and limits before adopting any managed service.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Troubleshoot common failures

Symptom Likely cause Fix
Cheerio returns an empty list The data is inserted after JavaScript runs, or the selector changed Inspect the raw HTML, validate the selector against a saved fixture, then use the browser fallback only if the field is absent from the response
Playwright cannot launch Browser binary or Linux dependency is missing Run the matching Playwright install command in the build image and keep package and browser versions current
Navigation times out Slow target, blocked resource, or an overly short budget Capture timing details, set bounded navigation and selector budgets, block nonessential resources where allowed, and retry only transient failures
Repeated 429 or 403 responses Traffic exceeds the site’s limits, or access requires authorization Stop or slow the queue, honor published rules and Retry-After, and obtain permission; do not treat proxies as a bypass
Records are duplicated after a restart No durable job state or URL canonicalization Persist idempotent job keys and completion state before acknowledging work
Rendered text is still missing The page needs a click, scroll, login, geolocation, or a later API response Reproduce the required browser action explicitly, wait for its response or selector, and confirm that the workflow is authorized

Collect responsibly

Before running a scraper, assess the target’s terms, access rules, published robots guidance, request rate, data type, and purpose. Robots.txt is useful operational guidance but is not, by itself, a complete legal permission or prohibition. Personal data, authentication boundaries, copyrighted material, and jurisdiction-specific rules can change the analysis; obtain advice for consequential projects. Identify yourself accurately, provide a contact address where appropriate, minimize retained data, and stop when the operator asks you to stop or responses show overload.

Practical checklist

  • Confirm the target and purpose are permitted.
  • Capture a raw response and prove whether each required field is server-rendered.
  • Use Axios with explicit timeout, status handling, and bounded retry/backoff.
  • Parse with Cheerio and test selectors against fixtures.
  • Escalate only missing client-rendered fields to Playwright.
  • Install and update the required browser binaries and OS dependencies.
  • Queue work with separate HTTP and browser concurrency limits.
  • Persist state, deduplicate URLs, classify failures, and expose metrics.
  • Use only trusted, authorized proxies and document their data-access implications.

Frequently Asked Questions

How can I test parser changes without contacting a live site?

Save representative HTML responses as fixtures and run the Cheerio parser against those files in automated tests. Add a new fixture when a permitted markup change is intentional.

What should I do when a site offers a documented API?

Prefer that API when it fits your purpose, then implement its authentication, pagination, quotas, and terms instead of scraping pages unnecessarily.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.