Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

Use a headless browser when the information you need appears only after a page runs JavaScript. For a new Node.js crawler, a practical starting point is Playwright: it opens pages in a real browser engine, waits for a signal relevant to the page, and lets your code read the resulting DOM. Use ordinary HTTP fetching instead when the initial HTML already contains the data; browser rendering adds setup and operational complexity and does not guarantee access or successful extraction.

Choose browser rendering only when the page needs it

An HTTP crawler requests a URL and parses the response body. That is usually the simpler route for static HTML, but it does not execute client-side JavaScript. Crawlee describes its CheerioCrawler as fast and efficient for plain HTTP/HTML work, while noting that it cannot handle JavaScript rendering. Its browser-backed choices include PlaywrightCrawler and PuppeteerCrawler; Crawlee recommends Playwright in its quick start for a new headless-browser project (Crawlee Quick Start).

Before building a browser crawler, compare the target text in the initial response HTML with what appears in a browser after the page loads. If the response already contains the fields you need, prefer an HTTP parser. If JavaScript creates or fills those fields, render the page. This keeps the browser work focused on pages that require it.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A browser-rendered page is still just the result of your crawler’s browser session. It does not prove that the target permits your access, that every user sees the same page, or that your crawler behaves like Googlebot. Google documents its own crawling and JavaScript processing separately (Google Crawling and Indexing).

Pick a Node.js browser stack

For this guide, use Playwright directly: it provides browser control and DOM access without requiring a crawler framework. Crawlee is useful when you want a framework that offers both HTTP-only and browser-backed crawler classes behind a common interface. Puppeteer is another supported browser-automation option, particularly if it is already part of your project. Google’s overview describes Puppeteer as a browser automation tool (Google Chrome for Developers: Puppeteer).

Choice Use it when Important constraint
HTTP parser, such as Crawlee’s CheerioCrawler The useful fields exist in returned HTML. It does not execute page JavaScript.
Playwright You need browser rendering and want to choose among browser engines. Browser binaries must match the installed Playwright release.
Puppeteer or Crawlee’s PuppeteerCrawler Your project already uses Puppeteer or you prefer its API. Crawlee describes its PuppeteerCrawler as controlling Chromium or Chrome.

Playwright documents support for Chromium, Firefox, and WebKit, with branded Chrome and Edge available when installed or through its CLI. These options do not guarantee identical rendering across browsers, so select an engine that matches your use case and validate it against the target (Playwright Browsers). Crawlee’s class comparison and installation guidance are at its quick-start page.

Install Node.js, Playwright, and a browser

Crawlee’s current quick start states Node.js 16 or later as its requirement; versions change, so check the installation documentation when setting up a new project. The commands below create a small Playwright project directly rather than scaffolding a Crawlee application.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  1. Install a supported Node.js release for your operating system. Confirm it is available in the terminal with node --version and npm --version.
  2. Create a project and install Playwright: mkdir rendered-crawler && cd rendered-crawler, then npm init -y and npm install playwright.
  3. Configure the project for ES modules by adding "type": "module" to the top-level fields of package.json. Alternatively, use a .mjs filename.
  4. Download the browser binary supported by your installed Playwright: npx playwright install chromium. Playwright’s release is coupled to particular browser versions; run the install command again after upgrading Playwright if needed. On supported Linux environments, its documentation also describes installing operating-system dependencies with npx playwright install-deps chromium (Playwright browser installation).

Build a cautious, bounded crawler

This example starts from one URL, stays on that URL’s host, visits at most 20 pages, waits for a target-specific selector, extracts a title and description, and writes JSON Lines to disk. It uses a selector because the appearance of the content is a better readiness signal than assuming that a generic navigation event means the application is finished. Change CONTENT_SELECTOR and the extraction selectors for the site you are authorized to crawl.

Save as crawler.js in the project directory, then run START_URL=https://example.com/ node crawler.js. Replace the example address with an appropriate target.

import { chromium } from 'playwright';
import { appendFile } from 'node:fs/promises';

const startUrl = process.env.START_URL;
const contentSelector = process.env.CONTENT_SELECTOR ?? 'main';
const maxPages = Number(process.env.MAX_PAGES ?? 20);
const delayMs = Number(process.env.DELAY_MS ?? 1000);

if (!startUrl) {
  throw new Error('Set START_URL to the first page to crawl.');
}
if (!Number.isInteger(maxPages) || maxPages < 1) {
  throw new Error('MAX_PAGES must be a positive integer.');
}

const start = new URL(startUrl);
const allowedHost = start.host;
const queue = [start.href];
const queued = new Set(queue);
const browser = await chromium.launch({ headless: true });

try {
  const context = await browser.newContext();
  const page = await context.newPage();
  page.setDefaultNavigationTimeout(30000);

  while (queue.length > 0 && queued.size <= maxPages) {
    const url = queue.shift();
    try {
      const response = await page.goto(url, { waitUntil: 'domcontentloaded' });
      if (!response) {
        console.warn(`No main document response: ${url}`);
        continue;
      }
      if (response.status() >= 400) {
        console.warn(`HTTP ${response.status()}: ${url}`);
        continue;
      }

      // Wait for the useful content, not an assumed universal "page ready" event.
      await page.locator(contentSelector).first().waitFor({ state: 'visible', timeout: 10000 });
      const record = await page.evaluate(() => ({
        title: document.querySelector('h1')?.textContent?.trim() ?? document.title,
        description: document.querySelector('meta[name="description"]')?.content ?? null,
        text: document.querySelector('main')?.innerText?.trim() ?? null,
        url: location.href,
        crawledAt: new Date().toISOString()
      }));
      await appendFile('pages.jsonl', `${JSON.stringify(record)}n`);
      console.log(`Saved ${record.url}`);

      const links = await page.locator('a[href]').evaluateAll(anchors =>
        anchors.map(anchor => anchor.href)
      );
      for (const href of links) {
        try {
          const next = new URL(href);
          next.hash = '';
          if (next.protocol.startsWith('http') &&
              next.host === allowedHost &&
              !queued.has(next.href) &&
              queued.size < maxPages) {
            queued.add(next.href);
            queue.push(next.href);
          }
        } catch {
          // Ignore malformed or non-URL href values.
        }
      }
    } catch (error) {
      console.warn(`Skipped ${url}: ${error.message}`);
    }

    if (queue.length > 0) {
      await new Promise(resolve => setTimeout(resolve, delayMs));
    }
  }
} finally {
  await browser.close();
}

The code extracts rendered text from main, even though the wait selector is configurable. If a site uses a different content container, update both the readiness selector and the extraction expression. The sample writes one JSON object per line to pages.jsonl; it does not attempt to model every site’s pagination, authentication, or page-specific data structure.

How the crawl is bounded

  • Host boundary: links are normalized and restricted to the starting host. This prevents the queue from wandering into unrelated sites; it does not distinguish every sub-area or content type on that host.
  • Page cap: MAX_PAGES limits the number of unique URLs added to the queue. Lower it for an initial test.
  • Delay: DELAY_MS pauses between completed navigations. This is a basic throttle, not a rate-limit guarantee; adjust it to the site’s rules and observed response behavior.
  • Timeout and failure handling: navigation has a 30-second default timeout, and missing content or failed navigation is logged and skipped rather than terminating the entire run.
  • Resource cleanup: the browser is closed in a finally block so it is shut down even if an error interrupts the loop.

Wait for the right signal and extract deliberately

domcontentloaded says the initial document has been parsed; it does not mean a client-side app has finished fetching and rendering its data. A fixed delay can work for a known site, but it is brittle: slow responses may need longer, while fast pages waste time. Prefer a selector tied to the content you need, as in the example. If the target has a reliable application-specific signal, use that instead.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Playwright and Puppeteer both expose page lifecycle and request events. Playwright’s Page API documents page events and request listeners (Playwright Page API); Puppeteer’s Page reference demonstrates navigation, screenshots, lifecycle, and request-event examples (Puppeteer Page class). These APIs help diagnose what a page is doing, but no one event is a universal indication that every application’s useful content is ready.

Keep extracted records small and intentional. Store the canonical source URL you actually visited, the fields needed for your task, and a crawl timestamp. Normalize whitespace and handle absent fields explicitly, as the sample does for the description. Avoid collecting unrelated personal or sensitive data.

Crawl responsibly; robots.txt is not access control

Read the site’s published crawl guidance and terms, keep request rates modest, and stop if the site signals that the activity is unwanted or causes operational problems. Google Search Central explains that robots.txt communicates which URLs a crawler may request, but rules cannot enforce crawler behavior. Disallowed URLs can still appear in search results if discovered elsewhere; robots.txt is not a way to protect private content. Use authentication and authorization controls for private material, and use documented search-visibility controls when the goal is to keep content out of results (Google’s robots.txt guide).

Robots policies and search indexing are distinct from whether your program can technically render a page. Do not assume that a page accessible in your browser is fair to crawl, or that a rendered copy reflects Google’s view of it.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Or skip the browser setup

If your immediate need is a screenshot or PDF rather than a custom Node.js extraction pipeline, ScreenshotNeo is a website screenshot API and MCP server for developers. One GET request can return PNG, JPEG, WebP, or PDF. Its clean-shot steps accept the cookie or consent banner like a visitor and remove more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and the response identifies page verdict and billing status in headers. An MCP server exposes screenshot and PDF tools to AI agents. These are screenshot capabilities, not a replacement for a crawler that extracts arbitrary page fields.

Example cURL call (see the ScreenshotNeo documentation for API details):

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python equivalent:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js equivalent:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. Every feature is on every plan. For AI workflows, its MCP server includes take_screenshot, get_page_info, and capture_pdf. Start with 1,000 free screenshots a month with no card.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Troubleshoot common failures

Playwright launches but cannot find Chromium

The Playwright package and browser binary may not both be installed, or the binary may not match the package release. Run npx playwright install chromium after installing or upgrading Playwright. On supported Linux setups, install the required operating-system packages using the documented dependency command.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The page loads but the selector times out

The selector may be wrong, the content may be outside the selected frame, or the page may not have rendered the expected state. Inspect the page in a browser, verify the selector against the live DOM, and determine whether the site requires a different readiness signal. Do not simply increase the timeout without checking the cause.

The crawler skips pages or saves empty fields

Check the logs for HTTP status, navigation failures, and the current selector. The sample only follows links on the starting hostname and extracts a small set of common fields; a site may use client-side routing, links that require interaction, or a different content structure. Add narrowly scoped handling for the actual site rather than broadening the crawler indiscriminately.

Results differ from what you see manually

Pages may vary by browser engine, session state, location, timing, or other conditions. The code creates a fresh browser context without custom authentication or locale. If you have permission and the task requires it, configure those conditions deliberately and record them with the output. Rendering alone does not establish parity with a search engine crawler.

The crawler runs slowly or leaves a browser process behind

Browser rendering requires browser setup and control, unlike plain HTTP parsing, but the cited documentation does not establish a general speed ratio. Keep the crawl limited to pages that need JavaScript, avoid unnecessary navigation and extraction, and retain the guaranteed cleanup path. Do not add concurrency until you have considered site load, policy, and your own memory and process limits.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Frequently Asked Questions

Can a rendered Node.js crawler prove what Google indexed?

No. A browser crawler reports what its own browser session rendered. Google documents its crawling and JavaScript processing as a separate system; a local rendering result is not evidence of indexing or search visibility.

Should I use Playwright or Puppeteer if I already have one installed?

Either can be reasonable when it meets the target site’s browser needs. Crawlee supports both through separate crawler classes; weigh your team’s existing familiarity and the required browser engine rather than assuming one is universally compatible.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.