Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

To crawl an infinite-scroll page with Node.js, load it in a real browser, scroll the correct element, wait for a measurable change, and stop after bounded rounds. A plain HTTP request often returns only the initial HTML because JavaScript creates later items in the browser. The example below uses Playwright, shows a Puppeteer equivalent, handles nested scroll containers, deduplicates records, retries transient waits, and records why crawling stopped.

Why ordinary HTTP fetching misses infinite-scroll items

An infinite list usually starts with a small HTML shell. JavaScript then requests another batch when the viewport reaches a trigger. fetch(), Axios, or another HTTP client can retrieve the initial response, but it does not execute that page code or render the list. You have two practical choices:

  • Use browser automation (Playwright or Puppeteer) to execute the page and scroll it.
  • Find the data endpoint used by the page and call it directly, subject to the site’s terms, authentication, and rate limits.

Browser automation is the safer general solution when requests depend on cookies, client-side state, or changing JavaScript. A direct endpoint can be faster and lighter when it is documented or clearly permitted, but its request format may change.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Before you crawl: access, scope, and data hygiene

  • Read robots.txt and the site’s terms of service. Robots.txt tells crawlers which URLs a site asks them to access and helps manage traffic; it is not an access-control or security mechanism.
  • Confirm that authentication, personal-data collection, copyright, and retention practices are lawful for your use case.
  • Use a reasonable rate, identify your crawler where appropriate, and honor server errors and rate-limit responses.
  • Define the fields you need, a maximum number of pages or rounds, and a storage format before launching a long crawl.

Playwright: a bounded infinite-scroll crawler

Install Playwright and its browser, then save the following as an ES module. Replace the URL and selectors with the target site’s actual markup.

npm install playwright
npx playwright install chromium
import { chromium } from 'playwright';
import { writeFile } from 'node:fs/promises';

const URL = 'https://example.com/list';
const ITEM = '.item';
const END = '.list-end, footer';
const MAX_ROUNDS = 40;
const MAX_STAGNANT = 3;

const browser = await chromium.launch({ headless: true });
const page = await browser.newPage({
  viewport: { width: 1440, height: 1000 }
});

const seen = new Set();
const rows = [];
let stagnantRounds = 0;
let stopReason = 'maximum rounds reached';

try {
  await page.goto(URL, { waitUntil: 'domcontentloaded', timeout: 45_000 });
  await page.locator(ITEM).first().waitFor({ state: 'attached', timeout: 15_000 }).catch(() => {});

  for (let round = 0; round < MAX_ROUNDS && stagnantRounds < MAX_STAGNANT; round++) {
    const before = await page.locator(ITEM).count();
    const end = page.locator(END).last();

    if (await end.count()) {
      await end.scrollIntoViewIfNeeded();
    } else {
      await page.mouse.wheel(0, 1200);
    }

    // Prefer a progress signal over a blind, long sleep.
    await page.waitForTimeout(500);
    const after = await page.locator(ITEM).count();
    stagnantRounds = after === before ? stagnantRounds + 1 : 0;

    const batch = await page.locator(ITEM).evaluateAll(nodes => nodes.map(node => ({
      id: node.getAttribute('data-id') || node.querySelector('a')?.href || null,
      text: node.textContent?.trim() || ''
    })));

    for (const row of batch) {
      if (row.id && !seen.has(row.id)) {
        seen.add(row.id);
        rows.push(row);
      }
    }

    const loadingVisible = await page.locator('.loading, [aria-busy="true"]').isVisible().catch(() => false);
    if (!loadingVisible && after === before) stopReason = 'no new items after repeated rounds';
    console.error({ round, before, after, unique: rows.length, loadingVisible });
  }

  if (stagnantRounds >= MAX_STAGNANT) stopReason = 'stagnation limit reached';
  await writeFile('items.json', JSON.stringify({ stopReason, count: rows.length, rows }, null, 2));
  await writeFile('final.html', await page.content());
} finally {
  await browser.close();
}

What the loop is doing

  1. It navigates with a finite timeout and waits for the DOM rather than every image or analytics request.
  2. It counts items before scrolling, scrolls a bottom sentinel when one exists, and falls back to a mouse-wheel event.
  3. It waits briefly, then measures item count again. A count that does not increase contributes to a stagnation counter.
  4. It extracts each visible batch and deduplicates by a stable data-id or canonical link. This matters when a virtualized list recycles DOM nodes.
  5. It exits after 40 rounds or three consecutive rounds without progress, writes structured records, and saves the final rendered HTML for auditing.

Waiting for the right progress signal

The 500 ms delay is only a fallback. Prefer a signal tied to the page you are crawling:

  • Wait for the item count to exceed its previous value.
  • Wait for a loading spinner to become hidden.
  • Wait for a “load more” button to disappear or become disabled.
  • Wait for a specific network response, if the endpoint and response shape are stable.
  • Compare document height or the target container’s scrollHeight.

Use a short, bounded wait and retry on timeout. An unbounded sleep can make a crawler hang forever when a request fails.

Scrolling a nested container instead of the window

Many applications keep the page itself fixed and scroll a nested div. In that case, changing window.scrollY or using a page-level wheel may not trigger loading. Locate the container and change its scroll position:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
const container = page.locator('.results-scrollbox');
await container.evaluate(el => {
  el.scrollTop = el.scrollHeight;
});
await page.waitForTimeout(500);

You can also scroll a child sentinel inside that container:

const sentinel = container.locator('.list-end').last();
if (await sentinel.count()) await sentinel.scrollIntoViewIfNeeded();

Inspect the page in developer tools to find the element whose scrollTop changes while you drag the list. That is the element your crawler must drive.

Puppeteer version

Puppeteer offers the same browser-driven approach. Its locator actions check that targets are in view and can generate the mouse-wheel scrolling needed by many lists.

import puppeteer from 'puppeteer';
import { writeFile } from 'node:fs/promises';

const browser = await puppeteer.launch({ headless: true });
const page = await browser.newPage();
const seen = new Set();
const rows = [];
let previousCount = 0;
let stagnant = 0;

try {
  await page.goto('https://example.com/list', {
    waitUntil: 'domcontentloaded',
    timeout: 45_000
  });

  for (let round = 0; round < 40 && stagnant < 3; round++) {
    const before = await page.locator('.item').count();
    const end = page.locator('.list-end, footer').last();
    if (await end.count()) {
      await end.scroll({ scrollTop: 1000 });
    } else {
      await page.mouse.wheel({ deltaY: 1200 });
    }

    await new Promise(resolve => setTimeout(resolve, 500));
    const current = await page.locator('.item').count();
    stagnant = current === before ? stagnant + 1 : 0;

    const batch = await page.locator('.item').evaluateAll(nodes => nodes.map(node => ({
      id: node.getAttribute('data-id') || node.querySelector('a')?.href || null,
      text: node.textContent?.trim() || ''
    })));
    for (const row of batch) {
      if (row.id && !seen.has(row.id)) {
        seen.add(row.id);
        rows.push(row);
      }
    }
    previousCount = current;
  }

  await writeFile('items.json', JSON.stringify(rows, null, 2));
  const html = await page.content();
  await writeFile('final.html', html);
} finally {
  await browser.close();
}

Use Puppeteer’s page content when you need the complete rendered HTML after loading, but prefer structured extraction while the locators are available. As with Playwright, select a nested scroll container when the application does not scroll the document.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Stopping rules that prevent runaway crawls

Combine at least one hard bound with one content-based condition:

Rule How to measure it When it helps
Maximum rounds Stop after a configured number such as 40 Protects against broken end markers and endless feeds
Maximum duration Abort when elapsed time exceeds your job budget Protects workers and queues from stalled pages
Repeated no-progress rounds Compare item count, unique IDs, or height Handles feeds that silently stop loading
End marker Detect “no more results” or a disabled control Cleanly ends finite result sets
Terminal response Recognize an API response indicating the final page Useful when network traffic is stable and permitted

Do not rely on only the number of DOM nodes: virtualization can keep that number constant while replacing records. Track unique IDs or canonical URLs as well.

Retries, failures, and auditability

Navigation and timeout failures

Retry transient navigation errors with exponential backoff, but cap attempts. Record the URL, attempt number, elapsed time, and final error. A timeout can mean a slow origin, a blocked resource, or a page that never reaches the requested load state; try domcontentloaded and then wait for the specific list selector rather than waiting for every resource.

Loading that never finishes

Bound every selector wait and spinner wait. If a spinner remains visible after the bound, capture the current HTML and stop or retry instead of scrolling indefinitely.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Duplicate or missing records

Deduplicate with a stable server ID whenever possible. If the site has no ID, normalize and use the canonical link plus a content hash. Save raw HTML or response data alongside parsed records so you can diagnose selector changes and rerun parsing without recrawling.

Network inspection

Listen for relevant responses to understand whether the page uses JSON, GraphQL, or HTML fragments. If you discover a direct endpoint, verify that calling it is allowed and preserve the browser’s required cookies, headers, pagination cursor, and rate limits. Do not assume an endpoint discovered in developer tools is public or stable.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Performance and reliability choices

  • Reuse one browser process and create isolated pages or contexts for jobs; launching a new browser for every URL is expensive.
  • Keep a realistic viewport. Some sites load different markup or omit items at mobile widths.
  • Block nonessential images, ads, or fonts only when doing so does not change the list’s behavior or violate site expectations.
  • Prefer event- or selector-based waits over large fixed delays, while retaining a small settle period for animations and debounced requests.
  • Persist checkpoints after each successful round for long jobs. A restart can resume from stored IDs or a saved cursor.
  • Log the final round, item count, unique count, duration, and termination reason. These fields make silent truncation visible.

Common problems and fixes

Symptom Likely cause Fix
Only the first batch appears JavaScript did not run, or the wrong element was scrolled Use Playwright/Puppeteer, identify the actual scroll container, and wait for a count or network change
Scroll position changes but no request occurs The trigger is a sentinel, threshold, or button rather than document height Scroll the sentinel into view or click the load control, then wait for its specific result
Item count stays constant while content changes Virtualized list recycles nodes Extract every round and deduplicate by stable ID or URL
Headless mode sees fewer items Responsive markup, consent UI, bot checks, or timing differences Set a known viewport, handle consent where permitted, capture screenshots/HTML for diagnosis, and use bounded retries
Browser memory keeps growing Pages or contexts are not closed, or the DOM grows without limit Close pages promptly, checkpoint results, and consider a direct permitted endpoint
Requests return 403 or 429 Access policy or rate limit Stop, review robots.txt and terms, reduce concurrency, respect retry-after guidance, and obtain authorization if needed

Or skip the browser setup

If your goal is a clean screenshot or PDF rather than extracting records, ScreenshotNeo provides a website screenshot API and MCP server. One request loads the URL and returns PNG, JPEG, WebP, or PDF. Before capture it accepts the cookie or consent banner like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each cleanup step can be turned off. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing status. AI agents can call its take_screenshot, get_page_info, and capture_pdf MCP tools.

See the full parameter list in the ScreenshotNeo documentation. cURL:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

The Free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000, and every feature is included on every plan. Create a free ScreenshotNeo account.

Frequently Asked Questions

Can I scroll an infinite page with only Node.js fetch?

Not when JavaScript creates the later items in the browser. Use Playwright or Puppeteer, or call a permitted data endpoint directly.

How do I know when an infinite list is finished?

Use a combination of a hard round or time limit and a progress signal such as unique-item count, hidden loading state, end marker, response cursor, or unchanged container height.

Why does window scrolling do nothing?

The page may scroll a nested container. Find the element whose scrollTop changes and scroll that locator or set its scrollTop directly.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.