Do these 3 things before closing this tab:
1Repair Windows errors before they cause bigger problems2Fix the driver behind crashes, sound loss and screen glitches3Clear out junk files and repair common Windows errorsSome links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
To crawl an infinite-scroll page with Node.js, load it in a real browser, scroll the correct element, wait for a measurable change, and stop after bounded rounds. A plain HTTP request often returns only the initial HTML because JavaScript creates later items in the browser. The example below uses Playwright, shows a Puppeteer equivalent, handles nested scroll containers, deduplicates records, retries transient waits, and records why crawling stopped.
Why ordinary HTTP fetching misses infinite-scroll items
An infinite list usually starts with a small HTML shell. JavaScript then requests another batch when the viewport reaches a trigger. fetch(), Axios, or another HTTP client can retrieve the initial response, but it does not execute that page code or render the list. You have two practical choices:
- Use browser automation (Playwright or Puppeteer) to execute the page and scroll it.
- Find the data endpoint used by the page and call it directly, subject to the site’s terms, authentication, and rate limits.
Browser automation is the safer general solution when requests depend on cookies, client-side state, or changing JavaScript. A direct endpoint can be faster and lighter when it is documented or clearly permitted, but its request format may change.
Before you crawl: access, scope, and data hygiene
- Read
robots.txtand the site’s terms of service. Robots.txt tells crawlers which URLs a site asks them to access and helps manage traffic; it is not an access-control or security mechanism. - Confirm that authentication, personal-data collection, copyright, and retention practices are lawful for your use case.
- Use a reasonable rate, identify your crawler where appropriate, and honor server errors and rate-limit responses.
- Define the fields you need, a maximum number of pages or rounds, and a storage format before launching a long crawl.
Playwright: a bounded infinite-scroll crawler
Install Playwright and its browser, then save the following as an ES module. Replace the URL and selectors with the target site’s actual markup.
#1 Best Overall
npm install playwright
npx playwright install chromium
import { chromium } from 'playwright';
import { writeFile } from 'node:fs/promises';
const URL = 'https://example.com/list';
const ITEM = '.item';
const END = '.list-end, footer';
const MAX_ROUNDS = 40;
const MAX_STAGNANT = 3;
const browser = await chromium.launch({ headless: true });
const page = await browser.newPage({
viewport: { width: 1440, height: 1000 }
});
const seen = new Set();
const rows = [];
let stagnantRounds = 0;
let stopReason = 'maximum rounds reached';
try {
await page.goto(URL, { waitUntil: 'domcontentloaded', timeout: 45_000 });
await page.locator(ITEM).first().waitFor({ state: 'attached', timeout: 15_000 }).catch(() => {});
for (let round = 0; round < MAX_ROUNDS && stagnantRounds < MAX_STAGNANT; round++) {
const before = await page.locator(ITEM).count();
const end = page.locator(END).last();
if (await end.count()) {
await end.scrollIntoViewIfNeeded();
} else {
await page.mouse.wheel(0, 1200);
}
// Prefer a progress signal over a blind, long sleep.
await page.waitForTimeout(500);
const after = await page.locator(ITEM).count();
stagnantRounds = after === before ? stagnantRounds + 1 : 0;
const batch = await page.locator(ITEM).evaluateAll(nodes => nodes.map(node => ({
id: node.getAttribute('data-id') || node.querySelector('a')?.href || null,
text: node.textContent?.trim() || ''
})));
for (const row of batch) {
if (row.id && !seen.has(row.id)) {
seen.add(row.id);
rows.push(row);
}
}
const loadingVisible = await page.locator('.loading, [aria-busy="true"]').isVisible().catch(() => false);
if (!loadingVisible && after === before) stopReason = 'no new items after repeated rounds';
console.error({ round, before, after, unique: rows.length, loadingVisible });
}
if (stagnantRounds >= MAX_STAGNANT) stopReason = 'stagnation limit reached';
await writeFile('items.json', JSON.stringify({ stopReason, count: rows.length, rows }, null, 2));
await writeFile('final.html', await page.content());
} finally {
await browser.close();
}
What the loop is doing
- It navigates with a finite timeout and waits for the DOM rather than every image or analytics request.
- It counts items before scrolling, scrolls a bottom sentinel when one exists, and falls back to a mouse-wheel event.
- It waits briefly, then measures item count again. A count that does not increase contributes to a stagnation counter.
- It extracts each visible batch and deduplicates by a stable
data-idor canonical link. This matters when a virtualized list recycles DOM nodes. - It exits after 40 rounds or three consecutive rounds without progress, writes structured records, and saves the final rendered HTML for auditing.
Waiting for the right progress signal
The 500 ms delay is only a fallback. Prefer a signal tied to the page you are crawling:
- Wait for the item count to exceed its previous value.
- Wait for a loading spinner to become hidden.
- Wait for a “load more” button to disappear or become disabled.
- Wait for a specific network response, if the endpoint and response shape are stable.
- Compare document height or the target container’s
scrollHeight.
Use a short, bounded wait and retry on timeout. An unbounded sleep can make a crawler hang forever when a request fails.
Scrolling a nested container instead of the window
Many applications keep the page itself fixed and scroll a nested div. In that case, changing window.scrollY or using a page-level wheel may not trigger loading. Locate the container and change its scroll position:
The Tool Desk
Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →const container = page.locator('.results-scrollbox');
await container.evaluate(el => {
el.scrollTop = el.scrollHeight;
});
await page.waitForTimeout(500);
You can also scroll a child sentinel inside that container:
const sentinel = container.locator('.list-end').last();
if (await sentinel.count()) await sentinel.scrollIntoViewIfNeeded();
Inspect the page in developer tools to find the element whose scrollTop changes while you drag the list. That is the element your crawler must drive.
Puppeteer version
Puppeteer offers the same browser-driven approach. Its locator actions check that targets are in view and can generate the mouse-wheel scrolling needed by many lists.
Rank #3
import puppeteer from 'puppeteer';
import { writeFile } from 'node:fs/promises';
const browser = await puppeteer.launch({ headless: true });
const page = await browser.newPage();
const seen = new Set();
const rows = [];
let previousCount = 0;
let stagnant = 0;
try {
await page.goto('https://example.com/list', {
waitUntil: 'domcontentloaded',
timeout: 45_000
});
for (let round = 0; round < 40 && stagnant < 3; round++) {
const before = await page.locator('.item').count();
const end = page.locator('.list-end, footer').last();
if (await end.count()) {
await end.scroll({ scrollTop: 1000 });
} else {
await page.mouse.wheel({ deltaY: 1200 });
}
await new Promise(resolve => setTimeout(resolve, 500));
const current = await page.locator('.item').count();
stagnant = current === before ? stagnant + 1 : 0;
const batch = await page.locator('.item').evaluateAll(nodes => nodes.map(node => ({
id: node.getAttribute('data-id') || node.querySelector('a')?.href || null,
text: node.textContent?.trim() || ''
})));
for (const row of batch) {
if (row.id && !seen.has(row.id)) {
seen.add(row.id);
rows.push(row);
}
}
previousCount = current;
}
await writeFile('items.json', JSON.stringify(rows, null, 2));
const html = await page.content();
await writeFile('final.html', html);
} finally {
await browser.close();
}
Use Puppeteer’s page content when you need the complete rendered HTML after loading, but prefer structured extraction while the locators are available. As with Playwright, select a nested scroll container when the application does not scroll the document.
Free tools Windows power users keep installed
One-click scans. No signup required.
Stopping rules that prevent runaway crawls
Combine at least one hard bound with one content-based condition:
| Rule | How to measure it | When it helps |
|---|---|---|
| Maximum rounds | Stop after a configured number such as 40 | Protects against broken end markers and endless feeds |
| Maximum duration | Abort when elapsed time exceeds your job budget | Protects workers and queues from stalled pages |
| Repeated no-progress rounds | Compare item count, unique IDs, or height | Handles feeds that silently stop loading |
| End marker | Detect “no more results” or a disabled control | Cleanly ends finite result sets |
| Terminal response | Recognize an API response indicating the final page | Useful when network traffic is stable and permitted |
Do not rely on only the number of DOM nodes: virtualization can keep that number constant while replacing records. Track unique IDs or canonical URLs as well.
Retries, failures, and auditability
Navigation and timeout failures
Retry transient navigation errors with exponential backoff, but cap attempts. Record the URL, attempt number, elapsed time, and final error. A timeout can mean a slow origin, a blocked resource, or a page that never reaches the requested load state; try domcontentloaded and then wait for the specific list selector rather than waiting for every resource.
Loading that never finishes
Bound every selector wait and spinner wait. If a spinner remains visible after the bound, capture the current HTML and stop or retry instead of scrolling indefinitely.
Duplicate or missing records
Deduplicate with a stable server ID whenever possible. If the site has no ID, normalize and use the canonical link plus a content hash. Save raw HTML or response data alongside parsed records so you can diagnose selector changes and rerun parsing without recrawling.
Best Value
Network inspection
Listen for relevant responses to understand whether the page uses JSON, GraphQL, or HTML fragments. If you discover a direct endpoint, verify that calling it is allowed and preserve the browser’s required cookies, headers, pagination cursor, and rate limits. Do not assume an endpoint discovered in developer tools is public or stable.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Performance and reliability choices
- Reuse one browser process and create isolated pages or contexts for jobs; launching a new browser for every URL is expensive.
- Keep a realistic viewport. Some sites load different markup or omit items at mobile widths.
- Block nonessential images, ads, or fonts only when doing so does not change the list’s behavior or violate site expectations.
- Prefer event- or selector-based waits over large fixed delays, while retaining a small settle period for animations and debounced requests.
- Persist checkpoints after each successful round for long jobs. A restart can resume from stored IDs or a saved cursor.
- Log the final round, item count, unique count, duration, and termination reason. These fields make silent truncation visible.
Common problems and fixes
| Symptom | Likely cause | Fix |
|---|---|---|
| Only the first batch appears | JavaScript did not run, or the wrong element was scrolled | Use Playwright/Puppeteer, identify the actual scroll container, and wait for a count or network change |
| Scroll position changes but no request occurs | The trigger is a sentinel, threshold, or button rather than document height | Scroll the sentinel into view or click the load control, then wait for its specific result |
| Item count stays constant while content changes | Virtualized list recycles nodes | Extract every round and deduplicate by stable ID or URL |
| Headless mode sees fewer items | Responsive markup, consent UI, bot checks, or timing differences | Set a known viewport, handle consent where permitted, capture screenshots/HTML for diagnosis, and use bounded retries |
| Browser memory keeps growing | Pages or contexts are not closed, or the DOM grows without limit | Close pages promptly, checkpoint results, and consider a direct permitted endpoint |
| Requests return 403 or 429 | Access policy or rate limit | Stop, review robots.txt and terms, reduce concurrency, respect retry-after guidance, and obtain authorization if needed |
Or skip the browser setup
If your goal is a clean screenshot or PDF rather than extracting records, ScreenshotNeo provides a website screenshot API and MCP server. One request loads the URL and returns PNG, JPEG, WebP, or PDF. Before capture it accepts the cookie or consent banner like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each cleanup step can be turned off. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing status. AI agents can call its take_screenshot, get_page_info, and capture_pdf MCP tools.
See the full parameter list in the ScreenshotNeo documentation. cURL:
Windows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallOutdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchcurl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python:
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
The Free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000, and every feature is included on every plan. Create a free ScreenshotNeo account.
Frequently Asked Questions
Can I scroll an infinite page with only Node.js fetch?
Not when JavaScript creates the later items in the browser. Use Playwright or Puppeteer, or call a permitted data endpoint directly.
How do I know when an infinite list is finished?
Use a combination of a hard round or time limit and a progress signal such as unique-item count, hidden loading state, end marker, response cursor, or unchanged container height.
Why does window scrolling do nothing?
The page may scroll a nested container. Find the element whose scrollTop changes and scroll that locator or set its scrollTop directly.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

