Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

To collect every record, first identify how the page reveals the next batch: a numbered link, a Load more control, or a scroll-triggered request. Drive that exact mechanism with a real browser, verify that each action adds new records, deduplicate results, and stop only when the control is exhausted. The same approach works for mixed pages that combine pagination with scrolling.

Choose the traversal pattern before writing code

Inspect the page as a user would. Look for these continuation mechanisms:

  • Numbered pagination: links or a Next control navigate to distinct URLs. Each page must be fetched and parsed.
  • Load more: a button appends another batch without replacing the current list. The button may disappear, become disabled, or remain visible while returning no new records.
  • Infinite scroll: reaching the bottom of a list or scroll container triggers another batch. Scroll the element that actually owns the list, not necessarily the whole window.
  • Mixed pattern: each numbered page may still lazy-load records. Traverse pages, then perform the page’s scrolling routine before extracting.

Do not substitute one pattern for another. Repeated scrolling will not advance a conventional paginated URL, and clicking a button is not proof that the site uses infinite scroll.

A reliable scraping workflow

  1. Map the controls. Record the list selector, item selector, continuation control, and any loading indicator. In browser developer tools, watch whether a click or scroll appends nodes, replaces them, or changes the URL.
  2. Define progress. Count item nodes or collect a stable record key after every action. An action that leaves the count and keys unchanged is not progress.
  3. Wait for evidence, not a fixed guess. Wait for a new item, a loading indicator to disappear, a URL change, or a short network-idle period. A fixed delay alone is fragile on slow and fast connections.
  4. Run a small trial. Stop after a few batches, inspect the records, and confirm that fields are complete and duplicates are absent before running the full collection.
  5. Stop deterministically. End when there is no next link, the load-more control is disabled or gone, scrolling produces no new keys, or a configured safety limit is reached.

Install Playwright and a browser

The example below uses Python and Chromium. Create an isolated environment, install the package, then install the browser binary:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
python -m venv .venv
# macOS/Linux
source .venv/bin/activate
# Windows PowerShell: .venvScriptsActivate.ps1
pip install playwright
playwright install chromium

Save the script as scrape_lists.py. Replace the URL and selectors with those observed on your target page. The selectors in the example target common, semantic markup and are intentionally easy to change.

Scrape an infinite-scroll list

This routine scrolls the list container, waits for the item count to increase, extracts fields, and deduplicates by a record URL (falling back to the complete text when no link exists).

import asyncio
from playwright.async_api import async_playwright, TimeoutError as PlaywrightTimeoutError

URL = "https://example.com/catalog"
LIST = "main"
ITEM = "article.product"
KEY = "a.product-link"
MAX_ROUNDS = 500

async def records(page):
    rows = await page.locator(ITEM).all()
    out = []
    for row in rows:
        link = row.locator(KEY).first
        href = await link.get_attribute("href") if await link.count() else None
        title = (await row.inner_text()).strip()
        out.append({"key": href or title, "text": title, "url": href})
    return out

async def scrape_infinite(page):
    seen = {}
    box = page.locator(LIST).first
    for _ in range(MAX_ROUNDS):
        before = len(seen)
        for item in await records(page):
            seen[item["key"]] = item
        await box.evaluate("el => el.scrollTo(0, el.scrollHeight)")
        try:
            await page.wait_for_function(
                "([sel, n]) => document.querySelectorAll(sel).length > n",
                [ITEM, before], timeout=10000
            )
        except PlaywrightTimeoutError:
            # No new nodes: one final extraction, then finish.
            for item in await records(page):
                seen[item["key"]] = item
            break
    return list(seen.values())

async def main():
    async with async_playwright() as pw:
        browser = await pw.chromium.launch(headless=True)
        page = await browser.new_page()
        await page.goto(URL, wait_until="domcontentloaded", timeout=90000)
        data = await scrape_infinite(page)
        print(f"collected {len(data)} records")
        for row in data:
            print(row)
        await browser.close()

asyncio.run(main())

If the page scrolls the window rather than a nested container, replace the container evaluation with await page.evaluate("window.scrollTo(0, document.body.scrollHeight)"). If content is virtualized, old nodes may be removed; extract each batch immediately and keep your own seen dictionary instead of relying on the final DOM.

Click a Load more control until it is exhausted

Count records before each click and wait for that count to grow. The loop also handles controls that become disabled or return an empty batch.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
async def scrape_load_more(page):
    seen = {}
    button = page.get_by_role("button", name="Load more")
    for _ in range(500):
        for item in await records(page):
            seen[item["key"]] = item
        before = await page.locator(ITEM).count()
        if not await button.count() or not await button.is_visible() or await button.is_disabled():
            break
        await button.click()
        try:
            await page.wait_for_function(
                "([sel, n]) => document.querySelectorAll(sel).length > n",
                [ITEM, before], timeout=15000
            )
        except PlaywrightTimeoutError:
            # The click produced no additional records; do not click forever.
            break
    for item in await records(page):
        seen[item["key"]] = item
    return list(seen.values())

Some sites replace the button after each request. Resolve it again inside the loop with a stable locator (role, text, or a data attribute) rather than retaining a stale element handle. If an explicit spinner exists, wait for it to become hidden as an additional condition.

Walk numbered pagination and Next links

Pagination is usually safer when each page has a unique URL. Capture a page, discover the next link, and stop when it is absent, disabled, or points to a URL already visited.

from urllib.parse import urljoin

async def scrape_pages(page):
    seen, visited = {}, set()
    for _ in range(500):
        current = page.url
        if current in visited:
            break
        visited.add(current)
        # Handle pages that lazy-load within each numbered page.
        page_rows = await scrape_infinite(page)
        for item in page_rows:
            seen[item["key"]] = item
        nxt = page.locator("a[rel='next'], a.next").first
        if not await nxt.count() or not await nxt.is_visible():
            break
        href = await nxt.get_attribute("href")
        if not href or href.startswith("javascript:"):
            break
        target = urljoin(current, href)
        if target in visited:
            break
        await page.goto(target, wait_until="domcontentloaded", timeout=90000)
    return list(seen.values())

If the site exposes page numbers but no reliable Next link, collect the discovered href values from the pager, normalize them with urljoin, remove duplicates, and visit them in order. Preview at least two URLs to verify that the control actually changes the result set.

Handle mixed pagination and scrolling

For a mixed interface, use the pagination loop as the outer traversal and run a page-specific scrolling function after every navigation. Do not assume the scroll position or loaded state carries across pages. Wait for the first page’s records after navigation, scroll its list to exhaustion, extract, then follow the next URL.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Selectors, extraction, and deduplication

  • Prefer stable attributes such as data-testid, semantic roles, or product IDs over generated class names.
  • Extract a canonical key: an ID in a data attribute, a normalized absolute URL, or a hash of fields that uniquely identify a record.
  • Normalize whitespace and URLs before deduplicating. Keep the first complete record, or merge later fields when cards are progressively enhanced.
  • Save raw HTML or JSON responses for a small sample. This makes selector changes and missing-field bugs diagnosable.

Performance, reliability, and cost controls

  • Use one browser context per account or proxy policy and one page per independent crawl; excessive parallel tabs can trigger throttling and increase memory use.
  • Set explicit navigation and action timeouts, but allow a longer timeout for the first page and known slow batches.
  • Persist checkpoints (last URL, item keys, and batch number) so a failure can resume without reprocessing everything.
  • Cap rounds and records. A broken selector or looping pager should fail safely instead of running indefinitely.
  • Block unnecessary images, fonts, or analytics only when doing so does not change the list’s behavior. Some sites load records through scripts tied to those resources.
  • Respect the site’s access rules, authentication requirements, and applicable laws. Do not bypass a CAPTCHA or bot challenge; treat it as a stop condition and review access with the site owner.

Troubleshooting common failures

The count never increases

You may be scrolling the wrong element, clicking a hidden duplicate button, or waiting on the wrong selector. Inspect which element’s scrollHeight changes, target the visible control, and wait for a specific new item or response.

Records repeat across batches

Virtualized lists can recycle DOM nodes, while APIs may overlap pages. Deduplicate by a stable ID or canonical URL and extract each batch before the next scroll.

The script finishes too early

A timeout can occur while a request is still pending. Wait for the site’s loading indicator to disappear or for a known item key to appear, then use a bounded retry before declaring exhaustion.

Pagination loops forever

Track visited URLs and stop when the next target is missing, identical to the current URL, or already visited. Also check that query parameters are not being reordered into equivalent URLs.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Headless output differs from a normal browser

Compare a headed run, viewport, locale, cookies, and authentication state. If a bot check or consent dialog blocks the list, solve the legitimate access requirement rather than attempting to evade it.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Search visibility for site owners

Google generally discovers pages through URLs in anchor href attributes. It does not click buttons and generally does not trigger JavaScript actions that require user interaction to reveal another page. If important content is behind only a Load more button or scroll event, provide crawlable sequential links and a unique URL for each page. You can keep the richer interface for people while exposing those links to crawlers.

Or skip the browser setup

If you need a visual capture of the rendered state rather than structured records, ScreenshotNeo returns a screenshot or PDF from one request. It accepts consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be disabled. Only clean shots are billed: bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are free, and response headers report the page verdict and billing status. Its MCP server exposes take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients. Every plan includes features such as full-page lazy-image loading, element capture, custom CSS and JavaScript, waits, request blocking, cookies and headers, device presets, PDF controls, caching, signed links, async webhooks, bulk capture, and a usage API.

See the ScreenshotNeo API documentation for parameters and authentication. cURL:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account to try it.

FAQ

Can I scrape an infinite list with plain HTTP requests?

Only when you can identify and call the underlying data endpoint legitimately. If records appear only after browser JavaScript runs, Playwright is the dependable first implementation.

How do I know I collected everything?

Log each batch’s count and keys, verify the final control state, and compare a limited run with what the page reports. A stable stop condition is stronger evidence than an arbitrary number of scrolls.

Should I store HTML or structured data?

Store structured fields for analysis and a small raw sample for auditing. Keeping both helps you detect layout changes without retaining unnecessary page content.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.