Driver FixRecommendedSound, Wi-Fi or graphics acting up? Check drivers firstFind missing or outdated drivers fast.Check DriversOctober DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsPC HealthRecommendedCrashes, freezes, slowdowns? Check your PC nowSpot repairable issues before they interrupt work.Check PC×
Skip to content
APIs

How to Scrape Paginated Lists, Load More Buttons, and Infinite Scroll

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The dependable way to scrape a long list is to identify how the site loads records, inspect the network request behind that interaction, and replay that request when possible. Use page, offset, or cursor parameters for ordinary pagination; repeat the Load more request until it is exhausted; and use a browser only when the data cannot be reproduced with HTTP. In every method, preserve filters, deduplicate by a stable record ID, record progress, and enforce hard limits.

Identify the loading pattern before writing a scraper

These interfaces look similar to a reader but require different traversal logic. Classify the page first:

Pattern What the user sees Typical scraper action
Numbered pages or Next A page number, a Next link, or a URL such as ?page=3 Follow the next URL or increment a documented page, offset, or cursor parameter.
Load more A button appends another batch while the URL may stay unchanged Inspect the button request, then repeat it with the returned cursor or offset until the button is gone, disabled, or produces no new records.
Infinite scroll More rows appear when the window or a list container reaches its end Reproduce the request fired by the scroll sentinel, or scroll the correct container in a browser and wait for the item count to increase.

Google treats pagination, Load more, and infinite scroll as separate interaction patterns. A page that looks like infinite scroll may actually have a hidden API with a simple cursor; finding that API is usually the shortest path to a complete export.

Inspect Network traffic first

  1. Open the page in a desktop browser and launch DevTools with Network.
  2. Enable “Preserve log,” clear existing entries, and filter to fetch or XHR.
  3. Change a filter or click Next/Load more, or scroll until new records appear.
  4. Open the request that returned the new records. Note its method, URL, query string, request body, response format, cursor or offset, and any required headers, cookies, or CSRF token.
  5. Use “Copy as cURL” and replay it outside the browser. Compare the returned records with those visible in the page before building a larger loop.

The underlying JSON request is often more stable than CSS selectors because it returns structured records without layout, animation, or virtualized-row issues. Reproduce only state you are authorized to use: filters, pagination parameters, and tokens needed for the same session. Do not assume a browser cookie or CSRF token is permanent; log when it expires and obtain a fresh one through the normal site flow.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Scrape ordinary pagination with HTTP

Offset and limit parameters

Many APIs accept offset and limit, and some report a total count. The safe loop advances by the number of records actually returned, not by a guessed page size. It stops on an empty page or when the reported total is reached.

import json
import time
from pathlib import Path
import requests

BASE = "https://example.com/api/products"
params = {"limit": 100, "category": "laptops", "sort": "updated"}
seen = set()
rows = []
offset = 0
reported_total = None

while True:
    query = {**params, "offset": offset}
    response = requests.get(BASE, params=query, timeout=30)
    response.raise_for_status()
    payload = response.json()
    batch = payload.get("items", [])
    if reported_total is None:
        reported_total = payload.get("total")

    if not batch:
        break

    for item in batch:
        key = item.get("id") or item.get("url")
        if key is not None and key not in seen:
            seen.add(key)
            rows.append(item)

    offset += len(batch)
    Path("progress.json").write_text(json.dumps({
        "offset": offset, "count": len(rows)
    }))
    if reported_total is not None and offset >= reported_total:
        break
    time.sleep(0.2)

Path("products.json").write_text(json.dumps(rows, ensure_ascii=False, indent=2))
print(f"saved {len(rows)} records")

Keep every filter in each request. Dropping a category, search term, sort order, or tenant parameter while incrementing offset can silently mix unrelated records. If the service uses one-based pages, send page=1, then increment until the response is empty or the Next link is absent. If it returns a cursor, send the response’s next cursor exactly as provided rather than constructing one.

Cursor pagination

A cursor is normally opaque. Treat it as a value to store and replay, not as a number to modify. Stop when the response has no next cursor, repeats a cursor already seen, or returns no records.

cursor = None
seen_cursors = set()
while True:
    query = {"limit": 100}
    if cursor:
        query["cursor"] = cursor
    data = requests.get(BASE, params=query, timeout=30).json()
    batch = data.get("items", [])
    # process and deduplicate batch here
    next_cursor = data.get("next_cursor")
    if not batch or not next_cursor or next_cursor in seen_cursors:
        break
    seen_cursors.add(next_cursor)
    cursor = next_cursor

Replay with cURL or Node.js

Once DevTools gives you a working request, cURL is useful for checking parameters independently of your application code:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
curl -G 'https://example.com/api/products' 
  -H 'Accept: application/json' 
  --data-urlencode 'limit=100' 
  --data-urlencode 'offset=0' 
  --data-urlencode 'category=laptops'

The equivalent Node.js request preserves the response as JSON for further processing:

const url = new URL('https://example.com/api/products');
url.search = new URLSearchParams({
  limit: '100', offset: '0', category: 'laptops'
});
const res = await fetch(url, { headers: { Accept: 'application/json' } });
if (!res.ok) throw new Error(`${res.status} ${res.statusText}`);
const payload = await res.json();
console.log(payload.items.length);

Scrape a Load more button

Prefer the button’s request

Click the control once with Network recording enabled. Record whether it uses GET or POST, where the offset or cursor is sent, and how the response signals exhaustion. Some sites return the next cursor in JSON; others embed it in a response header or in a data attribute.

A request-based loop follows the same rules as cursor pagination: append only unseen records, preserve all filters, detect a repeated cursor, and stop when the response is empty or explicitly complete. A Load more button can remain visible after the final batch, so “button still exists” is not proof that more data is available.

Use Playwright when the request is difficult to reproduce

Browser automation is appropriate when a token is generated in JavaScript, the list is assembled only in the DOM, or the click performs browser-only work. Use a semantic role or test ID rather than a fragile class name, and wait for a measurable change.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
from playwright.sync_api import sync_playwright, TimeoutError as PlaywrightTimeoutError

with sync_playwright() as p:
    browser = p.chromium.launch(headless=True)
    page = browser.new_page()
    page.goto("https://example.com/catalog", wait_until="domcontentloaded")
    seen = set()
    records = []

    while len(records) < 10000:
        cards = page.locator('[data-testid="product-card"]')
        before = cards.count()
        for i in range(before):
            card = cards.nth(i)
            key = card.get_attribute("data-id") or card.get_attribute("data-url")
            if key and key not in seen:
                seen.add(key)
                records.append({"id": key, "text": card.inner_text()})

        button = page.get_by_role("button", name="Load more")
        if await button.count() == 0 or not await button.is_enabled():
            break
        await button.click()
        try:
            page.wait_for_function(
                "(n) => document.querySelectorAll('[data-testid=product-card]').length > n",
                arg=before, timeout=10000
            )
        except PlaywrightTimeoutError:
            # No growth: avoid an endless loop and inspect the last response/log.
            break

    browser.close()
    print(f"saved {len(records)} records")

For production, wait for the specific response triggered by the click when possible. Waiting only a fixed delay is slower on fast runs and still flaky on slow ones. Save the last cursor, URL, and count after each successful batch so a crash can resume without starting over.

Scrape infinite scroll safely

Find the real scroll container

The page window is not always the scrolling element. A table may scroll inside a div, or an intersection-observer sentinel may trigger requests before the last row is visible. In DevTools, inspect which element’s scroll height changes and which request appears when its bottom enters view.

from playwright.sync_api import sync_playwright

with sync_playwright() as p:
    browser = p.chromium.launch(headless=True)
    page = browser.new_page()
    page.goto("https://example.com/feed", wait_until="networkidle")
    seen = set()
    records = []
    stagnant = 0

    for iteration in range(200):
        items = page.locator('[data-testid="feed-item"]')
        before = items.count()
        for i in range(before):
            node = items.nth(i)
            key = node.get_attribute("data-id")
            if key and key not in seen:
                seen.add(key)
                records.append({"id": key, "text": node.inner_text()})

        await items.last.scroll_into_view_if_needed()
        page.wait_for_timeout(500)
        after = items.count()
        stagnant = stagnant + 1 if after <= before else 0
        if stagnant >= 3 or len(records) >= 10000:
            break

    browser.close()

Replace the locator with the actual list item and scroll the container itself when necessary. Use several guards: a maximum iteration or item count, a wall-clock deadline, repeated cursors, unchanged item count, a missing sentinel, and a disabled control. Without these limits, a broken endpoint can leave a worker scrolling forever.

Deduplication, completeness, and resumability

  • Choose a stable key. Prefer an immutable record ID. A canonical URL is a fallback; visible text is unsafe because titles can repeat or change.
  • Keep first-seen order. Store IDs in a set and append a record only once, while retaining the source page or cursor for auditing.
  • Log progress. Record request parameters, response status, batch size, new-record count, cursor, and timestamp after each batch.
  • Validate totals. Compare your unique count with an API’s reported total, but treat totals as a changing snapshot rather than a guarantee when records are added or deleted during the run.
  • Resume deliberately. Persist the last successful cursor or offset and the IDs already written. For offset pagination, a changing dataset can shift rows; cursor pagination or a server-side snapshot is safer when available.

Choose the right implementation

Approach Best fit Trade-offs
Direct HTTP/API replay A JSON/XHR request is visible in Network tools Fast and structured, but you must reproduce pagination state, headers, cookies, and tokens.
Scrapy request spider Many URLs, retries, concurrency, and structured pipelines Excellent for HTTP workflows; it does not execute page JavaScript by itself.
Playwright or another headless browser Browser-only rendering, clicks, scrolling, or token generation Uses more CPU and memory and is slower than replaying JSON.
Hybrid Scrapy plus Playwright The initial page needs a browser but subsequent pages use an API Can reduce browser work, but requires coordination of cookies, tokens, and state.

Start with HTTP replay, then add a browser only for the part that genuinely requires it. This keeps selectors, timing, and browser resources out of the largest part of the crawl.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Performance, reliability, and responsible operation

  • Request the largest permitted page size, but respect documented limits and rate limits; a larger batch is not automatically safer.
  • Reuse a session so negotiated cookies and connection pools persist. Add bounded retries with backoff for transient 429 and 5xx responses, and never retry a malformed request indefinitely.
  • Set connect and read timeouts separately where your HTTP client supports them. A timeout should mark a batch incomplete, not advance the cursor.
  • Cache successful responses during development so selector and parser changes do not repeatedly hit the site.
  • Capture response status, content type, and a small diagnostic sample. A 200 response containing a login page or bot challenge is not a successful data batch.
  • Keep credentials and tokens out of logs. Restrict collection to data you are permitted to access, and follow the site’s terms, access controls, and applicable law.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Common failures and fixes

The scraper gets HTML instead of JSON

You probably lost a required cookie, authorization header, CSRF token, or content negotiation header. Copy the working request, compare headers, and confirm that the session is authenticated. Check for redirects to a login or challenge page.

Every Load more click returns the first batch

The next offset or cursor may be in the response, a hidden field, or a request body rather than the URL. Inspect the second request and send its updated state; do not keep clicking the same stale request.

The loop stops early

Check whether you stopped on a short page even though the API reports more records, discarded records whose IDs were missing, or dropped a filter on later requests. Log raw batch sizes and unique counts separately.

Infinite scroll never finishes

Virtualized lists may remove old DOM nodes, so DOM count can stay constant while new records arrive. Count stable IDs or intercept the network response instead. Also verify that you are scrolling the list container, not the window.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Playwright times out after a click

The click may be blocked by an overlay, the control may be outside the viewport, or the site may require a longer response wait. Use a role or test-ID locator, scroll it into view, wait for the specific response or item-count increase, and capture a screenshot or trace on failure.

Records are duplicated or missing

Deduplicate by an immutable key and inspect cursor repetition. Offset pagination over a live, changing dataset can skip or repeat rows; prefer a cursor or snapshot, or record the crawl time and accept that the result is not a transactionally consistent snapshot.

Or skip the browser setup

If your goal is a visual capture of each rendered page rather than structured record extraction, ScreenshotNeo provides a one-request screenshot API at ScreenshotNeo. It accepts consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each cleanup step can be disabled. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing result. Its MCP server lets Claude, Cursor, or another MCP client call take_screenshot, get_page_info, and capture_pdf.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com/list -o shot.webp

See the ScreenshotNeo API documentation for the 63 options, including full-page capture with lazy images loaded, CSS-selector element capture, dark mode, device presets and custom viewports, retina scale, PDF paper and page controls, custom CSS and JavaScript, pre-capture clicks, hidden selectors, selector/delay/network-idle waits, request and resource blocking, custom headers/cookies/user agents/Authorization, timezone and geolocation, transparent backgrounds, resizing, chosen cache TTLs, signed image links, asynchronous jobs with signed webhooks, bulk capture of up to 100 URLs per call, usage data, and an OpenAPI specification. Parameter names used by other screenshot APIs also work, which can simplify migration.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://example.com/list"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://example.com/list' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

The Free plan includes 1,000 shots per month without a card; paid plans start at $5 for 3,000 shots, with every feature on every plan. Create a free ScreenshotNeo account.

Frequently Asked Questions

How do I preserve the order of records after deduplication?

Keep an insertion-ordered collection keyed by the stable record ID. Ignore later duplicates unless your project explicitly needs the newest version, in which case update the existing value while retaining its original position.

What should I do when records change during a long crawl?

Treat the result as a time-bounded collection, record the crawl start and end times, and prefer a cursor or server-provided snapshot when consistency matters. Offset pagination on a live list can shift as rows are inserted or deleted.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Read next

Recommended PC Tool
Recommended PC Tool
Crashes, No Sound, or Screen Glitches?Free driver scan
PC Slower Than It Used to Be?Free scan - under a minute

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.