October DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsClean PCRecommendedOne scan can reveal what keeps slowing WindowsLook for cleanup and repair opportunities.Run ScanOctober DealsAmazon USDeal season is back - check today's better picksAmazon US: current deals, useful picks and tech finds.See Picks×
Skip to content
dynamic pagination

How to Scrape Websites with Dynamic Pagination (Infinite Scroll, Load More, and JavaScript)

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Find the request that delivers the next batch of records, then crawl that request and follow the site’s own continuation signal. Use browser automation only when the request cannot be reproduced reliably or the data exists only after browser interaction. A click, infinite scroll event, or JavaScript render is merely the trigger; the records usually arrive through a normal HTTP response you can inspect.

What dynamic pagination changes

Traditional pagination puts a next link and the records in the initial HTML. Dynamic pagination changes one or both: JavaScript requests another page after a click or scroll, and the browser inserts the response into the document. Common implementations include numbered API requests, offset or cursor parameters, a “Load more” button, and infinite scroll.

Do not assume that the browser’s load event means the results are ready. Later requests can still be pending and the DOM can still be changing. Your first job is to identify the request or browser state that actually produces the next records.

1. Establish where the records come from

Compare source HTML with the rendered page

  1. Fetch the URL with an HTTP client and save the response.
  2. Open the same URL in a browser, choose View page source, and inspect embedded JSON, script tags, and ordinary links.
  3. Compare that source with the rendered DOM in developer tools. If the records are already in the response, parse the HTML directly; no browser is needed.

Scrapy’s guidance recommends locating the source data before attempting to reproduce browser behavior: Selecting dynamically-loaded content.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Observe exactly one pagination action

  1. Open Developer Tools, select Network, enable Preserve log, and clear the log.
  2. Filter to Fetch/XHR (also check Doc if navigation replaces the page).
  3. Click Next or Load more, or scroll only far enough to trigger one batch.
  4. Open requests whose response contains the new records. Scrapy documents this workflow in Using your browser’s Developer Tools for scraping.

Inspect the request method, URL, query string, request body, relevant headers, cookies, and response schema. In the response preview, look for an array of records plus fields such as next, next_url, cursor, offset, or has_next. Use the browser’s “Copy as cURL” feature as a starting point, then remove values that are not required.

2. Replay the data request with an HTTP crawler

Replaying a reproducible data request is normally simpler than rendering every page. Keep only the parameters and authentication state the endpoint requires, and verify that your crawler receives the same records as the browser.

Generic cursor-based Python crawler

This example uses placeholders because every site names its fields differently. Replace the URL, parameters, record key, and continuation key after inspecting one real response.

import time
import requests

START_URL = "https://example.com/api/products"
HEADERS = {"Accept": "application/json", "User-Agent": "catalog-crawler/1.0"}

session = requests.Session()
cursor = None
seen = set()

while True:
    params = {"limit": 50}
    if cursor:
        params["cursor"] = cursor

    response = session.get(START_URL, params=params, headers=HEADERS, timeout=30)
    response.raise_for_status()
    payload = response.json()

    records = payload.get("items", [])
    for record in records:
        key = record.get("id") or record.get("url")
        if key is not None and key not in seen:
            seen.add(key)
            print(record)

    next_cursor = payload.get("next_cursor")
    has_next = payload.get("has_next")
    if not next_cursor or has_next is False:
        break

    cursor = next_cursor
    time.sleep(1)

If the endpoint uses a page number, increment page until the response omits its next link. For an offset API, increase offset by the server’s returned page size; do not guess that every response is full. If it returns a complete next URL, request that URL verbatim (subject to your allowed-domain and security checks).

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Validate every response

  • Check HTTP status and content type before parsing JSON.
  • Confirm the expected record array exists and has the expected shape.
  • Log the page, cursor, URL, status, item count, and error body.
  • Deduplicate with a stable item ID or canonical URL.
  • Stop only when the source says there is no continuation: a missing next link or cursor, has_next: false, or an equivalent documented signal.

Scrapy’s overview and dynamic-content examples show both next-link and Boolean continuation patterns: Scrapy at a glance and its developer-tools tutorial. A fixed page limit is a safety guard, not proof that the crawl is complete.

3. Scrapy implementation for paginated endpoints

Once you know the endpoint, a Scrapy spider can schedule requests and preserve crawl state. This example follows a JSON next_url; adapt the keys to the target.

import scrapy

class ProductsSpider(scrapy.Spider):
    name = "products"
    start_urls = ["https://example.com/api/products?limit=50"]

    def parse(self, response):
        if response.status != 200:
            self.logger.error("HTTP %s: %s", response.status, response.url)
            return
        payload = response.json()
        for item in payload.get("items", []):
            yield {
                "id": item.get("id"),
                "name": item.get("name"),
                "url": item.get("url"),
            }
        next_url = payload.get("next_url")
        if next_url:
            yield response.follow(next_url, callback=self.parse)

For POST-based searches, reproduce the JSON body with FormRequest or JsonRequest. Preserve a session only when cookies are genuinely required, and do not copy short-lived browser tokens into a long-running crawler without a plan to refresh them.

4. Browser automation when HTTP replay is not enough

Use Playwright (or another headless browser) when the endpoint depends on client-side state, a complex interaction, a signed request you cannot reproduce, or output that exists only in the rendered DOM. Browser automation costs more setup and is sensitive to UI and timing changes, so keep it as a fallback rather than the default.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Infinite scroll with a page-specific readiness check

import asyncio
from playwright.async_api import async_playwright

async def main():
    async with async_playwright() as p:
        browser = await p.chromium.launch(headless=True)
        page = await browser.new_page()
        await page.goto("https://example.com/catalog", wait_until="domcontentloaded")

        previous_count = 0
        while True:
            current_count = await page.locator("article.product").count()
            if current_count == previous_count:
                # Trigger the site’s own scroll handler.
                await page.evaluate("window.scrollTo(0, document.body.scrollHeight)")
            try:
                await page.wait_for_function(
                    "old => document.querySelectorAll('article.product').length > old || "
                    "document.querySelector('[data-end]') !== null",
                    arg=current_count,
                    timeout=10000,
                )
            except Exception:
                # A timeout is meaningful only after checking the end marker/count.
                pass

            new_count = await page.locator("article.product").count()
            end = await page.locator("[data-end]").count()
            if end or new_count <= current_count:
                break
            previous_count = new_count

        items = await page.locator("article.product").evaluate_all(
            "els => els.map(e => ({name: e.querySelector('h2')?.textContent?.trim(), "
            "url: e.querySelector('a')?.href}))"
        )
        for item in items:
            print(item)
        await browser.close()

asyncio.run(main())

Replace selectors with ones tied to the target’s result item and end marker. Waiting for a changed result count, a newly visible item, or a known “no more results” marker is safer than sleeping for an arbitrary duration. Playwright explains navigation and readiness behavior in Navigations and the Page API reference; its documentation cautions that generic network-idle is not a universal readiness condition.

“Load more” button

Locate the button, click it, and wait for the item count to increase. Disable or break when the button is absent, disabled, or replaced by an end marker. Capture the response event only if the request is stable; otherwise extract the newly rendered items after the page-specific wait.

5. Choosing the right approach

Consideration Replay the data request Browser automation
Data access Parse the response that contains records Read rendered DOM or interact with controls
Best fit Stable, understandable endpoint Complex browser state or browser-only output
Complexity Investigation up front; fewer rendering steps Browser installation, timing, and UI maintenance
Typical failures Changed parameters, schema, or token requirements Selectors, timing, dialogs, and browser state
Verification Status, schema, item count, continuation fields Observed result changes and end condition

These are implementation trade-offs, not measured performance claims. A hybrid is often best: discover the request in a browser, then crawl it with Scrapy or an HTTP client.

6. Reliability, politeness, and data quality

  • Use bounded retries with exponential backoff for transient 429 and 5xx responses; do not retry validation errors forever.
  • Persist the cursor or page state so a process can resume without duplicating all earlier records.
  • Keep a crawl manifest containing timestamps, request parameters, response status, and parser version.
  • Fail visibly when the schema changes. An empty list is not automatically “no more pages.”
  • Throttle requests, honor published limits, and avoid parallelism that overloads the service.
  • Review robots.txt, terms, authentication requirements, and applicable law before crawling. RFC 9309 defines robots.txt as crawler access rules and explicitly states: “These rules are not a form of access authorization.” Read the standard at RFC 9309.

7. Common failures and fixes

The HTML has no records

Cause: records arrive through XHR/fetch after load. Fix: inspect Network while performing one pagination action and replay the response request.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The crawler always gets the first page

Cause: the cursor, offset, or body value is not being updated, or the cursor is tied to a session. Fix: log the outgoing URL/body and the returned continuation token; preserve the required session and send the token exactly as returned.

It stops early

Cause: code treats an empty batch, timeout, or HTTP error as completion. Fix: distinguish errors from valid end signals, validate the schema, and retry transient failures within a limit.

It loops forever or duplicates records

Cause: a server repeats a cursor or the UI re-renders existing items. Fix: keep a set of seen cursors and stable item keys; abort on a repeated continuation value and investigate.

Playwright sees stale results

Cause: the script waits for generic navigation or a fixed sleep instead of evidence that this batch changed. Fix: wait for an increased item count, a specific new item, or an end marker, and handle consent dialogs only when they block the target interaction.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

403, CAPTCHA, or login wall

Cause: access controls or authentication. Fix: stop and use an authorized account or documented API. Do not attempt to bypass a CAPTCHA or access restriction.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup

If your goal is a reliable visual capture of each paginated state—not extraction of the underlying records—ScreenshotNeo provides a one-request website screenshot API. It can accept cookie/consent banners before capture and remove more than 60 known consent platforms, newsletter popups, and chat widgets; each cleanup step can be disabled. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers report the page verdict and billing result.

It also offers an MCP server for Claude, Cursor, and other MCP clients, with take_screenshot, get_page_info, and capture_pdf tools. Features include full-page lazy-image capture, CSS-selector element capture, device and viewport controls, custom CSS/JavaScript, click and wait conditions, request blocking, headers/cookies, timezone and geolocation, resizing, chosen-TTL caching, signed links, asynchronous webhooks, bulk capture of up to 100 URLs per call, a usage API, and an OpenAPI specification.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com/catalog -o shot.webp

See the ScreenshotNeo API documentation for parameters and response headers. A free plan includes 1,000 shots per month with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

8. Cost and operating notes

HTTP replay generally uses fewer rendering steps, but investigation and endpoint maintenance take engineering time. Browsers require Chromium or another engine, consume more CPU and memory, and need explicit waits and selector maintenance. In either approach, cache only when the site permits it, avoid unnecessary recrawls, and record enough metadata to explain missing or changed records. No universal speed or success percentage applies because the endpoint, page size, rate limits, and browser behavior differ by site.

FAQ

How do I know whether a site uses infinite scroll?

Scroll once while watching Network. If a new request returns records and the DOM grows without a URL change, it is an infinite-scroll pattern; the same request-inspection method applies to “Load more.”

Should I scrape the API or the rendered HTML?

Use the structured request when it is reproducible and authorized. Use rendered HTML when browser state or interaction is essential, or when reproducing the request would be less reliable than observing the page.

Is a missing next link always the end?

It is an end signal only when that is how the target represents completion. Confirm the response schema and distinguish a missing field caused by an error from a valid terminal response.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Is robots.txt permission to scrape?

No. RFC 9309 describes it as crawler guidance and says its rules are not access authorization. Check terms, permissions, and law separately.

Frequently Asked Questions

Can I scrape a JavaScript site without Selenium?

Often. If Network inspection reveals a reproducible JSON or HTML request, an HTTP client or Scrapy is usually sufficient; use a browser when the required state cannot be reproduced.

What should I store to resume a crawl?

Persist the current page, cursor or next URL, stable item keys, and enough request metadata to verify that a resumed response belongs to the same crawl.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Read next

Recommended PC Tool
Recommended PC Tool
Crashes, No Sound, or Screen Glitches?Free driver scan
PC Slower Than It Used to Be?Free scan - under a minute

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.