Do these 3 things before closing this tab:
1Repair Windows errors before they cause bigger problems2Fix the driver behind crashes, sound loss and screen glitches3Clear out junk files and repair common Windows errorsFind the request that delivers the next batch of records, then crawl that request and follow the site’s own continuation signal. Use browser automation only when the request cannot be reproduced reliably or the data exists only after browser interaction. A click, infinite scroll event, or JavaScript render is merely the trigger; the records usually arrive through a normal HTTP response you can inspect.
What dynamic pagination changes
Traditional pagination puts a next link and the records in the initial HTML. Dynamic pagination changes one or both: JavaScript requests another page after a click or scroll, and the browser inserts the response into the document. Common implementations include numbered API requests, offset or cursor parameters, a “Load more” button, and infinite scroll.
Do not assume that the browser’s load event means the results are ready. Later requests can still be pending and the DOM can still be changing. Your first job is to identify the request or browser state that actually produces the next records.
1. Establish where the records come from
Compare source HTML with the rendered page
- Fetch the URL with an HTTP client and save the response.
- Open the same URL in a browser, choose View page source, and inspect embedded JSON, script tags, and ordinary links.
- Compare that source with the rendered DOM in developer tools. If the records are already in the response, parse the HTML directly; no browser is needed.
Scrapy’s guidance recommends locating the source data before attempting to reproduce browser behavior: Selecting dynamically-loaded content.
Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchWindows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstall#1 Best Overall
Observe exactly one pagination action
- Open Developer Tools, select Network, enable Preserve log, and clear the log.
- Filter to Fetch/XHR (also check Doc if navigation replaces the page).
- Click Next or Load more, or scroll only far enough to trigger one batch.
- Open requests whose response contains the new records. Scrapy documents this workflow in Using your browser’s Developer Tools for scraping.
Inspect the request method, URL, query string, request body, relevant headers, cookies, and response schema. In the response preview, look for an array of records plus fields such as next, next_url, cursor, offset, or has_next. Use the browser’s “Copy as cURL” feature as a starting point, then remove values that are not required.
2. Replay the data request with an HTTP crawler
Replaying a reproducible data request is normally simpler than rendering every page. Keep only the parameters and authentication state the endpoint requires, and verify that your crawler receives the same records as the browser.
Generic cursor-based Python crawler
This example uses placeholders because every site names its fields differently. Replace the URL, parameters, record key, and continuation key after inspecting one real response.
import time
import requests
START_URL = "https://example.com/api/products"
HEADERS = {"Accept": "application/json", "User-Agent": "catalog-crawler/1.0"}
session = requests.Session()
cursor = None
seen = set()
while True:
params = {"limit": 50}
if cursor:
params["cursor"] = cursor
response = session.get(START_URL, params=params, headers=HEADERS, timeout=30)
response.raise_for_status()
payload = response.json()
records = payload.get("items", [])
for record in records:
key = record.get("id") or record.get("url")
if key is not None and key not in seen:
seen.add(key)
print(record)
next_cursor = payload.get("next_cursor")
has_next = payload.get("has_next")
if not next_cursor or has_next is False:
break
cursor = next_cursor
time.sleep(1)
If the endpoint uses a page number, increment page until the response omits its next link. For an offset API, increase offset by the server’s returned page size; do not guess that every response is full. If it returns a complete next URL, request that URL verbatim (subject to your allowed-domain and security checks).
Free tools Windows power users keep installed
One-click scans. No signup required.
Validate every response
- Check HTTP status and content type before parsing JSON.
- Confirm the expected record array exists and has the expected shape.
- Log the page, cursor, URL, status, item count, and error body.
- Deduplicate with a stable item ID or canonical URL.
- Stop only when the source says there is no continuation: a missing next link or cursor,
has_next: false, or an equivalent documented signal.
Scrapy’s overview and dynamic-content examples show both next-link and Boolean continuation patterns: Scrapy at a glance and its developer-tools tutorial. A fixed page limit is a safety guard, not proof that the crawl is complete.
3. Scrapy implementation for paginated endpoints
Once you know the endpoint, a Scrapy spider can schedule requests and preserve crawl state. This example follows a JSON next_url; adapt the keys to the target.
import scrapy
class ProductsSpider(scrapy.Spider):
name = "products"
start_urls = ["https://example.com/api/products?limit=50"]
def parse(self, response):
if response.status != 200:
self.logger.error("HTTP %s: %s", response.status, response.url)
return
payload = response.json()
for item in payload.get("items", []):
yield {
"id": item.get("id"),
"name": item.get("name"),
"url": item.get("url"),
}
next_url = payload.get("next_url")
if next_url:
yield response.follow(next_url, callback=self.parse)
For POST-based searches, reproduce the JSON body with FormRequest or JsonRequest. Preserve a session only when cookies are genuinely required, and do not copy short-lived browser tokens into a long-running crawler without a plan to refresh them.
4. Browser automation when HTTP replay is not enough
Use Playwright (or another headless browser) when the endpoint depends on client-side state, a complex interaction, a signed request you cannot reproduce, or output that exists only in the rendered DOM. Browser automation costs more setup and is sensitive to UI and timing changes, so keep it as a fallback rather than the default.
Infinite scroll with a page-specific readiness check
import asyncio
from playwright.async_api import async_playwright
async def main():
async with async_playwright() as p:
browser = await p.chromium.launch(headless=True)
page = await browser.new_page()
await page.goto("https://example.com/catalog", wait_until="domcontentloaded")
previous_count = 0
while True:
current_count = await page.locator("article.product").count()
if current_count == previous_count:
# Trigger the site’s own scroll handler.
await page.evaluate("window.scrollTo(0, document.body.scrollHeight)")
try:
await page.wait_for_function(
"old => document.querySelectorAll('article.product').length > old || "
"document.querySelector('[data-end]') !== null",
arg=current_count,
timeout=10000,
)
except Exception:
# A timeout is meaningful only after checking the end marker/count.
pass
new_count = await page.locator("article.product").count()
end = await page.locator("[data-end]").count()
if end or new_count <= current_count:
break
previous_count = new_count
items = await page.locator("article.product").evaluate_all(
"els => els.map(e => ({name: e.querySelector('h2')?.textContent?.trim(), "
"url: e.querySelector('a')?.href}))"
)
for item in items:
print(item)
await browser.close()
asyncio.run(main())
Replace selectors with ones tied to the target’s result item and end marker. Waiting for a changed result count, a newly visible item, or a known “no more results” marker is safer than sleeping for an arbitrary duration. Playwright explains navigation and readiness behavior in Navigations and the Page API reference; its documentation cautions that generic network-idle is not a universal readiness condition.
“Load more” button
Locate the button, click it, and wait for the item count to increase. Disable or break when the button is absent, disabled, or replaced by an end marker. Capture the response event only if the request is stable; otherwise extract the newly rendered items after the page-specific wait.
5. Choosing the right approach
| Consideration | Replay the data request | Browser automation |
|---|---|---|
| Data access | Parse the response that contains records | Read rendered DOM or interact with controls |
| Best fit | Stable, understandable endpoint | Complex browser state or browser-only output |
| Complexity | Investigation up front; fewer rendering steps | Browser installation, timing, and UI maintenance |
| Typical failures | Changed parameters, schema, or token requirements | Selectors, timing, dialogs, and browser state |
| Verification | Status, schema, item count, continuation fields | Observed result changes and end condition |
These are implementation trade-offs, not measured performance claims. A hybrid is often best: discover the request in a browser, then crawl it with Scrapy or an HTTP client.
6. Reliability, politeness, and data quality
- Use bounded retries with exponential backoff for transient 429 and 5xx responses; do not retry validation errors forever.
- Persist the cursor or page state so a process can resume without duplicating all earlier records.
- Keep a crawl manifest containing timestamps, request parameters, response status, and parser version.
- Fail visibly when the schema changes. An empty list is not automatically “no more pages.”
- Throttle requests, honor published limits, and avoid parallelism that overloads the service.
- Review
robots.txt, terms, authentication requirements, and applicable law before crawling. RFC 9309 defines robots.txt as crawler access rules and explicitly states: “These rules are not a form of access authorization.” Read the standard at RFC 9309.
7. Common failures and fixes
The HTML has no records
Cause: records arrive through XHR/fetch after load. Fix: inspect Network while performing one pagination action and replay the response request.
Recommended Free Tools
The crawler always gets the first page
Cause: the cursor, offset, or body value is not being updated, or the cursor is tied to a session. Fix: log the outgoing URL/body and the returned continuation token; preserve the required session and send the token exactly as returned.
It stops early
Cause: code treats an empty batch, timeout, or HTTP error as completion. Fix: distinguish errors from valid end signals, validate the schema, and retry transient failures within a limit.
It loops forever or duplicates records
Cause: a server repeats a cursor or the UI re-renders existing items. Fix: keep a set of seen cursors and stable item keys; abort on a repeated continuation value and investigate.
Playwright sees stale results
Cause: the script waits for generic navigation or a fixed sleep instead of evidence that this batch changed. Fix: wait for an increased item count, a specific new item, or an end marker, and handle consent dialogs only when they block the target interaction.
403, CAPTCHA, or login wall
Cause: access controls or authentication. Fix: stop and use an authorized account or documented API. Do not attempt to bypass a CAPTCHA or access restriction.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Or skip the browser setup
If your goal is a reliable visual capture of each paginated state—not extraction of the underlying records—ScreenshotNeo provides a one-request website screenshot API. It can accept cookie/consent banners before capture and remove more than 60 known consent platforms, newsletter popups, and chat widgets; each cleanup step can be disabled. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers report the page verdict and billing result.
It also offers an MCP server for Claude, Cursor, and other MCP clients, with take_screenshot, get_page_info, and capture_pdf tools. Features include full-page lazy-image capture, CSS-selector element capture, device and viewport controls, custom CSS/JavaScript, click and wait conditions, request blocking, headers/cookies, timezone and geolocation, resizing, chosen-TTL caching, signed links, asynchronous webhooks, bulk capture of up to 100 URLs per call, a usage API, and an OpenAPI specification.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com/catalog -o shot.webp
See the ScreenshotNeo API documentation for parameters and response headers. A free plan includes 1,000 shots per month with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
8. Cost and operating notes
HTTP replay generally uses fewer rendering steps, but investigation and endpoint maintenance take engineering time. Browsers require Chromium or another engine, consume more CPU and memory, and need explicit waits and selector maintenance. In either approach, cache only when the site permits it, avoid unnecessary recrawls, and record enough metadata to explain missing or changed records. No universal speed or success percentage applies because the endpoint, page size, rate limits, and browser behavior differ by site.
FAQ
How do I know whether a site uses infinite scroll?
Scroll once while watching Network. If a new request returns records and the DOM grows without a URL change, it is an infinite-scroll pattern; the same request-inspection method applies to “Load more.”
Should I scrape the API or the rendered HTML?
Use the structured request when it is reproducible and authorized. Use rendered HTML when browser state or interaction is essential, or when reproducing the request would be less reliable than observing the page.
Is a missing next link always the end?
It is an end signal only when that is how the target represents completion. Confirm the response schema and distinguish a missing field caused by an error from a valid terminal response.
Is robots.txt permission to scrape?
No. RFC 9309 describes it as crawler guidance and says its rules are not access authorization. Check terms, permissions, and law separately.
Frequently Asked Questions
Can I scrape a JavaScript site without Selenium?
Often. If Network inspection reveals a reproducible JSON or HTML request, an HTTP client or Scrapy is usually sufficient; use a browser when the required state cannot be reproduced.
What should I store to resume a crawl?
Persist the current page, cursor or next URL, stable item keys, and enough request metadata to verify that a resumed response belongs to the same crawl.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.
The Tool Desk
Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →




