Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

Use a Requests session to fetch each page, Beautiful Soup to extract its records, and a loop that follows the site’s real “Next” link until it disappears or produces no new records. Inspect one permitted page first, identify stable selectors, validate every field, de-duplicate records, save progress incrementally, and stop when the site signals that you should stop. If rows are added only after JavaScript runs, inspect the network requests for an official API or use browser automation rather than expecting Requests to execute the page.

What pagination scraping involves

Pagination is not one universal URL pattern. A site may expose ?page=2, numbered links, a cursor, a “Next” anchor, or a button that triggers an API request. The reliable workflow is:

  1. Inspect one page and locate the record container, fields, and pagination control.
  2. Fetch pages with a persistent HTTP session, timeout, status checking, and an identifiable User-Agent.
  3. Parse the returned HTML with Beautiful Soup.
  4. Extract and normalize fields defensively.
  5. Follow the discovered next URL, while tracking visited URLs and record IDs.
  6. Persist output page by page and stop on a missing next link, no new records, an explicit limit, or an access denial.

Before crawling, read the site’s robots.txt, terms, privacy requirements, and any published API documentation. Treat robots rules as an access and traffic-management signal, not as permission to bypass restrictions. Do not evade a 403, CAPTCHA, login wall, or 429 rate limit.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Inspect the first page before writing selectors

Find the record container

Open the permitted page in a browser, view its source or inspector, and identify the repeated element for one record: perhaps article.item, a table row such as table tbody tr, or a card with a stable class. Prefer semantic attributes, IDs, and stable class names over generated CSS classes or positional selectors.

Identify fields and pagination

For each record, note selectors for required fields and decide how missing values should be represented. Then inspect the pagination markup. An anchor such as <a rel="next" href="...">Next</a> is safer than guessing a page-number formula. Confirm whether the link is absolute or relative and whether it changes a cursor, query parameter, or path.

Check whether HTML contains the rows

Requests receives the server’s response; it does not run the page’s JavaScript. If the browser shows rows but the response source does not contain them, inspect the browser’s Network panel for an official JSON endpoint or embedded data. Use that documented endpoint when permitted. Use Playwright or Selenium only when browser execution is genuinely required.

Install the Python dependencies

python -m pip install requests beautifulsoup4 lxml

Beautiful Soup can use Python’s built-in html.parser, lxml, or html5lib. lxml is generally the faster choice; html5lib provides browser-like error recovery; the built-in parser avoids an extra parser dependency. The examples below use lxml.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A complete scraper that follows “Next”

Replace the example URL and selectors only after inspecting the target. This script writes each new record to a JSON Lines file, so an interruption does not erase earlier pages.

import json
import time
from pathlib import Path
from urllib.parse import urljoin

import requests
from bs4 import BeautifulSoup

START_URL = "https://example.com/items"       # replace after inspection
OUTPUT = Path("items.jsonl")
MAX_PAGES = 100
DELAY_SECONDS = 1.0

session = requests.Session()
session.headers.update({
    "User-Agent": "PaginationResearchBot/1.0 ([email protected])",
    "Accept": "text/html,application/xhtml+xml",
})

seen_urls = set()
seen_ids = set()
page_url = START_URL
pages_read = 0

with OUTPUT.open("a", encoding="utf-8") as out:
    while page_url and page_url not in seen_urls and pages_read < MAX_PAGES:
        seen_urls.add(page_url)
        response = session.get(page_url, timeout=20)

        if response.status_code in (403, 429):
            raise RuntimeError(
                f"Access was denied or rate-limited ({response.status_code}); stop and review the site's rules."
            )
        response.raise_for_status()

        soup = BeautifulSoup(response.text, "lxml")
        page_records = []

        for card in soup.select("article.item"):  # replace with the real record selector
            title_node = card.select_one("h2")
            link_node = card.select_one("a[href]")
            if not title_node or not link_node:
                continue

            title = title_node.get_text(" ", strip=True)
            record_url = urljoin(page_url, link_node["href"])
            record_id = record_url       # use a site's stable ID when available

            if not title or record_id in seen_ids:
                continue

            record = {"id": record_id, "title": title, "url": record_url}
            seen_ids.add(record_id)
            page_records.append(record)
            out.write(json.dumps(record, ensure_ascii=False) + "n")

        out.flush()
        pages_read += 1

        next_node = soup.select_one('a[rel="next"]')
        next_url = None
        if next_node and next_node.get("href"):
            next_url = urljoin(page_url, next_node["href"])

        if not page_records or not next_url or next_url in seen_urls:
            break

        page_url = next_url
        time.sleep(DELAY_SECONDS)

print(f"Saved {len(seen_ids)} unique records from {pages_read} pages to {OUTPUT}")

The selectors in this example are placeholders. A missing title_node is skipped instead of producing a misleading empty record. The URL set prevents a malformed site from cycling forever; the record-ID set prevents duplicates when pages overlap. If the site exposes a stable database ID, use it instead of the URL.

Extracting tables, cards, and optional fields

HTML tables

For a table, iterate over data rows and map cells to headers. Skip header rows explicitly and verify the expected number of cells before indexing them.

for row in soup.select("table.results tbody tr"):
    cells = [cell.get_text(" ", strip=True) for cell in row.select("td")]
    if len(cells) < 3:
        continue
    item = {"name": cells[0], "status": cells[1], "date": cells[2]}

Optional and nested values

Use select_one and test for None. Normalize whitespace with get_text(" ", strip=True), parse dates only after checking their format, and preserve the original text when a conversion fails. Use urljoin for relative links and normalize canonical URLs if query parameters create aliases.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Validation

Before writing a record, require its identifying field, check that URLs have the expected host, and reject impossible values. Log skipped records with their page URL so you can fix a selector without silently losing data.

Pagination strategies and when to use each

Follow discovered links

Following rel="next" or the site’s equivalent is the most resilient approach when URLs are irregular, cursors are opaque, or filters are encoded in a long query string. Resolve relative links against the current page.

Generate numbered URLs only after confirmation

If inspection proves that every page uses a stable pattern such as ?page=N, you can generate URLs. Still stop on an empty page, repeated records, a configured maximum, or a changed response. Do not assume page one plus an arbitrary number reaches the end.

Cursor and “load more” pagination

A cursor may appear in a link, hidden input, or JSON response. Carry it forward exactly as returned; do not increment it yourself. A “Load more” button often calls an endpoint visible in Network tools. Prefer that endpoint when authorized, and retain the request parameters that preserve filters and sort order.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

JavaScript-rendered pages

If the initial HTML has no records, Requests and Beautiful Soup cannot see what a browser creates later. First look for a documented API or embedded JSON. An API is usually lighter, easier to validate, and less fragile than simulating clicks. If no suitable endpoint exists, a browser tool can wait for a selector, click the control, and read the rendered DOM. Keep browser runs bounded, reuse a browser context where appropriate, and continue to obey access controls and rate limits.

Do not use browser automation to defeat bot checks, CAPTCHAs, authentication, or explicit denials. A blank page, challenge page, or redirect to a login form is a signal to stop and reassess authorization.

Saving, resuming, and scaling safely

Persist incrementally

JSON Lines, CSV append mode, or database upserts lets you resume after a network failure. Store the source URL, crawl timestamp, and a stable record ID. For a restart, load IDs already written and skip them.

Rate limits and retries

Use a deliberate delay and a timeout on every request. Retry only transient failures such as connection resets or selected 5xx responses, with exponential backoff and a maximum retry count. Do not blindly retry 403 or 429; wait according to the site’s guidance or stop. Cache responses during development so repeated selector changes do not repeatedly hit the server.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Memory and throughput

Streaming records to disk avoids holding a large crawl in memory. A single session reuses connections. Concurrency is appropriate only when the site permits it and you can enforce an aggregate request rate; parallel requests can overload a small service and trigger protective controls. Measure pages, records, skipped rows, duplicates, status codes, and elapsed time so a “successful” run with zero records is visible.

Common failures and fixes

Symptom Likely cause Fix
Zero records Wrong selector or rows rendered by JavaScript Inspect response HTML; correct the selector or locate an authorized API/browser workflow.
Only the first page is saved Next link selector, missing href, or a button-based paginator Inspect the exact pagination markup and Network requests; use urljoin for relative links.
Duplicate records Overlapping pages, unstable URLs, or a loop Deduplicate by a stable ID, track visited URLs, and stop when no new IDs appear.
403 or CAPTCHA Access policy, authentication, or bot protection Stop; obtain permission, use an official API, or contact the operator. Do not bypass the control.
429 Rate limit exceeded Stop sending requests, follow any retry guidance, reduce frequency, and resume only when allowed.
Timeouts or intermittent 5xx Network or server instability Use a finite timeout, bounded exponential backoff, caching, and incremental output.
Broken relative links Links resolved against the wrong page Pass the current page URL to urljoin; verify the resulting host and path.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup

If your goal is a visual capture of each page rather than extracting structured records, ScreenshotNeo provides a website screenshot API and MCP server. It can capture PNG, JPEG, WebP, or PDF output. Before capture it accepts cookie/consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each cleanup step can be disabled. Bot checks and CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and the response identifies the result with X-Page-Verdict and X-Billed headers.

One GET request is enough:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`HTTP ${res.status}`);
require('fs').writeFileSync('shot.webp', Buffer.from(await res.arrayBuffer()));

See the ScreenshotNeo documentation for options. Features include full-page captures with lazy images loaded, CSS-selector element capture, dark mode, 12 device presets plus custom viewports, retina scale, PDF paper settings and page ranges, custom CSS and JavaScript, pre-capture clicks, hidden selectors, waits for selectors/delays/network idle, request and resource blocking, headers, cookies, user agents, Authorization, timezone and geolocation, transparent backgrounds, resizing, chosen-TTL caching, signed image links, asynchronous jobs with signed webhooks, up to 100 URLs per bulk call, a usage API, an OpenAPI specification, and compatibility with parameter names used by other screenshot APIs. Its MCP server exposes take_screenshot, get_page_info, and capture_pdf to Claude, Cursor, and other MCP clients.

The Free plan includes 1,000 screenshots per month with no card. Starter is $5 for 3,000, Growth $15 for 15,000, Pro $39 for 60,000, Scale $99 for 250,000, and Business $249 for 1,000,000; yearly billing provides two months free, and every feature is on every plan. Create a free ScreenshotNeo account to start.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

FAQ

How do I know when pagination is finished?

Use the site’s next-control semantics first, then stop when that control disappears, no new record IDs are found, a URL repeats, or your explicit page limit is reached.

Should I use Requests or Selenium?

Use Requests and Beautiful Soup for server-rendered HTML. Choose browser automation only when authorized content is produced by JavaScript and no suitable API or embedded JSON is available.

Is scraping a paginated site legal?

It depends on the site’s terms, applicable law, the data involved, and your authorization. Review robots.txt and terms, minimize collection, protect personal data, and honor explicit denials.

Frequently Asked Questions

Can I scrape pages concurrently to finish faster?

Only if the site permits it and you can enforce a conservative aggregate rate. Sequential requests are safer for small sites; concurrency increases load and the chance of rate limiting.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What should I store to make a crawl auditable?

Store the source page URL, record identifier, crawl time, parsed fields, and error or skip logs. These let you explain where each record came from and resume without duplicates.

Why does my browser show more rows than response.text?

The browser has executed JavaScript after the initial response. Inspect network calls for an authorized API or use browser automation; Beautiful Soup cannot execute that JavaScript.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.