October DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsWindows FixRecommendedWindows errors stealing your time? Find the fix fastScan stability, cleanup and performance issues.Fix NowOctober DealsAmazon USDeal season is back - check today's better picksAmazon US: current deals, useful picks and tech finds.See Picks×
Skip to content
BeautifulSoup

How to Scrape E-Commerce Category Pages: Pagination, JavaScript, and Data Quality

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

To scrape an e-commerce category page reliably, first check the site’s access rules, then fetch its ordinary HTML and extract product cards with CSS selectors. Follow real pagination links or a permitted data endpoint to reach additional products; use a browser such as Playwright only when the data genuinely requires JavaScript or user interaction. Normalize and deduplicate the results, and validate them before treating the crawl as a complete catalog.

Plan the crawl before writing a scraper

A category page is usually a view of part of a catalog, not a complete product feed. Decide what you need to collect, how far the crawl may go, and how often it should run before sending requests.

Define records and boundaries

For each product, useful fields may include its URL, title, SKU or other exposed identifier, price, currency, availability, image URL, category path, and crawl timestamp. Some pages will not expose every field. Record missing values rather than silently substituting guesses. Set a category list, maximum page count, refresh cadence, and request concurrency that fit your need.

Decide how product variants should be represented. If a page exposes distinct sizes or colors with separate identifiers or prices, preserve those identifiers rather than collapsing them into a single product row. Keep the original text or response metadata alongside normalized values so a later parser change can be diagnosed.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Check permission and publisher controls

Before collecting data, read the site’s robots.txt, terms of service, and any applicable restrictions. Treat these as separate checks: robots rules are not a permission grant, and they do not settle questions about privacy, copyright, database rights, authentication, or contractual terms. Do not attempt to defeat login barriers, access controls, or anti-bot measures. Use a descriptive user agent, conservative request rates, timeouts, retries with backoff, caching, and a hard page cap.

Google explains that robots.txt manages crawler traffic and is not a way to hide URLs from search results. See its Robots.txt Introduction and Guide. A disallow rule should not be treated as evidence that a URL is secret or as a substitute for checking the site’s terms and applicable rules.

Choose an approach that matches the page

What the site exposes Approach Trade-off
Product cards and next-page links are present in the initial HTML HTTP client plus Scrapy selectors, lxml, or BeautifulSoup Fast and inexpensive, but will not see important content rendered only in the browser
Many categories need scheduled refreshes, retries, and job tracking Scrapy spider with item pipelines and persistent job state Offers crawl controls for larger jobs but requires framework setup
Products or prices appear only after JavaScript or an interaction First investigate whether a permitted JSON endpoint provides the data; otherwise use Playwright or another browser renderer More faithful to browser-visible content, but slower and more resource-intensive
A sitemap or merchant feed lists the catalog Discover product URLs there, then make targeted product requests Can make discovery more efficient, though feed fields may differ from page fields

Start with ordinary HTML rather than assuming a browser is necessary. Scrapy describes spiders as components that generate requests, parse responses, and return structured items; its documentation covers both spiders and selectors. For discovery, inspect navigation links, sitemaps, and feeds when the visible category links do not expose all relevant products. Google’s e-commerce structure guidance recommends direct links among menus, categories, subcategories, and products, and discusses sitemaps or feeds where links are incomplete.

Build a small HTML scraper with Python

The example below uses Requests and BeautifulSoup to fetch a category page, parse product cards, and follow a conventional next-page link. Install the dependencies with python -m pip install requests beautifulsoup4. Save the script as scrape_category.py, replace the starting URL and selectors with values observed on a site you are permitted to crawl, then run python scrape_category.py.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The CSS selectors in this example are illustrative: stores use different markup. Inspect the page source or browser developer tools to identify the repeated card container, product link, title, price, and next link. If a field is absent, the script leaves it empty rather than fabricating it.

import csv
import time
from urllib.parse import urljoin, urldefrag, urlparse

import requests
from bs4 import BeautifulSoup

START_URL = "https://example.com/category/shoes"
MAX_PAGES = 20
DELAY_SECONDS = 1.0

# Adapt these selectors to the store's ordinary HTML.
CARD_SELECTOR = ".product-card"
TITLE_SELECTOR = ".product-title"
PRICE_SELECTOR = ".price"
PRODUCT_LINK_SELECTOR = "a.product-link"
NEXT_SELECTOR = "a.next"

session = requests.Session()
session.headers.update({"User-Agent": "CategoryResearchBot/1.0 (contact: [email protected])"})


def clean_url(base, href):
    if not href:
        return ""
    absolute = urljoin(base, href)
    absolute, _fragment = urldefrag(absolute)
    return absolute


def text_or_empty(parent, selector):
    node = parent.select_one(selector)
    return node.get_text(" ", strip=True) if node else ""


def scrape():
    url = START_URL
    start_host = urlparse(START_URL).netloc
    seen_pages = set()
    seen_products = set()
    rows = []

    for _page_number in range(MAX_PAGES):
        url = clean_url(url, url)
        if not url or url in seen_pages:
            break
        if urlparse(url).netloc != start_host:
            raise ValueError("Next-page link left the configured site boundary")
        seen_pages.add(url)

        response = session.get(url, timeout=(5, 30))
        response.raise_for_status()
        soup = BeautifulSoup(response.text, "html.parser")

        for card in soup.select(CARD_SELECTOR):
            link = card.select_one(PRODUCT_LINK_SELECTOR)
            product_url = clean_url(url, link.get("href")) if link else ""
            if product_url and product_url in seen_products:
                continue
            if product_url:
                seen_products.add(product_url)
            rows.append({
                "product_url": product_url,
                "title": text_or_empty(card, TITLE_SELECTOR),
                "price_text": text_or_empty(card, PRICE_SELECTOR),
                "category_page": url,
                "crawled_at_utc": time.strftime("%Y-%m-%dT%H:%M:%SZ", time.gmtime()),
            })

        next_link = soup.select_one(NEXT_SELECTOR)
        next_url = clean_url(url, next_link.get("href")) if next_link else ""
        if not next_url:
            break
        time.sleep(DELAY_SECONDS)
        url = next_url

    with open("products.csv", "w", newline="", encoding="utf-8") as handle:
        fields = ["product_url", "title", "price_text", "category_page", "crawled_at_utc"]
        writer = csv.DictWriter(handle, fieldnames=fields)
        writer.writeheader()
        writer.writerows(rows)

    print(f"Wrote {len(rows)} product rows from {len(seen_pages)} pages to products.csv")


if __name__ == "__main__":
    scrape()

What to change before using it

  • Set START_URL to the category URL and replace the example user agent with one that identifies your crawler appropriately.
  • Inspect a representative page and adapt the card, title, price, product-link, and next-link selectors. A selector that matches no cards can produce a valid-looking but empty CSV.
  • Choose MAX_PAGES as a safety limit for your task. The script also stops when there is no next link or the same page URL recurs.
  • Increase the delay or reduce crawl scope if the publisher’s rate limits or terms require it. This simple example does not implement automatic retries or persistent job state; for a larger scheduled crawl, use a framework such as Scrapy and configure its crawl controls appropriately.

The example stores price as displayed text. For useful comparisons, parse the amount and currency separately with rules suited to the site’s locale. Do not assume a comma or period always has the same decimal meaning across markets.

Reach every product without clicking blindly

After parsing the first response, inspect how the category exposes further products. The simplest case is a real next-page link with a distinct URL. Follow it until the link disappears, product identifiers stop changing, or your configured page limit is reached. Keep a set of visited page URLs and product identifiers so duplicate links or repeated listings do not inflate the output.

Google recommends unique URLs for paginated sequences and warns that URL fragments are not reliable page numbers. See its pagination and incremental page loading guidance. Do not assume changing a fragment such as #page=2 fetches a distinct page; verify that the server actually returns different content.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Load-more buttons and infinite scroll

If a button or scroll event reveals more products, use browser developer tools to inspect the network requests made as the page loads more items. A stable JSON request may be a simpler and lighter source than rendering every browser interaction, but use it only if access is permitted and the endpoint is not an access-control bypass. Confirm that the response corresponds to the same category and preserves pagination state.

If there is no stable permitted endpoint and essential content appears only after JavaScript or a user action, a browser renderer such as Playwright is a reasonable fallback. This costs more time and resources than parsing initial HTML. Google notes that its crawlers generally do not click buttons or trigger JavaScript functions that require user actions to update a page; that is a search-crawling observation, not a guarantee about how a particular store’s own browser interface works.

Normalize, deduplicate, and validate the data

Scraping is not complete when a CSV has been written. Check whether the output represents distinct products and whether the values have a consistent meaning.

  • Canonicalize product URLs consistently, while preserving query parameters if they identify a real variant.
  • Prefer a stable exposed SKU or product ID for deduplication; otherwise use a stable product URL. Keep variant identifiers where separate variants matter.
  • Normalize availability labels into a documented set of values, but retain the original label for auditability.
  • Store parsed numeric prices with an explicit currency field, and retain raw displayed price text for locale or markup changes.
  • Track missing-field rates, duplicate rates, page counts, HTTP status distributions, and unexpected template changes.
  • Keep representative page fixtures and run parser regression checks when selectors or site markup change.

A category can omit products that exist elsewhere in the catalog, and a crawl can miss items because of pagination, merchandising changes, or a parser mismatch. If coverage matters, compare category results with permitted sitemap or feed discovery, and document the crawl boundary and timestamp with the dataset.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server made by Yorker Media. A screenshot is visual evidence, not a substitute for structured product extraction: use it to inspect how a category rendered or to keep a visual check alongside your scraper. One GET request can capture a page as PNG, JPEG, WebP, or PDF. Before capture, it can accept cookie or consent banners as a visitor and remove more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be turned off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and responses include X-Page-Verdict and X-Billed headers.

For example, capture a category page as WebP with cURL. Create an API key first, then replace the target URL with the category you are allowed to inspect. See the ScreenshotNeo API documentation for request options and response details.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com/category/shoes -o shot.webp

ScreenshotNeo also provides an MCP server for AI agents, including Claude, Cursor, and other MCP clients, with take_screenshot, get_page_info, and capture_pdf tools. Plans include 1,000 shots per month free with no card, Starter at $5 for 3,000, Growth at $15 for 15,000, Pro at $39 for 60,000, Scale at $99 for 250,000, and Business at $249 for 1,000,000; yearly billing gives two months free. Every feature is on every plan. Learn more at ScreenshotNeo, or sign up for 1,000 free screenshots a month with no card.

Troubleshoot common failures

The script returns zero products

Most often the card selector does not match the live markup, the page returned a consent or error screen, or products are inserted only after JavaScript runs. Inspect the response HTML and confirm the page status before changing selectors. If the initial response truly lacks the product data, inspect permitted network responses or use a browser renderer rather than repeatedly tuning selectors against an empty page.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Only the first batch appears

Check whether the category has a real next-page anchor, a load-more request, or an infinite-scroll request. Confirm that each next URL returns new products and is not merely a tracking or fragment variation. Set a page cap and stop on repeated page URLs or product IDs.

Prices or availability look wrong

Keep raw strings and inspect locale, variant selection, and page state. A displayed promotional price may coexist with a crossed-out list price, and stock wording may vary. Define which price and availability field your use case requires; do not treat one selector as semantically correct across every store.

Requests time out or return errors

Use bounded timeouts, lower concurrency, cache responses where appropriate, and retry transient failures with backoff rather than in a tight loop. Check the site’s rate limits and whether the crawl should pause. Do not respond to bot checks or authentication barriers by trying to circumvent them.

FAQs

Does robots.txt give permission to republish product data?

No. Robots.txt communicates crawler preferences; it does not by itself grant permission to collect, store, or republish data. Check the site’s terms and the rules applicable to your intended use.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Can I use a sitemap instead of crawling category pagination?

A sitemap or merchant feed can help discover product URLs when category browsing is incomplete. It may not contain the same fields as product pages, so verify whether it supplies the attributes your dataset needs.

Frequently Asked Questions

Does robots.txt give permission to republish product data?

No. Robots.txt communicates crawler preferences; it does not by itself grant permission to collect, store, or republish data. Check the site’s terms and the rules applicable to your intended use.

Can I use a sitemap instead of crawling category pagination?

A sitemap or merchant feed can help discover product URLs when category browsing is incomplete. It may not contain the same fields as product pages, so verify whether it supplies the attributes your dataset needs.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Read next

Recommended PC Tool
Recommended PC Tool
PC Slower Than It Used to Be?Free scan - under a minute
Crashes, No Sound, or Screen Glitches?Free driver scan

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.