Hardware FixRecommendedDevice not working? Your driver may be the problemCheck updates for common hardware issues.Fix DriversOctober DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsPC HealthRecommendedCrashes, freezes, slowdowns? Check your PC nowSpot repairable issues before they interrupt work.Check PC×
Skip to content
E-commerce

How to Build an E-Commerce Scraper That Survives Real Product Pages

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The dependable way to build an e-commerce scraper is to start with a site-specific data contract, reproduce a product page’s underlying HTTP or JSON request whenever it contains the data you need, and use a browser only for genuinely client-rendered interactions. Scrapy supplies crawling, pagination, retries, pipelines and exports; Playwright, connected through scrapy-playwright, handles JavaScript pages that cannot be extracted from stable requests. Add normalization, deduplication, robots.txt and terms checks, validation, monitoring and crawl timestamps before calling the scraper production-ready.

Define exactly what one product record means

Scraping becomes maintainable when every page produces the same contract. Decide which fields are required, which may be null, and how a variant differs from its parent product. Keep the original URL and retrieval time with every record so a price or availability change can be audited later.

Field Purpose Typical normalization
canonical_url Stable product identity and audit link Use the canonical link when supplied; remove fragments and normalize tracking parameters only under a documented rule
sku or product_id Deduplication and variant tracking Preserve the retailer’s identifier as a string, including leading zeroes
title, brand, category Catalog and search attributes Trim whitespace; retain Unicode; represent missing values as null
variant Size, color, storage or other purchasable option Store a structured object or a deterministic string, not display-only prose
price, currency Comparable monetary value Parse decimal separators and symbols separately; never silently convert currencies
availability Stock state at retrieval time Map site labels to a controlled vocabulary such as in_stock, out_of_stock or unknown
image_url Primary product image Resolve relative URLs and keep the source URL
rating, review_count Optional merchandising signals Collect only when permitted and distinguish absent from zero
retrieved_at Freshness and audit trail Store UTC in an unambiguous format

There is no universal selector set for online stores. Treat selectors, pagination rules and variant behavior as configuration for each site, not as a reusable global scraper.

Choose the least complicated architecture that works

Approach Use it when Trade-offs
Direct HTTP plus parser HTML or a JSON response already contains title, price and stock Lowest latency and simplest operations; fails when data exists only after JavaScript runs
Scrapy crawler You need pagination, link traversal, retries, item pipelines or feed exports Excellent crawl control, but selectors and site behavior require maintenance
Scrapy plus Playwright Prices, variants or stock appear only after client-side rendering or interaction Browser CPU, memory and operational complexity are higher
Hosted scraper API You need scheduling, browsers, proxies or dataset delivery without operating that infrastructure Introduces vendor cost, dependency and program-term considerations

Compare the choices against rendering requirements, crawl volume, freshness, selector stability, compliance constraints, infrastructure budget and tolerance for vendor dependency. Inspect a page and its network calls first. Reproducing an underlying request transfers less data and avoids browser overhead when that request contains the product data.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Build a small Scrapy spider first

Install and create the project

Use an isolated environment, then install Scrapy. Add scrapy-playwright only when a target actually needs a browser.

python -m venv .venv
. .venv/bin/activate
pip install scrapy
scrapy startproject shopcrawler
cd shopcrawler

Set a feed output while developing so you can inspect records without committing to a database.

scrapy crawl products -O products.json

A complete site-specific spider

The selectors below are deliberately examples. Replace them after inspecting the target store’s HTML or JSON response. The spider follows product links and pagination, emits a retrieval timestamp, and handles common price formats without assuming that every store uses the same currency.

import json
import re
from datetime import datetime, timezone
from urllib.parse import urljoin, urldefrag

import scrapy


def clean_text(value):
    return re.sub(r"\s+", " ", value or "").strip() or None


def parse_price(value):
    text = clean_text(value)
    if not text:
        return None
    number = re.sub(r"[^0-9,.-]", "", text)
    if "," in number and "." in number:
        # Treat the last separator as the decimal separator.
        if number.rfind(",") > number.rfind("."):
            number = number.replace(".", "").replace(",", ".")
        else:
            number = number.replace(",", "")
    elif "," in number:
        number = number.replace(",", ".")
    try:
        return float(number)
    except ValueError:
        return None


class ProductsSpider(scrapy.Spider):
    name = "products"
    allowed_domains = ["shop.example"]
    start_urls = ["https://shop.example/category/widgets"]

    custom_settings = {
        "ROBOTSTXT_OBEY": True,
        "DOWNLOAD_DELAY": 0.5,
        "CONCURRENT_REQUESTS": 8,
        "DOWNLOAD_TIMEOUT": 30,
        "RETRY_ENABLED": True,
        "FEEDS": {"products.json": {"format": "json", "overwrite": True}},
    }

    def parse(self, response):
        for href in response.css("a.product-card::attr(href)").getall():
            yield response.follow(href, callback=self.parse_product)

        next_href = response.css("a[rel='next']::attr(href)").get()
        if next_href:
            yield response.follow(next_href, callback=self.parse)

    def parse_product(self, response):
        canonical = response.css("link[rel='canonical']::attr(href)").get()
        canonical = urljoin(response.url, canonical or response.url)
        canonical = urldefrag(canonical)[0]

        currency = response.css("[itemprop='priceCurrency']::attr(content)").get()
        price_value = response.css("[itemprop='price']::attr(content)").get()
        if not price_value:
            price_value = response.css("[itemprop='price']::text").get()

        yield {
            "canonical_url": canonical,
            "sku": clean_text(response.css("[itemprop='sku']::attr(content), [itemprop='sku']::text").get()),
            "title": clean_text(response.css("h1[itemprop='name']::text, h1::text").get()),
            "brand": clean_text(response.css("[itemprop='brand']::text, [itemprop='brand']::attr(content)").get()),
            "category": clean_text(response.css("nav.breadcrumb a::text").getall()[-1] if response.css("nav.breadcrumb a::text").getall() else None),
            "variant": clean_text(response.css("[data-selected-variant]::attr(data-selected-variant)").get()),
            "price": parse_price(price_value),
            "currency": clean_text(currency),
            "availability": clean_text(response.css("[itemprop='availability']::attr href, .availability::text").get()),
            "image_url": urljoin(response.url, response.css("meta[property='og:image']::attr(content), img.product-image::attr(src)").get() or ""),
            "retrieved_at": datetime.now(timezone.utc).isoformat(),
            "source_url": response.url,
        }

In real work, prefer structured JSON or JSON-LD when it is complete, but validate it against the visible page. A retailer can publish stale structured data, omit a selected variant, or expose a promotional price that is not the checkout price. Keep missing fields as null and log the page rather than inventing values.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Normalize and deduplicate in a pipeline

Use the canonical URL as a fallback key and the SKU when it is stable. A pipeline should reject records with no identity, coerce controlled availability values, and record validation failures for review. Do not merge products solely because their titles look similar. Variant identifiers, currency and seller can make otherwise identical titles distinct.

Use Playwright only for genuinely dynamic pages

First inspect the browser’s network panel. If an XHR or fetch response contains the price, availability and variant data, reproduce that request with Scrapy instead of rendering the whole page. If the data appears only after JavaScript, a click, a selector wait or a client-side route change, integrate Playwright.

Minimal scrapy-playwright integration

Install the integration and its browser once:

pip install scrapy-playwright
playwright install chromium

Enable the download handler in Scrapy settings:

DOWNLOAD_HANDLERS = {
    "http": "scrapy_playwright.handler.ScrapyPlaywrightDownloadHandler",
    "https": "scrapy_playwright.handler.ScrapyPlaywrightDownloadHandler",
}
TWISTED_REACTOR = "twisted.internet.asyncioreactor.AsyncioSelectorReactor"
PLAYWRIGHT_BROWSER_TYPE = "chromium"

Request a rendered product page only where needed. Wait for a product selector, perform the minimum interaction, then let the normal callback parse the resulting HTML.

import scrapy
from scrapy_playwright.page import PageMethod

class RenderedProductsSpider(scrapy.Spider):
    name = "rendered_products"

    def start_requests(self):
        yield scrapy.Request(
            "https://shop.example/product/widget",
            meta={
                "playwright": True,
                "playwright_page_methods": [
                    PageMethod("wait_for_selector", "[data-product-price]"),
                    PageMethod("click", "button#accept-necessary", timeout=5000),
                ],
            },
            callback=self.parse_product,
        )

    def parse_product(self, response):
        yield {
            "title": response.css("h1::text").get(),
            "price": response.css("[data-product-price]::text").get(),
            "availability": response.css("[data-availability]::text").get(),
            "source_url": response.url,
        }

Keep browser concurrency low, close pages promptly, and avoid loading images or third-party resources unless they are part of the data contract. A browser is not a workaround for access controls: do not bypass CAPTCHAs, bot checks, authentication boundaries or technical restrictions.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Make crawling polite and compliant

Enable Scrapy’s ROBOTSTXT_OBEY setting so the crawler respects robots.txt. Also review the site’s terms, authentication boundaries, privacy obligations and applicable law before collecting or redistributing data. Robots.txt is an access preference, not a complete legal analysis.

  • Use conservative concurrency and download delays, especially on smaller stores.
  • Set request timeouts and retries with backoff; do not retry every client error indefinitely.
  • Cache responses during development and selector work to avoid repeated traffic.
  • Identify your crawler where the site’s policy permits and provide a contact route when appropriate.
  • Collect only the fields you need, and avoid personal data unless you have a clear lawful basis.

Validate, persist and monitor every run

Write validated items to a database or feed with the source URL and crawl timestamp. Before accepting a record, check that a price is non-negative, a currency is present when price is present, an image URL is absolute, and availability belongs to your controlled vocabulary. Keep the raw response or a content hash when audit requirements justify it.

Alert on selector drift and operational anomalies, not just exceptions:

  • zero products from a category that normally returns data;
  • an abrupt rise in null prices or unknown availability;
  • HTTP error spikes, timeouts or repeated redirects;
  • duplicate keys that exceed an expected threshold;
  • price changes outside a business-defined range.

Scrapy’s ecosystem includes Spidermon for monitoring, Scrapy Cloud for deployment and Zyte API for proxy or browser infrastructure. Commercial terms and availability can change, so verify current program details before committing. For recurring or multi-site jobs, schedule runs, partition work by store or category, and retain crawl provenance. A hosted scraper API can be sensible when operating browsers, proxies, scheduling and dataset delivery costs more engineering time than it saves.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Control freshness, performance and cost

Freshness is a product decision. Inventory monitoring may need frequent incremental requests, while a catalog export can run less often. Partition by category or changed URLs, use conditional requests where the retailer supports them, and avoid recrawling pages whose data cannot have changed.

Direct HTTP requests normally consume fewer CPU and memory resources than browser pages. Browser rendering adds startup, JavaScript execution and network cost, so reserve it for the subset of URLs that need it. Measure your own request rate, response sizes, browser minutes and failed pages; no universal performance number applies. Caching reduces load during development, but do not serve stale prices as current data without labeling their retrieval time.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Troubleshooting common failures

Symptom Likely cause Fix
HTML has no price Price is inserted by JavaScript or a different API call Inspect network requests; reproduce the data request, or use scrapy-playwright with a selector wait
Price is null or wildly wrong Locale separators, currency symbols or sale-price markup differ Parse numeric and currency fields separately; add a site-specific locale rule and validation bounds
Only the first page is collected Pagination is a button, cursor or JavaScript route Follow a stable next link, reproduce the cursor request, or render the required interaction
Many duplicate products Tracking parameters, fragments or variant URLs create multiple addresses Canonicalize URLs and deduplicate by SKU plus canonical URL; keep variant IDs distinct
Empty result set after a redesign Selectors drifted or the site returned a consent/interstitial page Save a sample response, alert on zero items, update selectors and handle the interstitial explicitly
Timeouts and 5xx responses Concurrency is too high or the store is temporarily unhealthy Lower concurrency, add bounded retries with backoff, increase timeout only when evidence supports it
403, CAPTCHA or bot check The site is enforcing an access control Stop and review permission and terms; do not attempt to defeat the control
Browser memory keeps growing Too many concurrent pages or unclosed contexts Reduce browser concurrency, block unnecessary resources and ensure pages are closed

Or skip the browser setup

If your job is to obtain a clean visual of a product page rather than parse its fields yourself, ScreenshotNeo provides a single website-screenshot API call. It accepts cookie or consent banners as a visitor and removes more than 60 known consent platforms, newsletter popups and chat widgets before capture; each step can be turned off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads and cache hits are not billed, and the response identifies the page verdict and billing status in headers. Its MCP server provides take_screenshot, get_page_info and capture_pdf tools for Claude, Cursor and other MCP clients.

See the ScreenshotNeo API documentation for all options, including full-page lazy-image loading, CSS-selector element capture, device presets, custom CSS and JavaScript, waits, request blocking, cookies, headers, geolocation, PDF output, signed links, asynchronous jobs, bulk capture and caching.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

cURL

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)

Node.js

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`HTTP ${res.status}`);
const fs = await import('node:fs/promises');
await fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer()));

The Free plan includes 1,000 shots per month with no card; paid plans start at $5 for 3,000 shots, and every feature is on every plan. Create a free ScreenshotNeo account to try it without a card.

Further reading

For a book-length introduction, Ryan Mitchell’s Web Scraping with Python, 2nd Edition (O’Reilly, ISBN 9781491985564) is a natural companion to this project, especially if you are new to request inspection, parsing and crawler design.

Frequently Asked Questions

Is an e-commerce scraper the same as a price-monitoring system?

No. A scraper extracts observations; a monitoring system adds scheduling, historical storage, change detection and alerts. Keep those responsibilities separate so extraction failures are not mistaken for price changes.

How should variants be represented?

Represent each purchasable variant with its own identifier, price and availability when the store supplies one. Keep a parent product reference if you also need product-level grouping.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

When should I stop scraping a page?

Stop when the data contract is complete, the page has no additional product links or pagination, or the site’s policy disallows further access. A bounded crawl is easier to audit than an open-ended link follower.

What should I retain for an audit?

At minimum retain the canonical source URL, retrieval timestamp, normalized record and validation status. For higher-assurance workflows, retain a raw response or content hash under an appropriate retention policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Read next

Recommended PC Tool
Recommended PC Tool
Crashes, No Sound, or Screen Glitches?Free driver scan
Windows Errors? Fix Them Before They SpreadFree repair scan

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.