Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Build a price scraper as a small, auditable pipeline: fetch a product page, identify the product and seller, parse price and currency, record availability and discount state, validate the result, then store a timestamped observation. Start with Python requests and Beautiful Soup for server-rendered HTML. Use Playwright when JavaScript or an AJAX request inserts the price after the initial response.

The examples below show both paths, including normalization, retries, scheduling, history, alerting and failure handling. They are designed for legitimate monitoring of pages you are allowed to access.

What a price scraper should capture

Define the output before writing selectors. A useful record has stable fields that can be compared over time and audited when a price alert looks wrong.

Field Purpose
product_url Canonical page URL that produced the observation.
sku SKU, product ID or another seller-provided identifier.
name Product name as displayed or supplied in structured data.
price_amount Normalized numeric value, never a formatted display string.
currency ISO-style currency code when the page provides one.
availability In stock, out of stock, preorder or an explicitly unknown state.
discount Sale state and, when available, the previous price.
retrieved_at UTC timestamp for the observation.
http_status, parser_version, error_state Operational context needed to distinguish a real change from a broken scrape.

Store the source URL and retrieval time with every row. If the target’s terms permit it, retain raw HTML or a content hash so you can inspect a selector failure without fetching the page again.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Check permission and crawl controls first

Fetch https://host/robots.txt for the applicable host before crawling. Google’s Crawling Infrastructure documentation says, “A robots.txt file lives at the root of your site.” The file can contain user-agent groups, allow, disallow and optional sitemap directives.

A robots file is a crawl instruction, not a complete legal permission. Read the site’s terms, authentication requirements and published rate limits, and consider the law that applies to you and the target. Legality varies by jurisdiction and by the target’s terms. Do not bypass login controls, CAPTCHAs or other access restrictions.

Choose the fetch method

Situation Recommended method Trade-off
The price is present in the initial HTML requests plus Beautiful Soup Fast, inexpensive and easy to operate.
The initial HTML has a shell, while JavaScript inserts the price Playwright with Chromium Higher CPU and memory use, but it observes the rendered DOM.
Many domains, browsers, proxies or recurring jobs are the bottleneck A managed scraping API Less infrastructure to run; verify current pricing, geography, data rights and partner terms before choosing one.

Inspect “view source” or the raw response first. If the price appears there, do not pay the cost of a browser. If it appears only after a script runs, use a browser and wait for the specific price element rather than sleeping for an arbitrary number of seconds.

Build a static HTML scraper in Python

Install dependencies

python -m venv .venv
. .venv/bin/activate
pip install requests beautifulsoup4

Use semantic data before CSS classes

Prefer JSON-LD, itemprop, aria-label and stable data-testid attributes. Generated class names often change during a deployment. The following script first looks for Product JSON-LD, then falls back to common semantic attributes. Replace the URL and selectors with values from the target page.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
from datetime import datetime, timezone
from decimal import Decimal, InvalidOperation
import json
import re
import requests
from bs4 import BeautifulSoup

URL = "https://example.com/product/widget"
HEADERS = {"User-Agent": "PriceMonitor/1.0 (contact: [email protected])"}


def parse_decimal(text):
    if not text:
        return None
    value = text.replace("\u00a0", " ").strip()
    # Keep digits, separators and a leading minus; currency is recorded separately.
    value = re.sub(r"[^0-9,.-]", "", value)
    if not value:
        return None
    if "," in value and "." in value:
        # Treat the last separator as the decimal separator.
        if value.rfind(",") > value.rfind("."):
            value = value.replace(".", "").replace(",", ".")
        else:
            value = value.replace(",", "")
    elif "," in value:
        tail = value.rsplit(",", 1)[-1]
        value = value.replace(",", ".") if len(tail) in (1, 2) else value.replace(",", "")
    try:
        return Decimal(value)
    except InvalidOperation:
        return None


def first_value(product, *keys):
    for key in keys:
        value = product.get(key)
        if value not in (None, ""):
            return value
    return None

response = requests.get(URL, headers=HEADERS, timeout=30)
response.raise_for_status()
soup = BeautifulSoup(response.text, "html.parser")
products = []
for node in soup.select('script[type="application/ld+json"]'):
    try:
        data = json.loads(node.string or node.get_text())
    except json.JSONDecodeError:
        continue
    candidates = data if isinstance(data, list) else data.get("@graph", [data]) if isinstance(data, dict) else []
    products.extend(x for x in candidates if isinstance(x, dict) and x.get("@type") in ("Product", ["Product"]))
product = products[0] if products else {}
offer = product.get("offers", {})
if isinstance(offer, list):
    offer = offer[0] if offer else {}

raw_price = first_value(offer, "price") or soup.select_one('[itemprop="price"]')
raw_currency = first_value(offer, "priceCurrency")
if hasattr(raw_price, "get_text"):
    raw_price = raw_price.get("content") or raw_price.get_text(" ", strip=True)
if not raw_currency:
    node = soup.select_one('[itemprop="priceCurrency"]')
    raw_currency = (node.get("content") or node.get_text(strip=True)) if node else None
if not raw_price:
    node = soup.select_one('[data-testid="price"], [aria-label*="price" i]')
    raw_price = node.get_text(" ", strip=True) if node else None

availability = first_value(offer, "availability")
record = {
    "product_url": response.url,
    "sku": product.get("sku") or product.get("mpn"),
    "name": product.get("name") or (soup.select_one("h1").get_text(" ", strip=True) if soup.select_one("h1") else None),
    "price_amount": str(parse_decimal(str(raw_price))) if parse_decimal(str(raw_price)) is not None else None,
    "currency": raw_currency,
    "availability": availability,
    "discount": {"sale": bool(offer.get("priceValidUntil") or product.get("offers"))},
    "retrieved_at": datetime.now(timezone.utc).isoformat(),
    "http_status": response.status_code,
    "parser_version": "2026-09-01",
    "error_state": None,
}
if record["price_amount"] is None or not record["currency"]:
    record["error_state"] = "missing_price_or_currency"
print(json.dumps(record, ensure_ascii=False, indent=2))

Do not turn an unavailable price into zero. Keep it as null (or your database’s equivalent) and mark the record as an error or unavailable state. Record the currency before stripping symbols, and test both decimal and thousands separators used by the target’s locale.

Handle JavaScript-rendered prices with Playwright

Install and launch Chromium

pip install playwright beautifulsoup4
playwright install chromium

Wait for the price selector that actually changes the DOM. A targeted wait is more reliable than a fixed delay, although a short delay can be useful for animations after the element exists.

from datetime import datetime, timezone
from decimal import Decimal
import re
from playwright.sync_api import sync_playwright, TimeoutError as PlaywrightTimeoutError

URL = "https://example.com/product/widget"
PRICE_SELECTOR = '[data-testid="price"]'

with sync_playwright() as p:
    browser = p.chromium.launch(headless=True)
    page = browser.new_page(locale="en-US")
    try:
        response = page.goto(URL, wait_until="domcontentloaded", timeout=60000)
        page.wait_for_selector(PRICE_SELECTOR, state="visible", timeout=30000)
        raw = page.locator(PRICE_SELECTOR).inner_text()
        value = re.sub(r"[^0-9.,-]", "", raw)
        if "," in value and "." in value:
            value = value.replace(",", "") if value.rfind(".") > value.rfind(",") else value.replace(".", "").replace(",", ".")
        price = Decimal(value)
        result = {
            "product_url": page.url,
            "name": page.locator("h1").first.inner_text() if page.locator("h1").count() else None,
            "price_amount": str(price),
            "currency": "USD",  # derive this from the page; do not assume it in production
            "retrieved_at": datetime.now(timezone.utc).isoformat(),
            "http_status": response.status if response else None,
            "error_state": None,
        }
        print(result)
    except PlaywrightTimeoutError:
        print({"product_url": URL, "error_state": "price_selector_timeout"})
    finally:
        browser.close()

Derive currency from JSON-LD, a currency meta tag or the visible locale; the example’s USD is intentionally marked as a value to replace. If a consent dialog blocks the page, handle it only when the target permits automated interaction. Treat a CAPTCHA, login page or empty product shell as a failed observation.

Normalize and validate every observation

  • Require a product identity, nonnegative price, recognized currency and a retrieval timestamp.
  • Accept an explicit unavailable or out-of-stock state; never coerce it to a numeric price.
  • Check bounds such as an expected maximum price and flag sudden distribution changes for review.
  • Compare the current value with the previous observation for the same product and seller, not merely with the last page fetched.
  • Store parser version, HTTP status and an error state so a selector miss cannot trigger a false discount alert.
  • Keep localized parsing rules per locale when a site serves different decimal conventions.

Schedule requests without overloading the target

A daily run may suit stable catalog prices. Faster-changing products need shorter intervals, but only within the target’s published limits. Use conservative pacing and bounded concurrency.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Retries and backoff

Retry transient network failures and selected 5xx responses with exponential backoff and jitter. Do not endlessly retry 401, 403, CAPTCHA, robots denials or a persistent selector miss. Cap attempts, record the final error and alert when the success rate falls.

Cron example

# Run at 06:15 UTC every day
15 6 * * * /srv/price-monitor/.venv/bin/python /srv/price-monitor/scrape.py >> /var/log/price-monitor.log 2>&1

For browser jobs, limit the number of concurrent Chromium contexts and close every browser in a finally block. Emit metrics for request count, latency, successful records, missing selectors and blocked pages.

Persist history and generate alerts

Use an append-only table keyed by product and seller rather than overwriting the current value. A relational schema can include product_id, seller, observed_at, amount, currency, availability, source_url, parser_version and error_state.

On each valid observation, compare the normalized amount with the immediately previous valid observation. Alert only after validation succeeds, and include the old value, new value, currency, URL and timestamps. Retain enough history to explain a notification and to distinguish a one-page glitch from a sustained change.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Scale deliberately

Self-hosted Requests and Beautiful Soup maximize control and have low operating cost for simple pages. Playwright adds rendering fidelity but requires browser CPU, memory and lifecycle management. At larger volume, browser hosting, proxy management and job orchestration may dominate engineering time.

Managed services can provide rendered-page fetching, proxy and browser operations, synchronous or asynchronous jobs, run polling, dataset export and recurring schedules. Scrapy.io documents those API capabilities; Decodo documents a managed eCommerce price-scraping API for rendered pages and protected targets. Verify current pricing, geographic coverage, data rights and partner terms before committing to either service.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Troubleshooting common failures

Symptom Likely cause Fix
Price is null in Requests JavaScript inserts it after load, or the selector changed. Inspect raw HTML; move to Playwright if absent, otherwise update a semantic selector and parser version.
Price parses as 12345 instead of 123.45 Locale-specific separators were handled incorrectly. Capture currency and locale, then apply a tested separator rule before conversion.
Every run reports a discount Sale detection is based on the presence of an offers object rather than old and current values. Store regular and sale prices separately when available and validate that the sale amount is lower.
Playwright times out Wrong selector, slow request, consent wall or blocked browser. Confirm the selector in rendered DevTools, wait for a specific element, inspect a screenshot/HTML dump, and classify CAPTCHA or login pages as failures.
HTTP 403 or CAPTCHA The target blocks the request or requires human verification. Stop aggressive retries, review terms and access requirements, and use an authorized integration or data feed if available.
False price-drop alerts Empty shells, unavailable prices or parser drift were stored as valid data. Enforce required-field and bounds validation, retain error states and alert on parser-success changes.
Runs become slow and expensive Too many browser launches or excessive concurrency. Reuse browser contexts, prefer static fetching where possible, pace requests and measure latency before adding workers.

Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server for developers. It accepts a URL and returns a PNG, JPEG, WebP or PDF; for a price-monitoring workflow, the captured image can provide a visual record alongside your structured parser. Before capture it accepts the cookie or consent banner like a visitor and removes more than 60 known consent platforms, newsletter popups and chat widgets, with each step switchable. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads and cache hits are not billed, and each response identifies the result with X-Page-Verdict and X-Billed headers.

Use the API key and URL shown in the ScreenshotNeo documentation:

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

ScreenshotNeo also offers an MCP server for Claude, Cursor and other MCP clients, with take_screenshot, get_page_info and capture_pdf tools. Options include full-page and CSS-selector captures, lazy-image loading, dark mode, device presets, arbitrary viewports, retina scale, PDF paper and page settings, custom CSS or JavaScript, clicks, selector or network-idle waits, request blocking, headers, cookies, user agents, authorization, timezone, geolocation, transparent backgrounds, resizing, chosen cache TTLs, signed links, asynchronous jobs with signed webhooks, bulk capture for up to 100 URLs per call, a usage API and an OpenAPI specification. Existing parameter names used by other screenshot APIs also work.

The Free plan includes 1,000 screenshots per month with no card. Paid plans start at $5 for 3,000 shots; every feature is on every plan, and yearly billing gives two months free. Create a free ScreenshotNeo account to try it.

Frequently Asked Questions

Can one scraper monitor several sellers for the same product?

Yes. Model each seller offer as a separate observation keyed by product and seller, then compare currency, availability and timestamps independently before producing a cross-seller view.

How should I test a parser after a site redesign?

Keep representative HTML fixtures or permitted content hashes, run them in continuous integration, and require the parser to produce a valid identity, currency and price before deployment.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Should the scraper save screenshots for every run?

Only when visual evidence is useful and the target’s terms permit it. A content hash is cheaper for routine debugging; retain a full image for failed or disputed observations according to your retention policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.