Driver FixRecommendedSound, Wi-Fi or graphics acting up? Check drivers firstFind missing or outdated drivers fast.Check DriversOctober DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsClean PCRecommendedOne scan can reveal what keeps slowing WindowsLook for cleanup and repair opportunities.Run Scan×
Skip to content
Automation

How to Build an Automated Price Tracker with Python Web Scraping

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Build a price tracker as a repeatable pipeline: identify a product and variant, retrieve a permitted source, extract and validate the price, store a timestamped observation, compare it with a baseline, and send an alert only when a defined condition is met. The Python example below uses an HTTP client, an HTML parser, SQLite history, robots.txt checks, and a command-line threshold alert. It is suitable for a small number of pages whose prices are present in the returned HTML; use an official API or feed when a retailer provides one.

Design the tracker before writing code

A reliable tracker is more than a loop around requests.get(). Define the data and failure behavior first.

Use an explicit product configuration

Store a stable product identifier, retailer, URL, currency, and extraction rule. A title alone is unsafe because color, size, storage capacity, seller, and bundle variants can have different prices. Keep one configuration per variant.

PRODUCTS = [
    {
        "product_id": "example-widget-blue-128gb",
        "retailer": "Example Store",
        "url": "https://example.com/products/widget",
        "currency": "USD",
        "price_selector": "[itemprop='price']",
    }
]

Choose the permitted source

Look for an official API, product feed, or other documented integration before parsing HTML. Read the retailer’s current terms and its robots.txt for the exact user agent and path. Python’s urllib modules handle URLs, and urllib.robotparser.RobotFileParser answers whether a user agent may fetch a URL under the site’s published robots rules. AWS also describes retrieving robots.txt as part of crawler setup in its crawler guidance. Robots rules do not settle every contractual or legal question; if access is disallowed, stop or use a permitted source.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Decide what an observation means

Record the source time, product and variant identity, currency, and price. If useful, also capture stock state, promotion text, seller, or location context. A displayed price is an observation from one source and time, not a guaranteed checkout total: tax, shipping, location, promotion, and availability can change later.

A complete small tracker in Python

Install the two third-party packages:

python -m pip install requests beautifulsoup4

Save this as tracker.py. It checks robots.txt, fetches conservatively, parses a currency-bearing price, validates the result, appends history to SQLite, and prints an alert when the latest price is at or below a target.

from __future__ import annotations

import re
import sqlite3
import time
from datetime import datetime, timezone
from decimal import Decimal, InvalidOperation
from urllib.parse import urljoin, urlparse
from urllib.robotparser import RobotFileParser

import requests
from bs4 import BeautifulSoup

USER_AGENT = "ExamplePriceTracker/1.0 (+replace-with-your-contact)"
DB_PATH = "prices.sqlite3"
PRODUCTS = [
    {
        "product_id": "example-widget-blue-128gb",
        "retailer": "Example Store",
        "url": "https://example.com/products/widget",
        "currency": "USD",
        "price_selector": "[itemprop='price']",
        "target_price": Decimal("49.99"),
    }
]

def allowed_by_robots(url: str) -> bool:
    parts = urlparse(url)
    robots_url = f"{parts.scheme}://{parts.netloc}/robots.txt"
    parser = RobotFileParser(robots_url)
    parser.read()
    return parser.can_fetch(USER_AGENT, url)

def parse_price(html: str, selector: str) -> Decimal:
    soup = BeautifulSoup(html, "html.parser")
    node = soup.select_one(selector)
    if node is None:
        raise ValueError(f"price selector not found: {selector}")
    raw = node.get("content") or node.get_text(" ", strip=True)
    cleaned = re.sub(r"[^0-9.,-]", "", raw)
    if not cleaned:
        raise ValueError(f"price text is empty: {raw!r}")
    # This example expects a dot decimal format, such as 49.99.
    cleaned = cleaned.replace(",", "")
    try:
        value = Decimal(cleaned)
    except InvalidOperation as exc:
        raise ValueError(f"cannot parse price: {raw!r}") from exc
    if value < 0 or value > Decimal("100000000"):
        raise ValueError(f"implausible price: {value}")
    return value

def init_db(connection: sqlite3.Connection) -> None:
    connection.execute("""
        CREATE TABLE IF NOT EXISTS observations (
            id INTEGER PRIMARY KEY,
            product_id TEXT NOT NULL,
            observed_at TEXT NOT NULL,
            price TEXT NOT NULL,
            currency TEXT NOT NULL,
            source_url TEXT NOT NULL
        )
    """)
    connection.commit()

def previous_price(connection: sqlite3.Connection, product_id: str):
    row = connection.execute(
        "SELECT price FROM observations WHERE product_id=? ORDER BY id DESC LIMIT 1",
        (product_id,),
    ).fetchone()
    return Decimal(row[0]) if row else None

def record(connection, product, price: Decimal) -> None:
    connection.execute(
        "INSERT INTO observations(product_id, observed_at, price, currency, source_url) VALUES (?, ?, ?, ?, ?)",
        (product["product_id"], datetime.now(timezone.utc).isoformat(),
         str(price), product["currency"], product["url"]),
    )
    connection.commit()

def fetch_product(session: requests.Session, product):
    url = product["url"]
    if not allowed_by_robots(url):
        raise PermissionError(f"robots.txt disallows {url}")
    response = session.get(url, timeout=(10, 30), allow_redirects=True)
    response.raise_for_status()
    content_type = response.headers.get("content-type", "").lower()
    if "html" not in content_type:
        raise ValueError(f"unexpected content type: {content_type}")
    return parse_price(response.text, product["price_selector"])

def main() -> None:
    with sqlite3.connect(DB_PATH) as connection:
        init_db(connection)
        with requests.Session() as session:
            session.headers.update({"User-Agent": USER_AGENT, "Accept": "text/html"})
            for index, product in enumerate(PRODUCTS):
                try:
                    price = fetch_product(session, product)
                    old = previous_price(connection, product["product_id"])
                    record(connection, product, price)
                    change = "first observation" if old is None else f"change {price - old:+.2f}"
                    print(f"{product['product_id']}: {price} {product['currency']} ({change})")
                    if price <= product["target_price"] and (old is None or old > product["target_price"]):
                        print(f"ALERT: {product['product_id']} reached {price} {product['currency']}")
                except Exception as exc:
                    # A failure is not a zero-price observation.
                    print(f"ERROR {product['product_id']}: {exc}")
                if index < len(PRODUCTS) - 1:
                    time.sleep(2)

if __name__ == "__main__":
    main()

Replace the example URL and selector with values documented or permitted by your retailer. Run it with python tracker.py. The first successful run creates prices.sqlite3; later runs append rows instead of overwriting history.

Make extraction and comparisons trustworthy

Prefer machine-readable price fields

Selectors such as itemprop="price" or a retailer’s documented JSON field are generally less ambiguous than a broad text search. Keep the currency alongside the number. Locale formats need an explicit rule: blindly removing commas turns some European values into the wrong amount. Validate that the selected node belongs to the intended product and variant, not a recommendation card or crossed-out “was” price.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Handle promotions and availability deliberately

Decide whether the tracker records sale price, regular price, or both. Capture promotion and stock context when the decision depends on it. Never convert a missing, blocked, empty, or malformed value to zero. Log the failure and investigate a changed page, consent wall, bot check, timeout, or selector.

Choose a comparison policy

  • Previous observation: report every validated change.
  • Baseline: compare with the first or user-selected reference.
  • Target: alert when a price crosses below a threshold, as the example does.
  • Percentage: alert only when the change exceeds a chosen percentage.

Persist an alert state or last-alerted value if notifications are added, so unchanged prices do not generate repeated messages. Email, a webhook, or a team-chat integration can consume the same event after the database write succeeds.

Scheduling, scale, and storage

Schedule according to need and permission

There is no universal correct polling interval. Use the least frequent schedule that meets your use case and the retailer’s permitted request volume. A system scheduler can run the script, for example with cron:

17 * * * * /usr/bin/python3 /opt/price-tracker/tracker.py >> /var/log/price-tracker.log 2>&1

For a few products, SQLite is adequate. As the catalog grows, retain the same logical columns in a relational database and add indexes on (product_id, observed_at). Keep retrieval logs separate from observations so operational failures cannot be mistaken for prices.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Reduce load and improve repeatability

  • Use one persistent HTTP session, explicit connect/read timeouts, and a descriptive user agent.
  • Space requests rather than firing a burst; do not bypass blocks or access controls.
  • Cache only when the retailer’s rules allow it and when stale data is acceptable.
  • Pin and update dependencies, record redirects and response status, and alert on repeated parse failures.
  • For client-rendered pages, first seek an official feed. A basic HTTP parser cannot see data that JavaScript adds after the response.

Common failures and fixes

Robots check fails or access is not permitted

Confirm the robots URL, user-agent string, and path. Re-read the retailer’s terms. Do not treat a publicly viewable page as automatic permission; switch to an approved API/feed or stop.

Selector not found

Save a diagnostic response, inspect whether the page is a consent or block page, and verify the variant URL. Prefer a stable attribute or documented field over generated class names. Update the configuration and add a test fixture before resuming alerts.

HTTP 403, 429, or a timeout

Respect the response. Reduce frequency, honor retry guidance, and check whether an official integration exists. Do not attempt to defeat CAPTCHA, bot detection, or other access controls.

Wrong currency or implausible value

Bind each product to an expected currency and region, then reject values outside a sensible range. Confirm that the selector is not reading a list price, installment amount, shipping charge, or another variant.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The price appears only after JavaScript runs

Inspect the retailer’s permitted API or embedded structured data first. A browser automation workflow adds complexity and should still obey access rules; it is not a reason to bypass them.

Duplicate or noisy alerts

Store the last alert condition and notify only on a threshold crossing or a meaningful change. A failed fetch must leave the prior valid observation untouched.

Monetization and Amazon-specific caution

If you publish a tracker or alerting site, read the current program terms for every data source. Amazon Associates’ Operating Policies state: “Unless otherwise agreed by Amazon, your Site must not have price tracking and/or price alerting functionality.” The same policy limits use of Program Content and disallows data mining, robots, or similar extraction tools for that content. An affiliate link or product-data permission should not be assumed to authorize a tracker. Verify any required agreement before monetizing one.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup

When you need screenshots of a rendered page for an audit, catalog, or visual check, ScreenshotNeo provides a one-call API. It accepts cookie and consent banners as a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and the response identifies the result with X-Page-Verdict and X-Billed headers. Its MCP server exposes take_screenshot, get_page_info, and capture_pdf to Claude, Cursor, and other MCP clients.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

See the ScreenshotNeo documentation for all options, including full-page lazy-image loading, CSS-selector element capture, device presets, retina scale, PDF ranges and margins, custom CSS or JavaScript, clicks, waits, request blocking, headers, cookies, user agents, authorization, timezone, geolocation, transparent backgrounds, resizing, TTL caching, signed links, asynchronous webhooks, bulk capture of up to 100 URLs per call, usage reporting, and the OpenAPI specification.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

The Free plan includes 1,000 screenshots per month with no card. Paid plans start at $5 for 3,000 shots; every feature is on every plan. Sign up free for ScreenshotNeo.

FAQ

Should I scrape every minute?

No. Set the least frequent interval that meets the use case and the retailer’s permitted volume; no universal interval is established.

Can a tracker guarantee the checkout price?

No. It records a timestamped observation that may differ at checkout because of tax, shipping, location, stock, seller, or promotion changes.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What should happen when parsing fails?

Log the failure, preserve the last valid observation, and investigate. Never store zero or send a price-drop alert for an unvalidated response.

Frequently Asked Questions

Should I scrape every minute?

No. Set the least frequent interval that meets the use case and the retailer’s permitted volume; no universal interval is established.

Can a tracker guarantee the checkout price?

No. It records a timestamped observation that may differ at checkout because of tax, shipping, location, stock, seller, or promotion changes.

What should happen when parsing fails?

Log the failure, preserve the last valid observation, and investigate. Never store zero or send a price-drop alert for an unvalidated response.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Read next

Recommended PC Tool
Recommended PC Tool
PC Slower Than It Used to Be?Free scan - under a minute
Outdated Drivers Are Slowing You DownFree scan - exact matches

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.