The Tool Desk
Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
To track a product price with Python, first choose a data source you are allowed to access, then retrieve and parse the price, normalize its amount and currency, and save a timestamped observation. Compare that observation with a previous price or a target threshold before sending an alert. For a small number of permitted, server-rendered pages, an HTTP client and HTML parser may be enough; recurring crawls, dynamic pages, or account-gated data call for a different approach.
A price tracker is only as trustworthy as its source and extraction rules. This guide builds a cautious workflow, explains the main implementation choices, and shows how to handle history, alerts, failures, and Amazon-specific access requirements without treating page scraping as universally permitted.
Choose a permitted source before writing a scraper
Price-tracking code does not make a source available or permitted. Check the site’s current terms and its robots.txt instructions before scheduling requests. Those checks are useful, but robots.txt is not a complete legal opinion: what is allowed can depend on the site’s terms, the facts, and the jurisdiction.
Recommended Free Tools
Python’s urllib.robotparser can read robots.txt and answer whether a particular user agent may fetch a URL under the published rules. The Python 3.14 documentation describes its RobotFileParser as a class that answers whether a user agent can fetch a URL on the site that published the file: Python urllib.robotparser documentation.
#1 Best Overall
Start with the least complex supported route
- Official API: Prefer one when it provides the data for your intended use and you qualify for access. API terms, account prerequisites, and permitted uses still apply.
- HTTP client and HTML parser: Suitable for a small number of permitted pages whose product data is present in the normal server response. Requests and Beautiful Soup are common Python tools for this kind of work; a parser does not bypass access restrictions or render dynamic content. See Real Python’s Python web scraping tutorials.
- Scrapy: Consider a project and spider when you need repeated crawling, pagination, link following, structured extraction, or exports. Its tutorial covers spiders, extraction, following links, and exporting: Scrapy Tutorial.
- Dynamic page: If the relevant data is not in the normal response, look for an official or otherwise permitted data route first. A rendering strategy may be appropriate only if site rules allow it; do not bypass bot checks or other access controls.
Compare routes by permission and official support, price freshness, geography and currency consistency, reliability, scale, maintenance, and cost. The cited documentation does not establish an apples-to-apples speed or cost benchmark, so the right choice depends on your workload rather than a universal ranking.
Build a small tracker with Python
The example below is a pattern for a page you have permission to fetch, not a claim that any live retailer uses its selectors or markup. Replace the example URL and CSS selectors after inspecting an allowed response. The script checks robots.txt, requests the page with a timeout, extracts a displayed price, stores an observation in CSV, and prints an alert when a configured threshold is reached.
Install the two third-party packages with python -m pip install requests beautifulsoup4. Save this as price_tracker.py and run it with python price_tracker.py.
Quick wins for a faster PC:
Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Clear out junk files and repair common Windows errorsFree Scan →Scan for outdated or missing drivers - takes under a minuteDriver Scan →from datetime import datetime, timezone
from decimal import Decimal, InvalidOperation
from pathlib import Path
from urllib.parse import urlparse
from urllib.robotparser import RobotFileParser
import csv
import re
import requests
from bs4 import BeautifulSoup
URL = "https://example.com/product"
USER_AGENT = "PriceTracker/1.0 (contact: [email protected])"
PRICE_SELECTOR = ".price" # Change to a selector on an allowed page
NAME_SELECTOR = "h1" # Change if the product name is elsewhere
CURRENCY = "USD" # Set from a reliable source; do not guess
TARGET_PRICE = Decimal("50.00")
CSV_PATH = Path("price_history.csv")
def allowed_by_robots(url: str, user_agent: str) -> bool:
parts = urlparse(url)
robots_url = f"{parts.scheme}://{parts.netloc}/robots.txt"
parser = RobotFileParser(robots_url)
parser.read()
return parser.can_fetch(user_agent, url)
def parse_amount(displayed: str) -> Decimal:
"""Parse a simple decimal display such as '$49.99'; adapt to the locale."""
cleaned = re.sub(r"[^0-9.,]", "", displayed).strip()
if not cleaned:
raise ValueError(f"No numeric price found in {displayed!r}")
# This example assumes a period is the decimal mark and commas are grouping.
# Change this rule for the page's locale; ambiguous values should be rejected.
cleaned = cleaned.replace(",", "")
try:
amount = Decimal(cleaned)
except InvalidOperation as exc:
raise ValueError(f"Malformed price: {displayed!r}") from exc
if amount <= 0:
raise ValueError(f"Price must be positive, got {amount}")
return amount
def append_observation(product: str, amount: Decimal, currency: str) -> None:
new_file = not CSV_PATH.exists()
with CSV_PATH.open("a", newline="", encoding="utf-8") as file:
writer = csv.DictWriter(
file,
fieldnames=["observed_at", "product", "amount", "currency", "source", "displayed_price"],
)
if new_file:
writer.writeheader()
writer.writerow({
"observed_at": datetime.now(timezone.utc).isoformat(),
"product": product,
"amount": str(amount),
"currency": currency,
"source": URL,
"displayed_price": "",
})
def main() -> None:
if not allowed_by_robots(URL, USER_AGENT):
raise SystemExit("robots.txt does not allow this user agent to fetch the URL")
try:
response = requests.get(
URL,
headers={"User-Agent": USER_AGENT},
timeout=(5, 20),
)
response.raise_for_status()
except requests.RequestException as exc:
raise SystemExit(f"Page request failed: {exc}") from exc
soup = BeautifulSoup(response.text, "html.parser")
name_node = soup.select_one(NAME_SELECTOR)
price_node = soup.select_one(PRICE_SELECTOR)
if name_node is None or price_node is None:
raise SystemExit("Product name or price selector was not found; review the page and selectors")
product = name_node.get_text(" ", strip=True)
displayed = price_node.get_text(" ", strip=True)
amount = parse_amount(displayed)
append_observation(product, amount, CURRENCY)
print(f"Recorded {product}: {CURRENCY} {amount} at {datetime.now(timezone.utc).isoformat()}")
if amount <= TARGET_PRICE:
print(f"ALERT: {product} is at or below {CURRENCY} {TARGET_PRICE}")
if __name__ == "__main__":
main()
The sample assumes a simple price format with a period as the decimal mark, a comma as the thousands separator, and a currency configured separately. Real pages may show ranges, unit prices, sale prices, membership prices, or localized formats. Do not silently convert an ambiguous string into a number: adapt the parser and record enough context to explain what the amount represents.
Rank #2
Check robots.txt before polling
The example reads the site’s robots.txt and calls can_fetch for the URL. If robots.txt cannot be read, the sample’s read() behavior should not be treated as a full permission check; investigate the site’s instructions and terms, and choose a conservative policy for unavailable rules. A permitted robots.txt result does not override other terms or legal requirements.
Inspect selectors and verify the extracted value
Use browser developer tools or inspect a permitted HTTP response to identify a stable product-name and price element. Validate a few responses manually before relying on the result. A selector can keep returning a value after the page changes meaning—for example, selecting a crossed-out list price rather than the current offer—so verify the displayed label and offer context as well as the digits.
Normalize prices and preserve useful history
A tracker should retain more than a number. At minimum, store the product identity, amount, currency, observation time, source URL, and relevant offer context. Keep the original displayed string or a separate raw field when practical; it helps diagnose parsing changes later.
Do these 3 things before closing this tab:
1Clear out junk files and repair common Windows errors2Scan for outdated or missing drivers - takes under a minute3Repair Windows errors before they cause bigger problems- Use decimal arithmetic: Python’s
Decimalavoids binary floating-point surprises in monetary values. - Make locale explicit:
1,299.00and1.299,00can mean different values. Choose a parsing rule for the source’s locale and flag ambiguous strings rather than guessing. - Separate amount and currency: A number without currency is not comparable across regions. If a page does not make currency clear, do not assign one based on assumption.
- Capture offer context: Distinguish regular price, sale price, unit price, membership-only price, shipping, and other conditions that matter to the buyer.
- Keep a history: Append observations rather than overwriting the last value. A CSV works for a small personal project; larger workflows can use SQLite, PostgreSQL, or another database. Real Python’s resource index covers these storage options along with CSV, JSON, and MongoDB.
Once the history outgrows a single script, a useful record often includes a stable product ID, normalized amount, currency, UTC timestamp, source, raw display text, and parse status. That makes it easier to detect duplicate observations and distinguish a real change from a broken selector.
Compare observations and send meaningful alerts
The sample checks whether the latest amount is at or below a user-defined target. Another common rule compares the new observation with the previous valid one. For a simple implementation, read the last row for the same product and currency, then alert only if the normalized amount differs. If you store several products, use their stable identifiers rather than product-name text alone.
Do not send an alert from an unvalidated parse. A missing selector, malformed amount, currency change, or sudden implausible jump should be logged for review rather than treated as a bargain. Keep alert delivery separate from extraction—for example, call an email, messaging, or notification service only after the observation has passed validation and the chosen condition is true.
Run on a schedule conservatively
Use a scheduler appropriate to your environment, such as cron or a task scheduler, but choose the interval based on the source’s rules and your actual freshness need. The cited sources do not establish a universal polling interval. Use modest request rates, retries for transient failures, and caching where suitable; avoid repeatedly fetching unchanged content. A retry should be bounded and should not turn a temporary error into a burst of requests.
Scale up when the workflow needs structure
Scrapy for recurring multi-page work
Scrapy provides a project structure for requests, callbacks, extraction, following links, and exporting data. That can be useful for a permitted crawl with pagination or many product pages. It also adds setup and maintenance: start a project when its structure solves a real need, not simply because the word “scraping” appears in the task. Follow the Scrapy tutorial for its documented spider and export workflow.
Managed scraping APIs
A hosted scraping API may suit a developer who wants a managed request and dataset workflow rather than maintaining all crawl infrastructure. Scrapy.io documents synchronous and asynchronous runs, dataset export, scheduling, and a Python SDK: Scrapy.io Web Scraping API documentation. Cost and suitability depend on the service and workload; the available documentation does not establish a comparative benchmark.
When a screenshot is useful—and when it is not
A screenshot can help inspect what a rendered page looks like, but an image is not a reliable substitute for structured price data when you need a normalized amount, currency, and offer context. ScreenshotNeo is a website screenshot API and MCP server for developers: ScreenshotNeo. Its one-call capture can be useful when a visual record is part of your workflow, but your tracker still needs a permitted, validated source for the actual price fields.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Amazon product data has specific access conditions
Do not assume that scraping Amazon product pages is automatically allowed or that it avoids API requirements. Amazon’s Selling Partner API Product Pricing API describes retrieval of catalog pricing and offer information for automated seller price management and repricing; that is a seller-oriented use case, not proof of universal access for a personal tracker. Consult the current Product Pricing API documentation.
Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchWindows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallAmazon Associates’ Product Advertising API has separate requirements. Its help page says a user must have an open Associates account, comply with the Associates Operating Agreement, apply for PA-API, and follow the API License Agreement: PA-API requirements. The cited help page also states an initial allowance of one request per second, with increases tied to shipped revenue attributed to the relevant account: Amazon’s request-rate guidance. These are documented conditions, not a guarantee that a particular reader will qualify or receive access. Check the current official terms before implementation; eligibility and applicable rules can change.
Best Value
Troubleshoot common failures
- Robots check returns false: Do not fetch the page with that user agent under the published robots rules. Reassess the source and choose a permitted route; do not attempt to evade the restriction.
- Request times out or returns an error: Confirm the URL and network access, retain a finite timeout, and handle transient failures with bounded retries. Do not repeatedly hammer a failing endpoint.
- Selector is missing: The page may have changed, the content may be dynamically rendered, or the selector may be wrong. Review an allowed response and the applicable source rules before adjusting the extraction method.
- Price parses incorrectly: Check decimal and grouping separators, currency symbols, ranges, and whether the selected element is a sale or regular price. Reject values that remain ambiguous.
- Price suddenly changes dramatically: Verify the source, product identity, currency, and offer conditions; a selector may have latched onto shipping, unit pricing, or unrelated text.
- History contains duplicates: Decide whether each scheduled observation should be stored or whether identical consecutive values should be coalesced. If coalescing, retain observation timestamps or a count so unchanged periods are not mistaken for missing checks.
- Alerts are noisy: Alert on a threshold crossing or a meaningful change, not every poll. Require a valid parse and consistent currency before evaluating the rule.
Or skip the browser setup
If a visual capture is useful alongside your price records, ScreenshotNeo can return a screenshot or PDF with one GET request. Its API accepts URL capture and has options such as waiting for a selector or network idle, choosing a viewport, and capturing a full page. Use an official data source or a permitted extraction method for the price itself; a screenshot does not normalize price values.
Example cURL request (replace the URL with a page you are allowed to capture and use your API key):
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
See the ScreenshotNeo API documentation for the request options. ScreenshotNeo removes cookie/consent banners, newsletter popups, and chat widgets before capture; bot checks, blank pages, and failed loads are not billed. Its MCP server lets AI agents use screenshot tools, and its free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 screenshots. Sign up for 1,000 free screenshots a month, with no card required.
Frequently Asked Questions
Is web scraping legal?
There is no universal answer for every site and jurisdiction. Review applicable terms and access rules, and get legal advice for a consequential use; robots.txt alone is not a legal determination.
Can a screenshot API extract a price for my tracker?
A screenshot is a visual image, not normalized product data. Use a permitted structured source and validate price, currency, and offer context; a screenshot may serve as a visual record.
Should I use CSV or a database?
CSV can suit a small single-script tracker. A database is more appropriate when you need multiple products, concurrent jobs, or structured querying; choose based on project needs.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.

