Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

Use a successful fetch, meaningful-text normalization, a SHA-256 digest, and a persisted previous snapshot to detect website changes reliably. The digest answers whether content changed; the saved normalized text lets Python show exactly which lines changed. Treat the first successful fetch as a baseline, never overwrite a good baseline after a failed request, and schedule the checker with cron or another recurring runner.

This implementation covers static and JavaScript-rendered pages, selected content regions, failure handling, unified diffs, retention, notifications, and an optional screenshot workflow.

The six-stage design

  1. Fetch: request the URL with a timeout and record the status code.
  2. Reduce: remove scripts, styles, navigation, footers, and other noise; extract visible text.
  3. Normalize: collapse repeated whitespace so formatting-only HTML changes do not trigger alerts.
  4. Fingerprint: encode the normalized text as UTF-8 and calculate a SHA-256 hexadecimal digest.
  5. Compare and persist: compare the new digest with the previous record, then save the new digest and text only after a successful, non-empty fetch.
  6. Report: print a baseline, unchanged result, or unified diff and send a notification if required.

SHA-256 is a one-way fingerprint, not a description of the change. A one-character input change produces a different digest, so retaining the normalized text is essential for a useful diff.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A complete Python tracker

The following script uses requests and beautifulsoup4. Install them in a virtual environment with python -m pip install requests beautifulsoup4. It accepts multiple URLs, stores one record per URL in a JSON state file, and exits with a failure status if any fetch fails.

#!/usr/bin/env python3
import argparse
import difflib
import hashlib
import json
import logging
import os
import tempfile
from datetime import datetime, timezone
from pathlib import Path

import requests
from bs4 import BeautifulSoup


def utc_now():
    return datetime.now(timezone.utc).isoformat()


def normalize_html(html, selector=None):
    soup = BeautifulSoup(html, 'html.parser')
    for tag in soup(['script', 'style', 'nav', 'footer', 'noscript']):
        tag.decompose()
    root = soup
    if selector:
        root = soup.select_one(selector)
        if root is None:
            raise ValueError(f'CSS selector matched nothing: {selector}')
    elif soup.body is not None:
        root = soup.body
    text = root.get_text(' ', strip=True)
    return ' '.join(text.split())


def load_state(path):
    if not path.exists():
        return {}
    with path.open('r', encoding='utf-8') as handle:
        value = json.load(handle)
    if not isinstance(value, dict):
        raise ValueError('state file must contain a JSON object')
    return value


def save_state(path, state):
    path.parent.mkdir(parents=True, exist_ok=True)
    fd, temporary = tempfile.mkstemp(prefix=path.name + '.', dir=path.parent)
    try:
        with os.fdopen(fd, 'w', encoding='utf-8') as handle:
            json.dump(state, handle, ensure_ascii=False, indent=2)
            handle.write('n')
        os.replace(temporary, path)
    except Exception:
        try:
            os.unlink(temporary)
        except FileNotFoundError:
            pass
        raise


def check_url(url, state, timeout, selector):
    try:
        response = requests.get(
            url,
            timeout=timeout,
            headers={'User-Agent': 'PythonWebsiteChangeTracker/1.0'},
        )
        response.raise_for_status()
        text = normalize_html(response.text, selector)
        if not text:
            raise ValueError('normalized response is empty')
    except Exception as exc:
        logging.error('%s: fetch failed: %s', url, exc)
        return False

    digest = hashlib.sha256(text.encode('utf-8')).hexdigest()
    previous = state.get(url)
    if previous is None:
        print(f'BASELINE {url} sha256={digest}')
    elif previous.get('digest') == digest:
        print(f'UNCHANGED {url} sha256={digest}')
    else:
        print(f'CHANGED {url}')
        old_lines = previous.get('text', '').splitlines()
        new_lines = text.splitlines()
        for line in difflib.unified_diff(
            old_lines,
            new_lines,
            fromfile='previous',
            tofile='current',
            lineterm='',
        ):
            print(line)

    state[url] = {
        'digest': digest,
        'text': text,
        'checked_at': utc_now(),
        'status_code': response.status_code,
        'content_type': response.headers.get('content-type', ''),
    }
    return True


def main():
    parser = argparse.ArgumentParser(description='Detect meaningful webpage changes')
    parser.add_argument('urls', nargs='+')
    parser.add_argument('--state', default='snapshots.json')
    parser.add_argument('--timeout', type=float, default=30)
    parser.add_argument('--selector', help='CSS region to monitor instead of the whole body')
    args = parser.parse_args()

    logging.basicConfig(level=logging.INFO, format='%(asctime)s %(levelname)s %(message)s')
    state_path = Path(args.state)
    state = load_state(state_path)
    all_ok = True
    for url in args.urls:
        if not check_url(url, state, args.timeout, args.selector):
            all_ok = False
    save_state(state_path, state)
    raise SystemExit(0 if all_ok else 1)


if __name__ == '__main__':
    main()

Save it as watch.py, then make the first run your baseline:

python watch.py https://example.com/pricing --state /var/lib/sitewatch/snapshots.json

The first successful observation prints BASELINE. A later identical digest prints UNCHANGED; a different digest prints a unified diff. The state record also keeps the UTC timestamp, HTTP status, and content type for diagnosis.

Choose what counts as a change

Monitor a focused region

Whole-page text includes unrelated updates such as rotating recommendations, timestamps, advertisements, and cookie controls. Pass a CSS selector to hash only the intended region:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
python watch.py https://example.com/pricing --selector '#plans' --state snapshots.json

If the selector is absent, the script treats that check as a failure and preserves the previous good record. For several regions, select a stable container that contains all of them or extend normalize_html to concatenate named selections in a fixed order.

Remove more volatile elements

The script removes script, style, nav, footer, and noscript. Add site-specific selectors before calling get_text when banners, clocks, ad slots, or recommendation widgets create noise. Keep the rule deterministic: the same source should produce the same normalized string on every run.

Text versus visual monitoring

Text hashing detects copy, prices, links, and other textual changes. It does not detect a color, image, layout, or font change when the visible text remains identical. For visual assurance, capture a rendered image separately and hash the image bytes or use an image-diff tool; expect dynamic animations and rotating content to require masking or a stable capture state.

Static HTML, JavaScript, and access controls

When requests is sufficient

Use the script when the response contains the article, price table, policy, or other content you intend to monitor. Inspect response.text or save a sample response if you are unsure.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

When the response is an application shell

Client-rendered sites may return a nearly empty HTML shell and load the real content through JavaScript. A successful HTTP status in that case is not evidence that the intended content was fetched. Use a browser-capable crawler, a headless browser, or an official API/change feed when one exists. Prefer an official change feed when it exposes the exact data you need; it is usually more stable than scraping presentation HTML.

Authentication and robots policies

Only monitor pages you are authorized to access. If a site requires authentication, use an approved API or carefully managed session cookies rather than embedding credentials in a cron command. Respect the site’s terms, rate limits, and robots policy. Never treat a bot-check or CAPTCHA response as an unchanged page.

Persisting history and sending alerts

Latest-state storage

The JSON file is a minimal latest-state database. Atomic replacement prevents a process interruption from leaving a truncated file. Protect it with filesystem permissions because the saved text may contain private or paywalled material.

Timestamped snapshots

For auditability, write each successful response to a directory such as history/<url-key>/2026-09-29T12-00-00Z.json containing the digest, normalized text, status, content type, and fetch time. Apply a retention policy, for example keeping daily snapshots for a defined period and deleting older files. Do not create a history entry for a timeout, HTTP error, empty normalization result, or selector mismatch.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Notifications

Put notification code after the new state has been written. Email, a chat webhook, or an incident system can receive the URL, timestamp, old digest, new digest, and unified diff. Send only after a successful fetch and persisted snapshot. If notification delivery fails, log that failure separately so it is not confused with a page fetch failure.

Run it on a schedule

Hourly cron

Create a dedicated virtual environment and use absolute paths in cron. Edit the crontab with crontab -e:

0 * * * * /opt/sitewatch/.venv/bin/python /opt/sitewatch/watch.py https://example.com/pricing --state /var/lib/sitewatch/snapshots.json >> /var/log/sitewatch.log 2>&1

The command runs at minute zero of every hour in the host’s local cron timezone. Use UTC on servers when you need timestamps and schedules that do not shift with daylight-saving changes. Ensure the cron user can read the script, write the state directory, and append to the log.

Avoid overlapping runs

If a slow page can exceed the interval, use a lock such as flock around the command or run through a queue with one worker per state file. Concurrent writers can otherwise race and overwrite a newer snapshot.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

In-process intervals

An always-on worker that sleeps between checks is suitable for a small deployment, but it must handle restarts, exceptions, and clock changes. Cron is simpler for one-shot checks because the operating system supplies restart behavior and execution history.

Troubleshooting

Every run reports a change

  • Inspect the stored diff for timestamps, ads, consent text, or rotating recommendations.
  • Pass a stable CSS selector and remove volatile descendants during normalization.
  • If the server varies content by location or user agent, pin those request attributes where permitted.

Every run reports unchanged although the page visibly changed

  • The change may be an image, CSS rule, layout, or client-side state not present in the fetched HTML.
  • Use a browser-capable fetcher for JavaScript content or add a visual capture check.
  • Confirm that your selector includes the changed region and that normalization did not remove it.

HTTP 403, 429, or CAPTCHA responses

Do not save the response as a new baseline. Slow the schedule, identify yourself appropriately, use an authorized API, or use a browser-capable crawler where allowed. A bot-check page is a failed observation, not a valid page snapshot.

Timeouts and connection errors

Increase the timeout only after checking network reliability, and record the exception. Retry with bounded backoff in a worker if transient failures are common. Keep the previous digest until a complete, non-empty response succeeds.

JSON state corruption or permission errors

Check ownership and free disk space, stop overlapping writers, and restore the last known-good copy. The script’s atomic write protects against partial writes but cannot fix an invalid file deliberately edited by another process.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Diffs are one giant line

The sample intentionally collapses whitespace, so the normalized document is one logical line. For line-oriented diffs, change normalization to preserve paragraph or block boundaries, for example by joining each selected element with a newline. Keep that policy stable or the next run will show formatting noise.

Performance, reliability, and cost considerations

SHA-256 computation is linear in the normalized text size and normally inexpensive compared with downloading or rendering the page. The expensive variables are network latency, browser startup, response size, and how much history you retain. Measure fetch duration, response size, error rate, false-positive rate, and storage growth in your own deployment rather than relying on a generic benchmark.

Hashing selected text reduces both false positives and stored data. Keep response metadata even when you do not retain raw HTML. For many URLs, process them sequentially first; move to bounded concurrency only after setting per-host limits and ensuring each worker has isolated or locked state handling.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup

When you need a rendered screenshot instead of scraped text, ScreenshotNeo provides a website screenshot API and MCP server. One GET request returns PNG, JPEG, WebP, or PDF. It accepts cookie and consent banners like a visitor, then removes more than 60 known consent platforms along with newsletter popups and chat widgets; each cleanup step can be disabled. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing result.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use the API directly (the ScreenshotNeo documentation lists all parameters):

curl -G 'https://api.screenshotneo.com/v1/shot' -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

For a Python capture:

import requests
r = requests.get('https://api.screenshotneo.com/v1/shot', params={'access_key': 'YOUR_API_KEY', 'url': 'https://stripe.com'}, timeout=90)
r.raise_for_status()
open('shot.webp', 'wb').write(r.content)

For Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`HTTP ${res.status}`);
require('fs').writeFileSync('shot.webp', Buffer.from(await res.arrayBuffer()));

To track visual changes, store the returned file with a timestamp, hash its bytes with SHA-256, and compare the digest on the next run. A visual hash is sensitive to legitimate layout changes, so use stable viewport, device, wait, and masking settings.

Capture controls useful for a tracker

  • Full-page captures can load lazy images; alternatively capture one element with a CSS selector.
  • Choose dark mode, any viewport, 12 device presets, and retina scale.
  • For documents, select PDF paper size, margins, landscape mode, and page ranges.
  • Render HTML/CSS to an image, inject custom CSS or JavaScript, click an element before capture, hide selectors, and wait for a selector, fixed delay, or network idle.
  • Block ads, trackers, selected requests, or resource types; supply custom headers, cookies, user agent, and Authorization.
  • Set timezone and geolocation, use a transparent background, resize images, and cache with a TTL you choose.
  • Create signed links for public <img> tags, submit asynchronous jobs with signed webhooks, capture up to 100 URLs per call, query usage, and use the OpenAPI specification.
  • Parameter names used by other screenshot APIs also work, which can simplify migration.

An MCP server exposes take_screenshot, get_page_info, and capture_pdf to Claude, Cursor, and other MCP clients, so an AI agent can perform the capture without custom browser orchestration.

Plans

Plan Allowance Price
Free 1,000 shots per month No charge, no card
Starter 3,000 shots $5
Growth 15,000 shots $15
Pro 60,000 shots $39
Scale 250,000 shots $99
Business 1,000,000 shots $249

Yearly billing provides two months free, and every feature is included on every plan. Create a free ScreenshotNeo account to get 1,000 screenshots a month without adding a card.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

FAQ

Can the tracker monitor a binary file such as a PDF or image?

Yes, but skip HTML parsing and hash the raw response bytes directly. Keep the content type, file size, and retrieval timestamp so a failed or truncated download cannot be mistaken for a legitimate revision.

What encoding should be used for a stable text digest?

Encode the final normalized string as UTF-8 before calling SHA-256. If different producers normalize Unicode differently, apply one documented normalization policy consistently before encoding.

How can I migrate the tracker to another host?

Copy the script, dependency lock information, and state or history directory while preserving file ownership and permissions. Because the digest and normalized text are portable JSON values, no database migration is required for the minimal design.

Frequently Asked Questions

Can the tracker monitor a binary file such as a PDF or image?

Yes. Hash the raw response bytes instead of parsing HTML, and retain content type, size, and retrieval time so incomplete downloads are rejected.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What encoding should be used for a stable text digest?

Encode the normalized string as UTF-8 and apply one consistent Unicode-normalization policy before hashing.

How can I migrate the tracker to another host?

Copy the script, dependency information, and state or history directory while preserving permissions; the JSON records are portable.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.