Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

You can scrape search-result HTML with Python’s requests and BeautifulSoup, but only when the site’s robots.txt and terms allow automated access. Use a slow, capped session, stop on 403, 429, 503, CAPTCHA, or robot-check responses, and never treat Amazon’s crawler rules as permission to collect customer-facing search pages.

The complete pattern below checks robots.txt, fetches with timeouts and bounded retries, extracts only the fields you need, paginates with a hard limit, deduplicates products, and writes an auditable CSV. The selectors are deliberately generic: Amazon’s markup changes, and you must adapt them only for a target you are authorized to access.

Start with permission, not code

Before making a request, identify the exact host and path you intend to collect. Read its robots.txt and terms of service, and obtain any additional permission your project requires. If a path is disallowed or automated access is prohibited, stop and use an official API, a licensed export, or another permissioned source instead.

Amazon documents separate user-agent strings for Amazonbot, Amzn-SearchBot, and Amzn-User. Those rules describe Amazon’s own crawlers and how they follow robots.txt and page-level directives; they do not authorize a third-party scraper to fetch customer search pages.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Build your first version against a practice site or an explicitly allowed endpoint. The example code uses example.com and generic selectors so it does not imply that a selector is guaranteed to work on Amazon.

The scraping pipeline

  1. Confirm scope: write down the allowed host, paths, fields, maximum pages, request rate, and retention period.
  2. Check robots.txt: fetch it before the first search request and fail closed if you cannot verify the rule.
  3. Fetch politely: identify your program honestly, set a timeout, and use only a small number of retries for transient network failures.
  4. Detect blocking: treat 403, 429, 503, CAPTCHA pages, and robot-check HTML as a stop signal, not a challenge to evade.
  5. Parse narrowly: select only the product fields you need and preserve the original text.
  6. Paginate explicitly: cap pages, follow a verified next link or documented page parameter, and stop when no new products appear.
  7. Validate and record: keep URLs, timestamps, status codes, parser misses, and an HTML hash so an anomalous result can be investigated.

Install the Python dependencies

python -m pip install requests beautifulsoup4

Requests handles HTTP retrieval and BeautifulSoup parses the returned HTML. Pin and review dependency versions in a real project, and run the script from a service account or job identity whose contact information is valid.

A permissioned, low-rate Python example

This runnable example uses a three-page cap, a two-second delay, a 15-second request timeout, and generic article.product cards. These values are conservative starting points for a practice site, not guarantees for Amazon or any other host.

import csv
import hashlib
import time
from datetime import datetime, timezone
from urllib import robotparser
from urllib.parse import urljoin

import requests
from bs4 import BeautifulSoup

BASE_URL = 'https://example.com'
SEARCH_PATH = '/search'
QUERY = 'python book'
MAX_PAGES = 3
DELAY_SECONDS = 2
TIMEOUT_SECONDS = 15
USER_AGENT = 'ResearchExampleBot/1.0 (contact: [email protected])'

session = requests.Session()
session.headers.update({'User-Agent': USER_AGENT, 'Accept': 'text/html,application/xhtml+xml'})


def robots_allows(url):
    robots_url = urljoin(BASE_URL, '/robots.txt')
    try:
        robots_response = session.get(robots_url, timeout=TIMEOUT_SECONDS)
        robots_response.raise_for_status()
    except requests.RequestException as exc:
        print(f'Cannot verify robots.txt: {exc}')
        return False
    parser = robotparser.RobotFileParser()
    parser.set_url(robots_url)
    parser.parse(robots_response.text.splitlines())
    return parser.can_fetch(USER_AGENT, url)


def get_page(url, params):
    for attempt in range(1, 3):
        try:
            response = session.get(url, params=params, timeout=TIMEOUT_SECONDS)
        except requests.RequestException as exc:
            if attempt == 2:
                print(f'Network failure: {exc}')
                return None
            time.sleep(attempt * 2)
            continue

        if response.status_code in {403, 429, 503}:
            print(f'Stop signal: HTTP {response.status_code} at {response.url}')
            return None
        if response.status_code != 200:
            print(f'Stopping on HTTP {response.status_code} at {response.url}')
            return None

        lowered = response.text.lower()
        markers = ('captcha', 'robot check', 'automated access', 'verify you are human')
        if any(marker in lowered for marker in markers):
            print('Stop signal: challenge or robot-check HTML detected')
            return None
        return response
    return None


first_url = urljoin(BASE_URL, SEARCH_PATH)
if not robots_allows(first_url):
    raise SystemExit('Robots policy could not be verified for this URL')

seen = set()
rows = []

for page in range(1, MAX_PAGES + 1):
    response = get_page(first_url, {'k': QUERY, 'page': page})
    if response is None:
        break

    soup = BeautifulSoup(response.text, 'html.parser')
    page_new = 0
    for card in soup.select('article.product'):
        title_node = card.select_one('.title')
        link_node = card.select_one('a[href]')
        if not title_node or not link_node:
            continue

        product_url = urljoin(response.url, link_node['href'])
        key = card.get('data-asin') or product_url
        if key in seen:
            continue
        seen.add(key)

        price_node = card.select_one('.price')
        rating_node = card.select_one('.rating')
        reviews_node = card.select_one('.review-count')
        rows.append({
            'url': product_url,
            'asin_or_key': key,
            'title': title_node.get_text(' ', strip=True),
            'price_text': price_node.get_text(' ', strip=True) if price_node else '',
            'rating_text': rating_node.get_text(' ', strip=True) if rating_node else '',
            'review_count_text': reviews_node.get_text(' ', strip=True) if reviews_node else '',
            'retrieved_at_utc': datetime.now(timezone.utc).isoformat(),
            'source_html_sha256': hashlib.sha256(response.content).hexdigest(),
        })
        page_new += 1

    print(f'page={page} new_products={page_new} status={response.status_code}')
    if page_new == 0:
        break
    time.sleep(DELAY_SECONDS)

with open('products.csv', 'w', newline='', encoding='utf-8') as output:
    writer = csv.DictWriter(output, fieldnames=[
        'url', 'asin_or_key', 'title', 'price_text', 'rating_text',
        'review_count_text', 'retrieved_at_utc', 'source_html_sha256'
    ])
    writer.writeheader()
    writer.writerows(rows)

print(f'Wrote {len(rows)} unique products to products.csv')

Replace BASE_URL, SEARCH_PATH, parameters, and selectors only after confirming that the target permits your use. Do not copy these selectors into an Amazon job and assume they will remain valid.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What the example deliberately does not do

  • It does not rotate IP addresses, forge browser fingerprints, solve CAPTCHAs, or retry a block indefinitely.
  • It does not use an infinite-scroll click path. If products appear only after interaction, a plain HTTP fetch may never see them.
  • It does not assume price, currency, rating, or review fields exist in every locale.
  • It stores raw-response hashes rather than silently overwriting an unexplained change.

Checking robots.txt safely

The robots_allows function downloads the file with the same identifying session, parses its user-agent and allow/disallow directives, and fails closed if the file cannot be retrieved. A missing, malformed, or temporarily unavailable file is a reason to pause and verify the site’s policy manually, not a reason to continue guessing.

Keep a copy of the policy and the retrieval time with your run metadata. Policies can differ by host and path, and a rule for a crawler operated by Amazon is not a blanket license for your program.

Pagination that does not loop or miss products

Use a documented page parameter

If the permitted endpoint documents a page parameter, increment it from a known starting value and enforce MAX_PAGES. Stop when a page returns no new product keys. Deduplicating by ASIN when it is present, or by canonical product URL otherwise, prevents repeated cards from inflating the result.

Follow a verified next link

Some sites expose a next-page anchor instead of a numeric parameter. Parse that link, resolve it with urljoin, verify that it remains on an allowed host and path, and stop if it is missing, repeats a previously visited URL, or exceeds your page cap.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Do not confuse interaction with pagination

Clicks, infinite scroll, and client-side requests can create links that a crawler never sees. AWS describes this as a common crawler limitation. If the permitted workflow requires a browser, document the interaction and use browser automation only within the site’s rules; do not treat hidden links as permission to increase request volume.

Parsing and validating Amazon-like result cards

Extract the smallest useful record: product URL, title, price text, rating text, review-count text, a stable product key when available, and retrieval time. Preserve the displayed strings instead of coercing them immediately to numbers. Currency symbols, decimal separators, localized wording, missing ratings, and unavailable prices vary by locale and edition.

Log a parsing miss with the page URL and a reason. A sudden run with zero titles may mean a markup change, a consent page, a regional redirect, or a block page. Compare the stored HTML hash with a prior successful run before changing selectors.

For reproducibility, keep the request URL (including query parameters), HTTP status, selected response headers, UTC timestamp, parser version, and a retention-limited copy or hash of the source HTML. Remove credentials and any personal data from logs.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Why requests often receives 503 or a robot-check page

HTTP 403 or 429

These responses commonly indicate that the host is refusing the request or enforcing a rate limit. Stop the run, review robots.txt and terms, reduce scope and frequency, and ask for an approved API or export. Replaying the same request faster is not a fix.

HTTP 503

A 503 can be a temporary service failure or an automated-access block. Your bounded retry policy should make at most a small number of attempts for network failures, then stop. Record the response and contact the site owner or use a permissioned source.

CAPTCHA or “robot check” in a 200 response

HTTP status alone is not enough. Inspect the body for challenge markers and stop when they appear. Do not attempt to bypass the challenge; it is an explicit signal that the requested access is not being granted.

Empty results with status 200

Check whether the response is a consent page, regional redirect, JavaScript shell, or changed markup. Compare the final URL, content type, HTML hash, and parser-miss logs. If results are rendered after interaction, reassess whether browser automation is allowed and whether an official feed is more appropriate.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

TLS, JA3, or fingerprinting problems at scale

Industry guidance reports TLS/JA3 fingerprinting and 503 blocking as scale-related failure modes. Treat that as a signal to evaluate a compliant API or managed data service, not as an invitation to disguise your client.

Choosing an access method

Approach Permission and terms Reliability and content Maintenance and cost
Requests + BeautifulSoup You control the rate and must verify policy yourself. Fast for server-rendered HTML; misses interaction-generated content. Low infrastructure cost; selectors and locale differences require maintenance.
Browser automation Still requires permission; a browser does not override terms. Can execute permitted interactions and lazy loading, but is slower and heavier. Higher CPU, memory, latency, and operational complexity.
Official API or export Uses the provider’s documented access model. Usually the most stable schema, subject to its fields and quotas. Often the lowest maintenance; availability and pricing depend on the provider.
Managed scraping/data API Review the provider’s authorization, retention, and target coverage. May handle retries, localization, and rendering, but coverage and fidelity vary. Transfers infrastructure cost to a usage fee; compare total operating cost.

Choose by permission first, then by required fidelity, pagination behavior, geographic coverage, latency, request volume, maintenance effort, and total cost. A small, occasional, server-rendered extract may fit direct requests. A business-critical feed should be evaluated against an official API, export, or a managed service before you build a fragile selector layer.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Performance, reliability, and operating cost

  • Rate: one request every few seconds is a conservative prototype setting. Increase only when the owner’s policy explicitly permits it.
  • Timeouts: use finite connect/read timeouts and bounded retries so a stalled page cannot consume the whole job.
  • Concurrency: avoid parallel requests until you have written permission and measured the permitted load. More workers increase pressure and block risk.
  • Caching: cache responses when your terms allow it, and avoid re-fetching unchanged pages.
  • Scope: cap pages and fields; do not download assets you do not need.
  • Monitoring: alert on rising 403/429/503 rates, zero-new-product pages, parser misses, and locale-specific field loss.
  • Cost: direct requests have little software overhead but still consume bandwidth, storage, engineering time, and review effort. Official and managed options add provider-specific fees; compare them with the cost of maintaining selectors and handling blocks.

Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server, so it is useful when you need a visual record of a search page rather than structured product fields. A single GET returns PNG, JPEG, WebP, or PDF. It accepts the cookie/consent banner like a visitor, then removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each cleanup step can be disabled.

Only clean shots are billed. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and the response identifies the result with X-Page-Verdict and X-Billed headers. This does not grant permission to capture Amazon pages: confirm the target’s rules first.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

One-call capture with cURL

See the parameter reference in the ScreenshotNeo documentation.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://www.amazon.com/s?k=python+book -o shot.webp

The same call in Python

import requests

r = requests.get(
    'https://api.screenshotneo.com/v1/shot',
    params={'access_key': 'YOUR_API_KEY', 'url': 'https://www.amazon.com/s?k=python+book'},
    timeout=90,
)
r.raise_for_status()
open('shot.webp', 'wb').write(r.content)

Node.js

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://www.amazon.com/s?k=python+book' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`HTTP ${res.status}`);
const fs = await import('node:fs/promises');
await fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer()));

ScreenshotNeo also supports full-page capture with lazy images loaded, CSS-selector element capture, dark mode, 12 device presets plus custom viewports, retina scale, PDF paper size/margins/landscape/page ranges, HTML/CSS-to-image, custom JavaScript and CSS, clicks, waits for selectors or network idle, ad/tracker/request/resource blocking, custom headers/cookies/user agents and Authorization, timezone and geolocation, transparent backgrounds, resizing, configurable-TTL caching, signed public-image links, asynchronous jobs with signed webhooks, bulk capture for up to 100 URLs per call, a usage API, an OpenAPI specification, and familiar parameter names for easier migration.

An MCP server exposes take_screenshot, get_page_info, and capture_pdf to Claude, Cursor, and other MCP clients. Plans include 1,000 shots per month free with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account to try it within the target’s permission boundaries.

FAQ

Can I reuse the CSV as a price database?

Only if your agreement with the site and the applicable law permit that reuse; the script’s output format does not change those obligations.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What if two locales return different product sets?

Run each permitted locale as a separate dataset and retain the locale, currency text, final URL, and retrieval timestamp with every row.

Is a screenshot a substitute for structured extraction?

No. A screenshot preserves visual evidence; it does not reliably provide normalized titles, prices, ratings, or ASINs. Use an authorized API or parser when downstream systems need fields.

Frequently Asked Questions

Can I reuse the CSV as a price database?

Only if your agreement with the site and applicable law permit that reuse; the output format does not change those obligations.

What if two locales return different product sets?

Run each permitted locale as a separate dataset and retain locale, currency text, final URL, and retrieval timestamp with every row.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Is a screenshot a substitute for structured extraction?

No. It preserves visual evidence, not normalized product fields; use an authorized API or parser for downstream data.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.