PC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minuteThe reliable way to scrape ecommerce data is to start with the least complex source the merchant permits: an authorized API, product feed, or export. Use direct HTML parsing only for pages you are allowed to access, and use browser automation such as Playwright when the required fields appear only after normal JavaScript rendering. Store every observation with its variant, seller, currency, availability, and UTC timestamp. A block, challenge, or explicit prohibition is a stop signal—not a puzzle to bypass.
Define exactly what you need before you collect anything
“Scrape the catalog” is not a specification. Write down the fields, refresh interval, markets, and permitted use first. A price-monitoring job may need fewer pages than a complete catalog import.
Separate catalog attributes from offer attributes
Titles, descriptions, model numbers, options, and images are comparatively stable catalog attributes. Price, stock, seller, promotion, delivery promise, and currency are volatile offer attributes. Treating an offer as a permanent product fact creates misleading comparisons.
- Identity: source URL, product ID or SKU when legitimately exposed, and a stable model or handle.
- Variant: size, color, capacity, pack count, region, or any other option that changes the offer.
- Seller and offer: merchant, marketplace seller, condition, promotion text, and shipping terms where relevant.
- Value: numeric price, displayed price text, currency, and whether the value is a sale, subscription, unit, or total.
- Availability: the exact stock or availability text and a normalized status such as in-stock, out-of-stock, preorder, or unknown.
- Observation: UTC time, retrieval outcome, response status, and parser version.
Keep the original displayed text alongside normalized fields. It lets you audit currency symbols, sale labels, “from” prices, and locale-specific number formats later.
#1 Best Overall
Choose an authorized source before writing a crawler
Merchant feed, export, or API
Use a product feed, export, or API when the merchant offers one and your collection and reuse fit the permission granted. Shopify’s catalog documentation describes discoverable product data such as titles, descriptions, options, images, prices, and availability, with data described as continuously updated. That does not grant unrestricted collection: Shopify’s API terms restrict scraping and systematic automated collection, prohibit bypassing API restrictions, and limit collection to the permissions and purposes granted. Those terms are Shopify-specific; read the target provider’s own documentation and contract.
Ask the merchant or platform for the intended route when you need historical data, bulk access, or fields not exposed publicly. An official feed usually gives you clearer identifiers and fewer layout changes than page parsing.
Direct page retrieval
For a permitted public page whose needed values are in the returned HTML, an HTTP client and parser are usually simpler than a browser. Begin with a narrow URL list, identify the selectors, parse structured data when it is present, and save the response outcome. Do not assume that a page visible in a browser is automatically unrestricted for automated collection.
Browser rendering with Playwright
Use Playwright when the product details appear only after ordinary, permitted browser rendering or interaction. Browser startup, JavaScript execution, cookie state, and selector maintenance add cost and failure modes, so use it only when a simpler authorized request cannot supply the fields. Playwright renders a page; it does not provide authorization.
The Tool Desk
Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Build a bounded, incremental collection plan
Discover a finite URL set
Prefer a merchant-provided feed, sitemap, category pages, or an explicitly documented catalog endpoint. Deduplicate canonical product URLs and cap the number of pages per run. Avoid generating every combination of filters, sorting orders, query terms, and pagination unless that space is required and permitted.
Refresh changes instead of recrawling everything
Keep the last successful observation and schedule refreshes according to the decision you are supporting. A daily assortment report, an hourly stock alert, and a one-time market snapshot have different needs. Faster polling is not automatically better: Salesforce notes that request cost varies by path, and uncached pages, combined filters, and pages that fan out into internal calls can cost much more than a typical page.
- Use a queue with a maximum concurrency and a delay between requests.
- Back off after 429 responses, timeouts, challenges, or rising latency.
- Cache responses only when the site’s rules allow it, and avoid repeatedly downloading unchanged pages.
- Record failures separately from “out of stock”; a timeout is not an inventory fact.
- Stop a job when the owner signals that automated access is not allowed.
Why “polite” traffic can still be expensive
A low request rate does not guarantee a low storefront load. A search URL with many filters can trigger expensive backend work, and aggregate traffic to that pattern can remain harmful even when every individual IP is below a per-client threshold. Salesforce recommends understanding page costs, load testing, regulating traffic, and configuring rate limits and firewall rules. The collector’s practical response is to narrow requests and respect those controls, not to evade them.
robots.txt is guidance, not permission or security
Read robots.txt as the publisher’s crawler instructions. Google explains that crawlers must honor the rules voluntarily, syntax can be interpreted differently, and a disallowed URL can still be indexed if other sites link to it. RFC 9309 formalizes the Robots Exclusion Protocol, but it is not an authentication system or a security boundary.
Therefore, a robots check is one input to a collection decision, not legal clearance. A URL can be publicly discoverable while automated access is contractually restricted. Conversely, a missing rule is not an invitation to overload a site. Check the terms, API license, authentication requirements, and any explicit owner instructions as well.
“robots.txt is advisory: crawlers must honor it voluntarily, and it has no enforcement mechanism.” — Salesforce Developers, Bot Mitigation Best Practices for Flash Sales.
A practical Python baseline for permitted static pages
The following example intentionally uses a small, explicit URL list and records outcomes. Replace the selectors and URLs only for a site whose access and reuse you have permission to conduct. It does not attempt to defeat a challenge or retry indefinitely.
import csv
import time
from datetime import datetime, timezone
from decimal import Decimal, InvalidOperation
from urllib.parse import urlparse
import requests
from bs4 import BeautifulSoup
URLS = [
"https://example.com/products/widget",
]
HEADERS = {
"User-Agent": "CatalogObserver/1.0 (contact: [email protected])",
"Accept": "text/html,application/xhtml+xml",
}
def parse_money(text):
if not text:
return None
cleaned = text.replace(",", "").strip()
digits = "".join(ch for ch in cleaned if ch.isdigit() or ch == ".")
if not digits:
return None
try:
return str(Decimal(digits))
except InvalidOperation:
return None
def fetch(url):
observed_at = datetime.now(timezone.utc).isoformat()
try:
response = requests.get(url, headers=HEADERS, timeout=20)
except requests.RequestException as exc:
return {"url": url, "observed_at": observed_at,
"outcome": "request_error", "error": str(exc)}
if response.status_code in (401, 403, 429):
return {"url": url, "observed_at": observed_at,
"outcome": f"access_{response.status_code}", "error": "stop_and_review"}
if response.status_code >= 400:
return {"url": url, "observed_at": observed_at,
"outcome": f"http_{response.status_code}"}
soup = BeautifulSoup(response.text, "html.parser")
title_node = soup.select_one("h1")
price_node = soup.select_one("[data-price]")
availability_node = soup.select_one("[data-availability]")
title = title_node.get_text(" ", strip=True) if title_node else ""
displayed_price = price_node.get_text(" ", strip=True) if price_node else ""
availability = availability_node.get_text(" ", strip=True) if availability_node else ""
return {
"url": url,
"product_id": "", # fill only when legitimately available
"variant": "",
"seller": "",
"currency": "", # derive from the page; do not guess
"title": title,
"displayed_price": displayed_price,
"price": parse_money(displayed_price),
"availability": availability,
"observed_at": observed_at,
"outcome": "ok",
"parser_version": "1.0",
}
rows = []
for url in URLS:
rows.append(fetch(url))
time.sleep(2) # choose a delay appropriate to the site’s instructions
with open("observations.csv", "w", newline="", encoding="utf-8") as output:
fields = sorted({key for row in rows for key in row})
writer = csv.DictWriter(output, fieldnames=fields)
writer.writeheader()
writer.writerows(rows)
Install the dependencies with python -m pip install requests beautifulsoup4. If the values are absent from the response HTML, do not “fix” the script by hammering the site; move to an authorized feed or a permitted browser-rendering workflow.
Rank #3
When to use browser automation
Choose browser rendering only after confirming that normal page access is permitted and static retrieval cannot provide the fields. Keep the browser workflow bounded: one context, a small queue, explicit waits for the selector you need, and a hard timeout. Selectors should target semantic attributes or stable test IDs where the site supplies them, not long chains of generated classes.
Handle these cases distinctly:
- Consent dialog: follow the site’s normal visitor flow when permitted, and document what state your collector uses.
- Login or paywall: use only credentials and automation expressly authorized for the account and purpose.
- Bot challenge or CAPTCHA: stop. Do not rotate identities or automate a bypass.
- Missing lazy-loaded images: wait for the required element or use the merchant’s feed instead of scrolling every page.
- Selector failure: record “parser_error,” preserve the page metadata you are allowed to retain, and update the parser deliberately.
Compare collection methods by the decision you must make
| Method | Best fit | Strengths | Costs and risks |
|---|---|---|---|
| Merchant feed or export | Catalog import, approved analytics, scheduled updates | Clear permission, stable identifiers, structured fields | Coverage and freshness depend on what the merchant supplies |
| Authorized API | Ongoing product or offer queries | Documented fields, predictable request semantics | Terms, scopes, quotas, and reuse limits apply; never bypass restrictions |
| Static HTML parsing | Small, permitted public-page collection | Low implementation overhead and no browser startup | Breaks when content is client-rendered or markup changes |
| Playwright rendering | Permitted pages whose data appears after JavaScript | Can reproduce ordinary browser rendering and interaction | Slower, more resource-intensive, selector-sensitive, and not authorization |
Evaluate each option on permission and terms, field and variant coverage, freshness, geography and seller coverage, implementation effort, request load, markup stability, and usage limits. There is no meaningful universal winner without those requirements.
Legal and policy boundaries depend on context
There is no blanket rule that ecommerce scraping is always legal or always illegal. In hiQ Labs, Inc. v. LinkedIn Corporation, the Ninth Circuit considered whether collecting publicly viewable profile information was access “without authorization” under the CFAA in that dispute. The decision is not a universal license to collect retailer data and does not resolve contracts, copyright, database rights, privacy, or laws outside that context.
The U.S. Department of Justice’s CFAA Justice Manual says a prosecution may not be based solely on violating a contractual access restriction or terms of service for a generally available public website. That is prosecution guidance about a particular statute, not a ruling that removes civil claims or other obligations.
Do these 3 things before closing this tab:
1Clear out junk files and repair common Windows errors2Scan for outdated or missing drivers - takes under a minute3Repair Windows errors before they cause bigger problems- Read the target’s terms, API license, and documented access policy.
- Confirm whether authentication, personal data, restricted content, or non-public endpoints are involved.
- Define how you will use, retain, combine, and publish the data.
- Consider copyright, database, privacy, consumer-protection, and contractual issues in every relevant geography.
- Obtain jurisdiction-specific legal advice for commercial, large-scale, or personal-data projects.
Cloudflare’s sample terms illustrate language aimed at AI-related scraping, but the sample is expressly not legal advice or a guarantee of an outcome. Treat it as an example of one provider’s approach, not a general template or statement of law.
Troubleshooting common failures
HTTP 403, 401, or a challenge page
Cause: access control, a firewall rule, authentication, or a bot-management decision. Fix: stop retries, check the documented API or feed, contact the owner, and obtain permission. Do not add proxy rotation, CAPTCHA solving, or identity spoofing to defeat the control.
HTTP 429 or rising latency
Cause: rate limits or expensive request patterns. Fix: reduce concurrency, increase delay, remove combinatorial filters, cache where allowed, and switch to incremental updates.
The HTML has no price or stock
Cause: client-side rendering, a locale-dependent response, or a selector that no longer matches. Fix: compare the permitted API/feed first; otherwise use a bounded Playwright flow and wait for a specific element. Record parser failures separately from unavailable stock.
Free tools Windows power users keep installed
One-click scans. No signup required.
Prices do not compare correctly
Cause: different variants, sellers, currencies, units, or promotions. Fix: make those dimensions explicit in the record, preserve displayed text, and only normalize currencies or units with a documented conversion policy.
A catalog run becomes unexpectedly large
Cause: duplicate URLs, faceted navigation, infinite pagination, or repeated full crawls. Fix: bound discovery, canonicalize URLs, deduplicate product identifiers, cap pages, and schedule deltas instead of rebuilding the entire catalog.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Or skip the browser setup
ScreenshotNeo provides a website screenshot API and MCP server for developers. One GET request can return a PNG, JPEG, WebP, or PDF. Before capture it accepts the cookie or consent banner like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be turned off. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and each response identifies the result with X-Page-Verdict and X-Billed headers.
For a visual record of a permitted product page, call the API directly:
Recommended Free Tools
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python:
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
open("shot.webp", "wb").write(r.content)
Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
const fs = await import('node:fs/promises');
await fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer()));
See the ScreenshotNeo documentation for request parameters. The service supports full-page capture with lazy images loaded, CSS-selector element capture, dark mode, 12 device presets or any viewport, retina scale, PDF paper size and page ranges, HTML/CSS-to-image, custom CSS and JavaScript, clicks, selector waits, delays, network-idle waits, ad/tracker/request blocking, headers, cookies, user agents, Authorization, timezone and geolocation, transparent backgrounds, resizing, chosen cache TTLs, signed image links, asynchronous jobs with signed webhooks, bulk capture of up to 100 URLs per call, a usage API, and an OpenAPI specification. Parameter names used by other screenshot APIs also work, which can simplify migration.
Best Value
An MCP server exposes take_screenshot, get_page_info, and capture_pdf to Claude, Cursor, and other MCP clients. Every feature is available on every plan:
| Plan | Included shots | Price |
|---|---|---|
| Free | 1,000 per month | No card required |
| Starter | 3,000 | $5 |
| Growth | 15,000 | $15 |
| Pro | 60,000 | $39 |
| Scale | 250,000 | $99 |
| Business | 1,000,000 | $249 |
Yearly billing gives two months free. Cookie banners, popups, and chat widgets are removed before the shot; bot checks, blank pages, and failed loads are never billed; and AI agents can take screenshots through the MCP server. Start with 1,000 free screenshots a month—no card required.
Final checklist before you schedule a job
- Have you selected a feed, API, export, static page, or browser route that the owner permits?
- Did you read the applicable terms, API scope, robots instructions, and access controls?
- Are product, variant, seller, currency, availability, and observation time separate fields?
- Is discovery bounded and deduplicated, with incremental refresh instead of needless full crawls?
- Do retries stop on challenges, prohibitions, repeated 403/401 responses, and sustained 429s?
- Can you explain every normalized price and retain the original displayed value?
- Are parser versions, failures, and timestamps retained so results can be audited?
Frequently Asked Questions
Should I convert currencies during collection?
Store the page’s displayed currency and value first. Apply a separate, dated conversion policy only when your comparison requires it, and retain the original amount so the result remains auditable.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
How should I handle a product sold by several marketplace sellers?
Create one observation per seller or offer, keeping seller identity, condition, shipping terms, variant, currency, and timestamp explicit. Do not collapse unlike offers into one product price.
Is a screenshot a substitute for structured product data?
No. A screenshot preserves visual evidence, while a feed, API, or parsed record supplies fields for filtering and analysis. Use screenshots as an audit or presentation artifact alongside structured observations.
The Bottom Line
Use the most direct authorized source available, collect only the fields and pages your decision requires, and treat robots instructions, rate limits, challenges, and terms as boundaries rather than obstacles. Accurate timestamps and explicit offer context matter as much as the parser.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

