Quick wins for a faster PC:
Repair Windows errors before they cause bigger problemsFix Now →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →To scrape an e-commerce category page reliably, first check the site’s access rules, then fetch its ordinary HTML and extract product cards with CSS selectors. Follow real pagination links or a permitted data endpoint to reach additional products; use a browser such as Playwright only when the data genuinely requires JavaScript or user interaction. Normalize and deduplicate the results, and validate them before treating the crawl as a complete catalog.
Plan the crawl before writing a scraper
A category page is usually a view of part of a catalog, not a complete product feed. Decide what you need to collect, how far the crawl may go, and how often it should run before sending requests.
Define records and boundaries
For each product, useful fields may include its URL, title, SKU or other exposed identifier, price, currency, availability, image URL, category path, and crawl timestamp. Some pages will not expose every field. Record missing values rather than silently substituting guesses. Set a category list, maximum page count, refresh cadence, and request concurrency that fit your need.
Decide how product variants should be represented. If a page exposes distinct sizes or colors with separate identifiers or prices, preserve those identifiers rather than collapsing them into a single product row. Keep the original text or response metadata alongside normalized values so a later parser change can be diagnosed.
Free tools Windows power users keep installed
One-click scans. No signup required.
#1 Best Overall
Check permission and publisher controls
Before collecting data, read the site’s robots.txt, terms of service, and any applicable restrictions. Treat these as separate checks: robots rules are not a permission grant, and they do not settle questions about privacy, copyright, database rights, authentication, or contractual terms. Do not attempt to defeat login barriers, access controls, or anti-bot measures. Use a descriptive user agent, conservative request rates, timeouts, retries with backoff, caching, and a hard page cap.
Google explains that robots.txt manages crawler traffic and is not a way to hide URLs from search results. See its Robots.txt Introduction and Guide. A disallow rule should not be treated as evidence that a URL is secret or as a substitute for checking the site’s terms and applicable rules.
Choose an approach that matches the page
| What the site exposes | Approach | Trade-off |
|---|---|---|
| Product cards and next-page links are present in the initial HTML | HTTP client plus Scrapy selectors, lxml, or BeautifulSoup | Fast and inexpensive, but will not see important content rendered only in the browser |
| Many categories need scheduled refreshes, retries, and job tracking | Scrapy spider with item pipelines and persistent job state | Offers crawl controls for larger jobs but requires framework setup |
| Products or prices appear only after JavaScript or an interaction | First investigate whether a permitted JSON endpoint provides the data; otherwise use Playwright or another browser renderer | More faithful to browser-visible content, but slower and more resource-intensive |
| A sitemap or merchant feed lists the catalog | Discover product URLs there, then make targeted product requests | Can make discovery more efficient, though feed fields may differ from page fields |
Start with ordinary HTML rather than assuming a browser is necessary. Scrapy describes spiders as components that generate requests, parse responses, and return structured items; its documentation covers both spiders and selectors. For discovery, inspect navigation links, sitemaps, and feeds when the visible category links do not expose all relevant products. Google’s e-commerce structure guidance recommends direct links among menus, categories, subcategories, and products, and discusses sitemaps or feeds where links are incomplete.
Build a small HTML scraper with Python
The example below uses Requests and BeautifulSoup to fetch a category page, parse product cards, and follow a conventional next-page link. Install the dependencies with python -m pip install requests beautifulsoup4. Save the script as scrape_category.py, replace the starting URL and selectors with values observed on a site you are permitted to crawl, then run python scrape_category.py.
The CSS selectors in this example are illustrative: stores use different markup. Inspect the page source or browser developer tools to identify the repeated card container, product link, title, price, and next link. If a field is absent, the script leaves it empty rather than fabricating it.
import csv
import time
from urllib.parse import urljoin, urldefrag, urlparse
import requests
from bs4 import BeautifulSoup
START_URL = "https://example.com/category/shoes"
MAX_PAGES = 20
DELAY_SECONDS = 1.0
# Adapt these selectors to the store's ordinary HTML.
CARD_SELECTOR = ".product-card"
TITLE_SELECTOR = ".product-title"
PRICE_SELECTOR = ".price"
PRODUCT_LINK_SELECTOR = "a.product-link"
NEXT_SELECTOR = "a.next"
session = requests.Session()
session.headers.update({"User-Agent": "CategoryResearchBot/1.0 (contact: [email protected])"})
def clean_url(base, href):
if not href:
return ""
absolute = urljoin(base, href)
absolute, _fragment = urldefrag(absolute)
return absolute
def text_or_empty(parent, selector):
node = parent.select_one(selector)
return node.get_text(" ", strip=True) if node else ""
def scrape():
url = START_URL
start_host = urlparse(START_URL).netloc
seen_pages = set()
seen_products = set()
rows = []
for _page_number in range(MAX_PAGES):
url = clean_url(url, url)
if not url or url in seen_pages:
break
if urlparse(url).netloc != start_host:
raise ValueError("Next-page link left the configured site boundary")
seen_pages.add(url)
response = session.get(url, timeout=(5, 30))
response.raise_for_status()
soup = BeautifulSoup(response.text, "html.parser")
for card in soup.select(CARD_SELECTOR):
link = card.select_one(PRODUCT_LINK_SELECTOR)
product_url = clean_url(url, link.get("href")) if link else ""
if product_url and product_url in seen_products:
continue
if product_url:
seen_products.add(product_url)
rows.append({
"product_url": product_url,
"title": text_or_empty(card, TITLE_SELECTOR),
"price_text": text_or_empty(card, PRICE_SELECTOR),
"category_page": url,
"crawled_at_utc": time.strftime("%Y-%m-%dT%H:%M:%SZ", time.gmtime()),
})
next_link = soup.select_one(NEXT_SELECTOR)
next_url = clean_url(url, next_link.get("href")) if next_link else ""
if not next_url:
break
time.sleep(DELAY_SECONDS)
url = next_url
with open("products.csv", "w", newline="", encoding="utf-8") as handle:
fields = ["product_url", "title", "price_text", "category_page", "crawled_at_utc"]
writer = csv.DictWriter(handle, fieldnames=fields)
writer.writeheader()
writer.writerows(rows)
print(f"Wrote {len(rows)} product rows from {len(seen_pages)} pages to products.csv")
if __name__ == "__main__":
scrape()
What to change before using it
- Set
START_URLto the category URL and replace the example user agent with one that identifies your crawler appropriately. - Inspect a representative page and adapt the card, title, price, product-link, and next-link selectors. A selector that matches no cards can produce a valid-looking but empty CSV.
- Choose
MAX_PAGESas a safety limit for your task. The script also stops when there is no next link or the same page URL recurs. - Increase the delay or reduce crawl scope if the publisher’s rate limits or terms require it. This simple example does not implement automatic retries or persistent job state; for a larger scheduled crawl, use a framework such as Scrapy and configure its crawl controls appropriately.
The example stores price as displayed text. For useful comparisons, parse the amount and currency separately with rules suited to the site’s locale. Do not assume a comma or period always has the same decimal meaning across markets.
Reach every product without clicking blindly
After parsing the first response, inspect how the category exposes further products. The simplest case is a real next-page link with a distinct URL. Follow it until the link disappears, product identifiers stop changing, or your configured page limit is reached. Keep a set of visited page URLs and product identifiers so duplicate links or repeated listings do not inflate the output.
Google recommends unique URLs for paginated sequences and warns that URL fragments are not reliable page numbers. See its pagination and incremental page loading guidance. Do not assume changing a fragment such as #page=2 fetches a distinct page; verify that the server actually returns different content.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Rank #3
Load-more buttons and infinite scroll
If a button or scroll event reveals more products, use browser developer tools to inspect the network requests made as the page loads more items. A stable JSON request may be a simpler and lighter source than rendering every browser interaction, but use it only if access is permitted and the endpoint is not an access-control bypass. Confirm that the response corresponds to the same category and preserves pagination state.
If there is no stable permitted endpoint and essential content appears only after JavaScript or a user action, a browser renderer such as Playwright is a reasonable fallback. This costs more time and resources than parsing initial HTML. Google notes that its crawlers generally do not click buttons or trigger JavaScript functions that require user actions to update a page; that is a search-crawling observation, not a guarantee about how a particular store’s own browser interface works.
Normalize, deduplicate, and validate the data
Scraping is not complete when a CSV has been written. Check whether the output represents distinct products and whether the values have a consistent meaning.
- Canonicalize product URLs consistently, while preserving query parameters if they identify a real variant.
- Prefer a stable exposed SKU or product ID for deduplication; otherwise use a stable product URL. Keep variant identifiers where separate variants matter.
- Normalize availability labels into a documented set of values, but retain the original label for auditability.
- Store parsed numeric prices with an explicit currency field, and retain raw displayed price text for locale or markup changes.
- Track missing-field rates, duplicate rates, page counts, HTTP status distributions, and unexpected template changes.
- Keep representative page fixtures and run parser regression checks when selectors or site markup change.
A category can omit products that exist elsewhere in the catalog, and a crawl can miss items because of pagination, merchandising changes, or a parser mismatch. If coverage matters, compare category results with permitted sitemap or feed discovery, and document the crawl boundary and timestamp with the dataset.
The Tool Desk
Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server made by Yorker Media. A screenshot is visual evidence, not a substitute for structured product extraction: use it to inspect how a category rendered or to keep a visual check alongside your scraper. One GET request can capture a page as PNG, JPEG, WebP, or PDF. Before capture, it can accept cookie or consent banners as a visitor and remove more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be turned off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and responses include X-Page-Verdict and X-Billed headers.
For example, capture a category page as WebP with cURL. Create an API key first, then replace the target URL with the category you are allowed to inspect. See the ScreenshotNeo API documentation for request options and response details.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com/category/shoes -o shot.webp
ScreenshotNeo also provides an MCP server for AI agents, including Claude, Cursor, and other MCP clients, with take_screenshot, get_page_info, and capture_pdf tools. Plans include 1,000 shots per month free with no card, Starter at $5 for 3,000, Growth at $15 for 15,000, Pro at $39 for 60,000, Scale at $99 for 250,000, and Business at $249 for 1,000,000; yearly billing gives two months free. Every feature is on every plan. Learn more at ScreenshotNeo, or sign up for 1,000 free screenshots a month with no card.
Troubleshoot common failures
The script returns zero products
Most often the card selector does not match the live markup, the page returned a consent or error screen, or products are inserted only after JavaScript runs. Inspect the response HTML and confirm the page status before changing selectors. If the initial response truly lacks the product data, inspect permitted network responses or use a browser renderer rather than repeatedly tuning selectors against an empty page.
PC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchOnly the first batch appears
Check whether the category has a real next-page anchor, a load-more request, or an infinite-scroll request. Confirm that each next URL returns new products and is not merely a tracking or fragment variation. Set a page cap and stop on repeated page URLs or product IDs.
Best Value
Prices or availability look wrong
Keep raw strings and inspect locale, variant selection, and page state. A displayed promotional price may coexist with a crossed-out list price, and stock wording may vary. Define which price and availability field your use case requires; do not treat one selector as semantically correct across every store.
Requests time out or return errors
Use bounded timeouts, lower concurrency, cache responses where appropriate, and retry transient failures with backoff rather than in a tight loop. Check the site’s rate limits and whether the crawl should pause. Do not respond to bot checks or authentication barriers by trying to circumvent them.
FAQs
Does robots.txt give permission to republish product data?
No. Robots.txt communicates crawler preferences; it does not by itself grant permission to collect, store, or republish data. Check the site’s terms and the rules applicable to your intended use.
Can I use a sitemap instead of crawling category pagination?
A sitemap or merchant feed can help discover product URLs when category browsing is incomplete. It may not contain the same fields as product pages, so verify whether it supplies the attributes your dataset needs.
Frequently Asked Questions
Does robots.txt give permission to republish product data?
No. Robots.txt communicates crawler preferences; it does not by itself grant permission to collect, store, or republish data. Check the site’s terms and the rules applicable to your intended use.
Can I use a sitemap instead of crawling category pagination?
A sitemap or merchant feed can help discover product URLs when category browsing is incomplete. It may not contain the same fields as product pages, so verify whether it supplies the attributes your dataset needs.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




