Recommended Free Tools
To scrape multiple pages reliably, define the records you need, fetch each page, extract and normalize the same fields, then save one validated record per item. For pagination, follow the page’s next link until it disappears. Use Requests and Beautiful Soup for small server-rendered jobs, Scrapy for larger or branching crawls, and Playwright when the data requires a browser to render.
Plan the crawl before fetching pages
Start by deciding what one output record represents: a product, article, listing, or another item. Write down the fields and their expected types before writing selectors. A product record, for example, might contain name, price, url, and source_page.
Choose a stable key—often a source ID or canonical URL—for deduplication. Decide how missing values should be represented, and normalize whitespace, dates, prices, and URLs consistently. If later auditing matters, retain the raw response or add provenance such as the source URL and crawl time.
Before crawling, review the site’s robots.txt, terms, authentication boundaries, privacy obligations, and copyright constraints. A crawler can observe robots rules, but compliance with a site’s rules and applicable law is a separate question. Scrapy provides robots.txt support and configurable request controls in its settings.
#1 Best Overall
Choose a method that matches the pages
| Method | Best fit | Trade-off |
|---|---|---|
| Requests + Beautiful Soup | A small, straightforward set of server-rendered pages | Simple and explicit, but you must build the loop, retries, and output handling. Beautiful Soup is forgiving of imperfect markup; Scrapy notes its parser approach is slower than lxml-backed selectors. Scrapy selector guide |
| Scrapy | Many pages, pagination, branching links, or repeatable crawls | Provides scheduling, duplicate filtering, asynchronous processing, pipelines, and feed exports, but requires a project and spider structure. Scrapy overview |
| Playwright or browser-rendering integration | Pages where content appears only after browser JavaScript runs | Uses more browser machinery than a direct HTTP request. Prefer an underlying JSON or API request if one is available. Playwright network events |
A useful diagnostic is to inspect the initial HTML response. If the desired text and links are already present, a browser may be unnecessary. If not, check the browser’s network requests for a data endpoint before resorting to rendering. When a browser is necessary, distinguish a completed HTTP exchange from a successful page: Playwright documents that HTTP 404 or 503 responses still count as successful responses from the HTTP standpoint. Playwright request documentation
Scrape a small set with Requests and Beautiful Soup
This example visits a known set of listing pages, extracts one record per product card, validates required fields, and writes JSON Lines. Replace the example domain and selectors with those confirmed from the target site’s HTML. Install the dependencies with python -m pip install requests beautifulsoup4.
import json
import time
from urllib.parse import urljoin
import requests
from bs4 import BeautifulSoup
START_URLS = [
"https://example.com/catalog?page=1",
"https://example.com/catalog?page=2",
]
OUTPUT = "products.jsonl"
session = requests.Session()
session.headers.update({"User-Agent": "ExampleResearchBot/1.0 (contact: [email protected])"})
def clean(text):
return " ".join((text or "").split())
def scrape_page(url):
response = session.get(url, timeout=(5, 30))
response.raise_for_status()
soup = BeautifulSoup(response.text, "html.parser")
records = []
for card in soup.select("article.product"):
link = card.select_one("a[href]")
name = card.select_one("h2")
price = card.select_one(".price")
record = {
"name": clean(name.get_text()) if name else "",
"price": clean(price.get_text()) if price else "",
"url": urljoin(response.url, link["href"]) if link else "",
"source_page": response.url,
}
if record["name"] and record["url"]:
records.append(record)
return records
seen = set()
with open(OUTPUT, "w", encoding="utf-8") as out:
for page_url in START_URLS:
for record in scrape_page(page_url):
key = record["url"]
if key not in seen:
seen.add(key)
out.write(json.dumps(record, ensure_ascii=False) + "n")
time.sleep(1) # Set a respectful interval for this site.
print(f"Wrote {len(seen)} unique records to {OUTPUT}")
The request timeout is a connect/read timeout tuple, not a guarantee that every server or operation completes within one combined wall-clock limit. raise_for_status() makes HTTP error statuses visible instead of silently parsing an error page as data. For larger jobs, add bounded retries for transient failures, structured logs, and checkpoints so a stopped run need not start from zero.
When the page list is not known in advance
For pagination, parse the next link from each response, resolve relative links against the response URL, and stop when there is no next link. Track visited URLs: sites sometimes link back to a page or expose duplicate links. A reliable crawl should also have a maximum-page or maximum-item guard in case the site’s pagination is malformed.
Use Scrapy for pagination and larger crawls
Scrapy is designed around spiders: classes that define where requests begin and how responses produce items or further requests. The official tutorial demonstrates yielding extracted items and following a relative next-page link with response.follow. Scrapy tutorial
Install Scrapy with python -m pip install scrapy, then create a project using scrapy startproject catalog_crawl. Save this spider as catalog_crawl/spiders/catalog.py:
import scrapy
class CatalogSpider(scrapy.Spider):
name = "catalog"
allowed_domains = ["example.com"]
start_urls = ["https://example.com/catalog"]
def parse(self, response):
for card in response.css("article.product"):
name = card.css("h2::text").get(default="").strip()
href = card.css("a::attr(href)").get()
price = card.css(".price::text").get(default="").strip()
if name and href:
yield {
"name": name,
"price": price,
"url": response.urljoin(href),
"source_page": response.url,
}
next_href = response.css("a.next::attr(href)").get()
if next_href:
yield response.follow(next_href, callback=self.parse)
Run it from the project directory and export records as JSON Lines with scrapy crawl catalog -O products.jsonl. Scrapy supports JSON, CSV, and XML feed exports, as well as item pipelines for validation or transformation. Feed exports
Scrapy schedules yielded requests asynchronously and filters duplicate URLs by default. Its concurrency, download-delay, and AutoThrottle controls let you balance throughput with load on the target site; select conservative per-domain values rather than assuming maximum concurrency is appropriate. AutoThrottle and settings
Do these 3 things before closing this tab:
1Fix the driver behind crashes, sound loss and screen glitches2Clear out junk files and repair common Windows errors3Scan for outdated or missing drivers - takes under a minuteRank #3
Handle JavaScript-rendered pages carefully
First determine whether the data is available from an ordinary request. Browser developer tools can show network requests made when the page loads or pagination is clicked; if a permitted, stable JSON endpoint supplies the same data, requesting it directly is usually simpler than rendering the entire page.
If browser execution is needed, Playwright can listen for request, response, completion, and failure events. A request finishing does not mean the response was successful: inspect its status, and separately wait for the page element or application state that signals the data is ready. Playwright network events The exact wait condition depends on the site’s behavior; a fixed sleep can be either wasteful or too short.
For a crawler that otherwise uses Scrapy, browser-rendering integrations can bridge browser pages into the crawl, but their setup and compatibility depend on the chosen integration and versions. Scrapy’s ecosystem documentation discusses JavaScript rendering options. Scrapy dynamic content
Validate, pace, and make runs recoverable
- Test representative pages: Save a few responses and verify selectors on empty, typical, and edge-case pages before scaling up.
- Validate the schema: Check required fields and types before writing records. Track missing-field counts so a changed page layout is visible.
- Normalize and deduplicate: Normalize whitespace, dates, prices, and absolute URLs; deduplicate with a stable source key instead of display text.
- Limit load: Configure per-domain concurrency and delay. Scrapy supports these controls and AutoThrottle; a faster crawl is not automatically a better one. AutoThrottle
- Recover cleanly: Use request timeouts, bounded retries, structured logs, and checkpoints. Store the input URL or crawl state needed to resume interrupted work.
- Preserve provenance: Keep the source URL and, where useful, the raw response or retrieval time so an anomalous record can be traced to its origin.
There is no authoritative, comparable page-per-second figure established here. Actual throughput depends on the target, network, page weight, rendering needs, concurrency, and permitted request rate; measure your own crawl under the intended settings rather than applying a generic speed claim.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Troubleshoot common failures
The output is empty or fields are missing
Inspect the saved response and confirm that it contains the content. If the page returns a consent screen, an error page, or a JavaScript shell, selectors for the expected content will return nothing. Re-check selectors against the actual markup and test one page before running the full crawl.
Pagination stops too early or loops
Inspect the next-link selector and the actual href on the last page. Resolve relative hrefs against the response URL, track visited URLs, and impose a crawl limit. If the site uses a button or script rather than a link, identify the resulting network request or browser interaction.
The script parses an error page as normal content
Check HTTP status before parsing. With Requests, call raise_for_status() or branch on the status code. With Playwright, inspect response status: 404 and 503 are completed HTTP responses, not successful application results. Playwright request documentation
Requests time out or fail intermittently
Set explicit timeouts, log the failing URL and exception, and retry only transient failures a bounded number of times with a delay. Do not retry indefinitely; checkpoint completed work and review whether concurrency is too high for the host.
The Tool Desk
Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Best Value
Records contain duplicates or inconsistent values
Deduplicate by a stable identifier or canonical URL, not by a mutable title. Normalize formatting before validation, and retain raw values when transformations such as currency parsing could otherwise obscure the source.
Or skip the browser setup
For a one-off screenshot or a browser-based capture step in a workflow, ScreenshotNeo provides a website screenshot API and MCP server. It does not replace a structured crawler when you need records from many pages, but it can avoid setting up browser capture yourself. A single GET request can return PNG, JPEG, WebP, or PDF; see the ScreenshotNeo API documentation.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp
ScreenshotNeo accepts cookie or consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each of those steps can be turned off. Bot checks and CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and the response includes X-Page-Verdict and X-Billed headers. Its MCP server offers take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients. The Free plan includes 1,000 screenshots per month without a card; paid plans start at $5 for 3,000 screenshots.
Sign up for 1,000 free screenshots a month, no card required.
Free tools Windows power users keep installed
One-click scans. No signup required.
Frequently asked questions
Should I scrape pages into CSV or JSON?
Choose the format that fits the next step: CSV is convenient for rows with a fixed set of columns, while JSON Lines works well for incremental writing and records with structured or optional fields.
How do I know when a crawl is complete?
For link-following pagination, completion usually means the parsed next-page link is absent and the crawler has no remaining scheduled requests. Also check that the final output passes your expected record-count and required-field checks.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

