Do these 3 things before closing this tab:
1Repair Windows errors before they cause bigger problems2Fix the driver behind crashes, sound loss and screen glitches3Clear out junk files and repair common Windows errorsSome links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
Reliable web scraping is a pipeline, not a browser script: define an authorized scope, locate the source that actually carries the data, acquire it at a tolerable rate, extract and validate records, persist state, and monitor for change. Start with a documented API, export, or ordinary HTTP request. Reproduce a browser network request when that is practical; use a headless browser only when the interaction or rendered DOM is genuinely required.
This approach lowers resource use, reduces parsing work, and makes failures diagnosable. The examples below show how to combine direct requests, Scrapy, and Playwright without treating robots.txt, retries, or anti-bot responses as permission to push harder.
1. Define scope, permission, and success criteria
Write down the target domains and paths, fields to collect, intended use, retention period, and expected request volume before writing a spider. Identify whether the site publishes an API, feed, bulk export, or search endpoint. An export that supplies the needed records is usually less work for both you and the site than crawling every page.
The Tool Desk
Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Check the site’s terms, authentication requirements, privacy obligations, intellectual-property constraints, and any contractual limits that apply to your jurisdiction and use case. RFC 9309 (the September 2022 Robots Exclusion Protocol standard) is explicit that robots.txt instructions are not access authorization. A public URL is not automatically permission to collect, retain, or republish every field it exposes.
#1 Best Overall
The European Data Protection Board’s Guidelines 03/2026 page is an open consultation (feedback is listed for 8 July through 30 October 2026) focused on scraping in generative-AI contexts. It is draft consultation material, not final law or a universal rule for every crawler. Obtain appropriate legal and privacy review for production workloads.
2. Find the real data source before rendering a browser
Inspect the ordinary response
Request one representative URL with a normal HTTP client and save the response. Search for the fields you need, JSON-LD, embedded state, pagination links, or a script that points to an API. HTML selectors should be your fallback when the data is already present in the document, not your first assumption.
Inspect network requests for dynamic pages
If the initial HTML lacks the records, open browser developer tools, reload the page, and filter the Network panel for Fetch/XHR. Record the request method, URL, query parameters or body, required headers, cookies, and pagination token. Reproduce that request directly and parse its JSON, HTML, or XML response. Preserve only headers that are actually required; do not copy volatile browser headers blindly.
Recommended Free Tools
import requests
endpoint = "https://target.example/api/items"
params = {"page": 1, "limit": 100}
r = requests.get(endpoint, params=params, timeout=30)
r.raise_for_status()
data = r.json()
for item in data.get("items", []):
print(item.get("id"), item.get("name"))
Direct acquisition normally transfers less data and avoids browser startup. A headless browser is justified when the request cannot reasonably be reproduced, when a login or multi-step interaction is required and authorized, or when the rendered DOM itself is the deliverable.
3. Choose the smallest tool that fits
| Need | Better starting point | Trade-off |
|---|---|---|
| Documented records, search, or bulk export | Official API or export | Verify its terms, authentication, quotas, and update semantics. |
| Data exposed through a browser network request | Direct HTTP request, optionally inside Scrapy | Usually lighter and more structured; you must reproduce request details accurately. |
| Many pages, link discovery, scheduling, retries, and deduplication | Scrapy | Requires crawler configuration and target-specific parsing logic. |
| Browser interaction, rendered DOM, or screenshots | Playwright | Full browser automation consumes more CPU and memory and adds integration complexity. |
Compare candidates on completeness, request volume, rendering fidelity, throughput, maintenance cost, observability, and fit with the publisher’s documented access method. No tool is universally fastest; workload and target behavior determine the result.
4. Build a controlled Scrapy crawler
Scrapy supplies scheduling, middleware, duplicate filtering, retries, and crawl-level controls. Keep extraction code focused on the target schema and put operational policy in settings.
import scrapy
class ProductSpider(scrapy.Spider):
name = "products"
allowed_domains = ["target.example"]
start_urls = ["https://target.example/catalog"]
custom_settings = {
"ROBOTSTXT_OBEY": True,
"USER_AGENT": "catalog-monitor/1.0 ([email protected])",
"CONCURRENT_REQUESTS_PER_DOMAIN": 2,
"DOWNLOAD_DELAY": 0.75,
"AUTOTHROTTLE_ENABLED": True,
"RETRY_HTTP_CODES": [429, 500, 502, 503, 504],
"FEEDS": {"items.jsonl": {"format": "jsonlines", "overwrite": True}},
}
def parse(self, response):
for card in response.css("article.product"):
yield {
"url": response.urljoin(card.css("a::attr(href)").get()),
"name": card.css("h2::text").get(default="").strip(),
"price": card.css(".price::text").get(default="").strip(),
}
next_url = response.css("a[rel='next']::attr(href)").get()
if next_url:
yield response.follow(next_url, callback=self.parse)
Enable Scrapy’s robots middleware and set the user agent used for matching. If you connect Playwright to Scrapy, retain Scrapy’s scheduling, robots, retry, and duplicate-filtering behavior through an integration such as scrapy-playwright instead of bypassing crawler controls.
5. Use Playwright only for browser-required work
Playwright’s Python library provides synchronous and asynchronous APIs and can launch Chromium, Firefox, or WebKit. The asynchronous form is convenient when several independent browser tasks must be coordinated.
import asyncio
from playwright.async_api import async_playwright
async def main():
async with async_playwright() as p:
browser = await p.chromium.launch(headless=True)
page = await browser.new_page()
await page.goto("https://target.example/catalog", wait_until="networkidle")
await page.locator(".load-more").click()
await page.wait_for_selector("article.product")
records = await page.locator("article.product").evaluate_all(
"els => els.map(e => ({name: e.querySelector('h2')?.textContent?.trim(), "
"price: e.querySelector('.price')?.textContent?.trim()}))"
)
print(records)
await browser.close()
asyncio.run(main())
Limit browser concurrency, close contexts promptly, and reuse a browser process where appropriate. Do not launch a browser merely to parse an API response that an HTTP client can obtain directly. Browser automation is also not a license to defeat CAPTCHAs, access controls, or identity checks; obtain authorization or use the publisher’s supported interface.
6. Respect robots.txt and control load
Interpret the protocol correctly
Rules belong at /robots.txt. After a successful fetch, follow the parseable rules for the matching user agent. Under RFC 9309, a 4xx response makes the file unavailable and may permit access under that protocol, while server or network errors make it unreachable and require complete disallow under the standard. Those are protocol behaviors, not a legal permission decision.
Rank #3
Translate crawl directives yourself
Current Scrapy documentation notes that Scrapy does not automatically act on Crawl-delay or Request-rate. Convert applicable expectations into per-domain delay and concurrency settings. Begin conservatively, raise concurrency gradually, and measure the result.
Free tools Windows power users keep installed
One-click scans. No signup required.
Use server signals as a brake
- Treat 429 and 503 responses, rising retry counts, increasing latency, and explicit block pages as reasons to slow down or pause.
- Honor
Retry-Afterwhen supplied, and add exponential backoff with jitter so many workers do not retry simultaneously. - Prefer a published API or export when one exists.
- Do not rotate identities or proxies to keep pushing through a rate limit or access control. That changes the technical symptom, not your authorization.
7. Make extraction and validation dependable
Parse JSON with a schema-aware model when possible. For HTML or XML, isolate selectors in versioned parser modules. Treat markup, embedded scripts, and field order as variable input.
- Define required fields, allowed types, units, and normalization rules (for example, decimal currency and UTC timestamps).
- Reject or quarantine records missing identifiers instead of silently emitting partial rows.
- Track counts for discovered, fetched, parsed, valid, rejected, and duplicate records.
- Store the source URL and retrieval time with each record so corrections can be traced.
- For PDFs, locate the underlying file resource first; apply format-specific extraction and use OCR only for image-based pages where necessary.
Validate pagination explicitly. APIs may use cursor tokens rather than page numbers; persist the next cursor only after the current batch has been validated and written. Stop when the service signals completion, not when a guessed page limit is reached.
8. Persist state, retries, and cache deliberately
Separate crawl state from site-specific parsing. A durable queue should record the URL or API request, attempt count, first-seen time, last error, and completion state. Make output writes idempotent with a stable key so a retry cannot create a second copy of the same record.
Cache responses during development and parser testing when the content is not time-sensitive. In production, choose a retention period that matches freshness requirements and the target’s terms. Record status code, response size, latency, and a content hash; a sudden hash change with a stable status often indicates a template or schema change.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Rank #4
Retry transient transport failures and selected 5xx responses, not validation errors caused by your own parser. Cap attempts, use exponential backoff, and send exhausted jobs to a review queue. A timeout should have a separate budget for connection, response, and browser navigation where your client supports it.
9. Monitor drift and operational health
Dashboards should show request rate, status distribution, latency percentiles, retry rate, timeout rate, bytes transferred, and records per successful request. Data-quality charts should show field missingness, type failures, duplicate rate, and record volume by source.
Alert on sustained 429/503 responses, a sharp rise in empty pages, selector match rates falling below their normal range, or a cursor that stops advancing. Keep a small fixture set of representative responses and run parser tests against it on every code change. When a site changes, pause broad collection, capture a new fixture, update the parser, and replay it before resuming.
10. Troubleshooting common failures
| Symptom | Likely cause | Fix |
|---|---|---|
| HTML contains no records but the browser shows them | Data arrives through a Fetch/XHR request or client-side script. | Inspect Network requests and reproduce the structured request; use Playwright only if reproduction is impractical. |
| Every request returns 429 | Concurrency or frequency exceeds the target’s tolerance. | Reduce per-domain concurrency, increase delay, honor Retry-After, and request an official quota if available. |
| Intermittent 503 and rising latency | Overload, maintenance, or an upstream dependency. | Pause or back off; cap retries and inspect status trends before resuming. |
| Scrapy follows links outside the intended site | Broad selectors or missing domain/path limits. | Set allowed domains, restrict start paths, and validate every followed URL. |
| Duplicate records after a restart | Output writes are not idempotent or the duplicate filter state was lost. | Use a stable record key, durable job state, and a deduplication constraint in the sink. |
| Playwright waits forever | The page never reaches the chosen load condition or selector. | Use an explicit timeout, wait for a meaningful selector or response, and capture diagnostics before retrying. |
| Parser suddenly emits empty fields | Markup or JSON schema drift. | Compare saved fixtures and content hashes, version the parser, and quarantine affected output. |
| robots.txt cannot be fetched | 4xx availability differs from a server/network failure. | Apply RFC 9309 handling for the response class, then make a separate legal and authorization decision. |
11. Capture rendered pages without maintaining a browser service
When your deliverable is a clean screenshot or PDF rather than extracted records, ScreenshotNeo is the first screenshot API to try: it removes consent banners, newsletter popups, and chat widgets before capture, bills only clean shots, and has the lowest paid plan listed here.
Or skip the browser setup
One GET request returns PNG, JPEG, WebP, or PDF output. The API accepts full-page capture, CSS-selector element capture, device and viewport settings, dark mode, retina scale, waits, custom CSS and JavaScript, headers and cookies, blocking rules, PDF options, caching TTLs, signed links, asynchronous jobs, webhooks, and bulk capture. Its MCP server exposes take_screenshot, get_page_info, and capture_pdf to Claude, Cursor, and other MCP clients.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
See the ScreenshotNeo API documentation for all parameters and response headers. Responses identify the page verdict and whether the shot was billed: bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing.
Best Value
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
The Free plan includes 1,000 shots per month with no card. Paid plans start at $5 for 3,000 shots; yearly billing provides two months free, and every feature is included on every plan. Create a free ScreenshotNeo account to begin.
12. A practical operating checklist
- Confirm scope, authorization, retention, and expected volume.
- Check for an API, export, or feed.
- Inspect one ordinary response and then the browser’s network traffic if needed.
- Choose direct HTTP, Scrapy, or Playwright based on the actual requirement.
- Enable robots handling, identify your crawler, and set conservative per-domain limits.
- Add bounded retries, caching, durable state, and idempotent writes.
- Validate fields and monitor missingness, latency, status codes, and drift.
- Pause on overload or access-control signals and investigate before changing the design.
Frequently asked questions
Should I run one crawler for every domain?
Usually not. Keep shared queue, state, and observability components, but isolate per-domain policies and parsers so one site’s delay, schema, or failure mode cannot silently affect another.
Quick wins for a faster PC:
Clear out junk files and repair common Windows errorsFree Scan →Scan for outdated or missing drivers - takes under a minuteDriver Scan →Is a successful HTTP status proof that a record is usable?
No. A 200 response can contain an error object, an empty application state, a consent wall, or a changed schema. Validate the payload and required fields before marking the request successful.
When should a crawl become a scheduled data pipeline?
When freshness, retries, and historical corrections matter. At that point, persist raw responses or hashes, version parsers, make runs resumable, and alert on both transport and data-quality failures instead of treating the job as a one-off script.
Frequently Asked Questions
Can robots.txt grant permission to scrape a site?
No. RFC 9309 defines crawler instructions and expressly says they are not access authorization; review the site’s terms, controls, and applicable law separately.
How can I tell whether browser rendering is really necessary?
Inspect the initial response and the browser’s Fetch/XHR traffic. If the needed data comes from a reproducible request, use that request; reserve Playwright for interactions or rendered output that cannot reasonably be reproduced.
What is the safest response to repeated 429 or 503 errors?
Reduce concurrency, increase delay, honor Retry-After, and pause if errors continue. Do not rotate identities to keep forcing requests through a limit.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

