A reliable web-scraping data pipeline is a sequence of separable stages: define what you may collect, schedule URLs, download politely, parse fields, clean and validate items, persist raw and structured data, and orchestrate and monitor every run. Start with a self-hosted Scrapy pipeline for control; add browser rendering only for pages that create their data in JavaScript; use Airflow when scrapes must trigger downstream transformations or analytics.
Start with a data contract and source policy
Before writing a spider, document the output and the boundaries of the crawl. This prevents a parser from quietly changing your dataset when a website redesigns its HTML.
- Allowed scope: domains, URL patterns, robots.txt directives, terms, authentication boundaries and any pages that must never be requested.
- Fields: names, data types, required versus optional values, units, timezone and an example of a valid record.
- Freshness: how often each source needs updating and how long raw responses and cleaned records are retained.
- Provenance: source URL, retrieval timestamp, parser version and (where useful) a content hash.
- Failure policy: retry limits, acceptable null rates, duplicate rules and what should page an operator.
Robots.txt is a signal to honor alongside site terms and applicable law; it is not authorization by itself. Translate Crawl-delay and Request-rate into your own downloader settings because Scrapy does not apply those directives automatically.
Use a pipeline with explicit stages
Scrapy’s documented flow is engine, scheduler, downloader, spider and item pipeline. Keep those responsibilities separate so you can change storage or scheduling without rewriting extraction code.
PC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware match#1 Best Overall
- Discovery and scheduling: seed URLs, assign priorities, deduplicate requests and enforce per-domain concurrency.
- Downloading: apply timeouts, retries, caching and measured concurrency. Record status codes and latency.
- Parsing: select fields with CSS or XPath and yield typed items rather than writing directly to a database.
- Item processing: normalize, validate, deduplicate and attach provenance.
- Persistence: retain lawful raw responses for replay, then write cleaned rows to a database, warehouse or object store.
- Orchestration: schedule runs, trigger dependent jobs and alert on drift or freshness failures.
Build the first Scrapy project
Create an item contract
Define the fields your spider promises. A small dataclass makes type and validation rules visible to reviewers.
from dataclasses import dataclass
from datetime import datetime
from typing import Optional
@dataclass
class Product:
url: str
name: str
price: Optional[float]
retrieved_at: datetime
parser_version: str = "products-v1"
Configure polite downloading
In settings.py, begin conservatively and tune from measurements rather than guessing:
ROBOTSTXT_OBEY = True
DOWNLOAD_DELAY = 2.0
CONCURRENT_REQUESTS_PER_DOMAIN = 2
DOWNLOAD_TIMEOUT = 30
RETRY_ENABLED = True
RETRY_TIMES = 2
HTTPCACHE_ENABLED = True
FEEDS = {
"s3://your-bucket/products/%(time)s.json": {
"format": "json",
"encoding": "utf8",
"store_empty": False,
"overwrite": False,
}
}
Map each site’s published delay or request-rate guidance to DOWNLOAD_DELAY and concurrency. If 429 or 503 responses, latency, or ban-page signals increase, lower concurrency and add delay before considering more retries.
Parse and yield typed items
import scrapy
from datetime import datetime, timezone
from .items import Product
class ProductSpider(scrapy.Spider):
name = "products"
allowed_domains = ["example.com"]
start_urls = ["https://example.com/catalog"]
def parse(self, response):
for card in response.css("article.product"):
raw_price = card.css(".price::text").get()
price = float(raw_price.replace("$", "").replace(",", "")) if raw_price else None
yield Product(
url=response.urljoin(card.css("a::attr(href)").get()),
name=card.css("h2::text").get(default="").strip(),
price=price,
retrieved_at=datetime.now(timezone.utc),
)
yield from response.follow_all(response.css("a.next::attr(href)"), self.parse)
Use selectors that tolerate harmless whitespace but fail loudly when required fields disappear. Treat parser changes as schema changes: version extractors and monitor field-level null rates.
Do these 3 things before closing this tab:
1Fix the driver behind crashes, sound loss and screen glitches2Repair Windows errors before they cause bigger problems3Scan for outdated or missing drivers - takes under a minuteClean, validate and deduplicate in item pipelines
Scrapy’s item pipeline is the natural place to cleanse fields, validate records, drop duplicates and persist items. Keep these operations deterministic and ordered.
from scrapy.exceptions import DropItem
class ValidateProductPipeline:
seen = set()
def process_item(self, item, spider):
item["name"] = " ".join(item["name"].split())
if not item["url"] or not item["name"]:
raise DropItem("missing required field")
if item["price"] is not None and item["price"] < 0:
raise DropItem("negative price")
key = (item["url"], item["name"])
if key in self.seen:
raise DropItem("duplicate")
self.seen.add(key)
return item
For multi-process or multi-day jobs, replace the in-memory set with a database uniqueness constraint or a durable key-value store. Store retrieval time, source URL and parser version with every cleaned record so an analyst can explain where a value came from.
Rank #2
Choose an export and storage boundary
Feed exports can write JSON, CSV or XML directly to storage backends such as Amazon S3. A practical pattern is immutable raw snapshots in object storage plus cleaned, queryable rows in a database or warehouse. Keep the raw layer only where collection and retention are lawful, and encrypt credentials and sensitive fields.
Handle JavaScript-heavy pages selectively
First determine whether the required data exists in the initial HTML or an XHR/JSON response. Prefer the response when possible: it is faster and easier to retry. When content is genuinely rendered in a browser, integrate a renderer such as scrapy-playwright for only those requests.
- Route browser work by URL or an explicit request flag instead of enabling it globally.
- Wait for a meaningful selector or network-idle condition, not an arbitrary long sleep.
- Set a browser timeout and cap concurrent pages; browsers consume substantially more memory than HTTP requests.
- Capture the final URL, status and a screenshot or HTML snapshot for debugging when a selector fails.
Do not use browser automation for every page by default. A mixed queue—ordinary HTTP first, browser fallback only when a page requires it—usually gives better throughput and lower operating cost than an all-browser crawler.
Schedule recurring runs with Airflow
When a scrape must trigger transformations, quality checks or analytics, put it in an Airflow DAG. Airflow documents ETL/ELT as a core use case and supports datasets, object storage and extensible providers. Its 2023 survey reports that 90% of respondents use Airflow for ETL/ELT powering analytics use cases; that figure describes survey respondents, not every Airflow installation.
from datetime import datetime
from airflow import DAG
from airflow.operators.bash import BashOperator
with DAG(
dag_id="products_scrape",
start_date=datetime(2024, 1, 1),
schedule="0 3 * * *",
catchup=False,
tags=["scraping"],
) as dag:
scrape = BashOperator(
task_id="scrape",
bash_command="scrapy crawl products -s JOBDIR=/tmp/products_job",
)
transform = BashOperator(
task_id="transform",
bash_command="python jobs/transform_products.py",
)
quality = BashOperator(
task_id="quality_check",
bash_command="python jobs/check_products.py",
)
scrape >> transform >> quality
For production, use durable worker storage rather than a local temporary directory, make each task idempotent, and publish a dataset or object-store completion event for downstream consumers. Separate backfills from regular schedules so a historical replay does not overload the source.
Make reliability measurable
Emit metrics per run and per domain:
- requests attempted, completed, retried and abandoned;
- HTTP status distribution, latency percentiles and timeout count;
- items yielded, validation drops, duplicate rate and field-level null rates;
- freshness of the newest accepted record and parser version;
- browser-page count, render timeout rate and memory use.
Alert on sustained 429/503 responses, a sudden rise in blank or ban pages, a drop in item yield, or a freshness deadline breach. Keep retries bounded: retrying a blocked request indefinitely increases pressure without improving data.
Free tools Windows power users keep installed
One-click scans. No signup required.
Choose an operating model
| Approach | JavaScript rendering | Rate and retry control | Scheduling and dependencies | Operating trade-off |
|---|---|---|---|---|
| Self-hosted Scrapy | HTTP by default; add a browser integration when needed | Direct control of delays, concurrency and retries | Pair with a scheduler or Airflow | Maximum control; you operate workers, storage and monitoring |
| Browser-augmented Scrapy | Suitable for client-rendered pages | Control remains in your code, with higher resource demand | Same orchestration options | Useful fallback, but slower and more memory-intensive |
| Hosted scraping API | Depends on the provider’s rendering and extraction features | Provider exposes API-level controls | Often includes asynchronous runs or schedules; verify each service | Less infrastructure to operate, with provider dependency and data-residency questions to assess |
| Airflow-based orchestration | Not a renderer; it schedules your scraper | Delegated to the scraper task | Strong dependency handling and dataset-triggered workflows | An orchestrator, not a replacement for a crawler |
Common failures and fixes
429 or 503 responses rise
Cause: concurrency or retry pressure is too high. Lower per-domain concurrency, increase delay, honor published request limits, and reduce retry attempts. Check whether a ban page is being returned with a 200 status.
Selectors suddenly return null
Cause: a template or parser contract changed. Save a failing response, compare its structure with a known-good snapshot, update and version the extractor, then run a null-rate quality check before deploying.
Only part of the page is present
Cause: the data is client-rendered or loaded after the initial response. Inspect network responses first; if no usable endpoint exists, route that URL to a bounded Playwright-rendered request and wait for a specific selector.
Duplicate rows appear across runs
Cause: the deduplication key is only in memory or is not stable. Define a canonical key, enforce uniqueness at the persistence layer, and make upserts idempotent.
Airflow task succeeds but data is missing
Cause: the task checked process exit status but not output quality. Make the task fail when freshness, minimum row count or required-field checks fail, and publish an explicit completion marker only after validation.
Credentials or private pages leak into logs
Cause: verbose request logging or unredacted URLs. Store secrets in the orchestrator’s connection manager, redact authorization headers and query parameters, and restrict raw-response access.
Rank #4
Or skip the browser setup
When your job needs a rendered visual snapshot rather than a maintained browser worker, ScreenshotNeo provides a website screenshot API and MCP server. One GET request returns PNG, JPEG, WebP or PDF; it can load lazy images, target a CSS-selected element, set dark mode, viewport and retina scale, wait for a selector, delay or network idle, run custom CSS or JavaScript, click an element, hide selectors, set headers, cookies, user agent, authorization, timezone or geolocation, block ads, trackers, requests or resource types, resize images, cache with a TTL, create signed links, run asynchronous jobs with signed webhooks, capture up to 100 URLs per bulk call, expose usage and OpenAPI endpoints, and accept parameter names used by other screenshot APIs.
Before capture it accepts cookie or consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups and chat widgets; each step can be disabled. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads and cache hits are not billed, and response headers report the page verdict and whether the request was billed.
Quick wins for a faster PC:
Clear out junk files and repair common Windows errorsFree Scan →Scan for outdated or missing drivers - takes under a minuteDriver Scan →curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
See the ScreenshotNeo API documentation for parameters and response formats. The same request in Python:
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
And in Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`HTTP ${res.status}`);
const fs = await import('node:fs/promises');
await fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer()));
ScreenshotNeo’s MCP server exposes take_screenshot, get_page_info and capture_pdf to Claude, Cursor and other MCP clients, so an AI agent can request captures without custom browser glue. Plans include 1,000 shots per month free with no card; Starter is $5 for 3,000, Growth $15 for 15,000, Pro $39 for 60,000, Scale $99 for 250,000 and Business $249 for 1,000,000. Yearly billing gives two months free, and every feature is on every plan. Create a free ScreenshotNeo account to start with the 1,000 monthly shots.
Final implementation checklist
- Policy, authorization, robots.txt and retention decisions are recorded.
- Scheduler, downloader, parser, pipeline, storage and orchestration have separate responsibilities.
- Delays, concurrency and retries reflect observed server responses.
- Raw evidence and provenance are retained where lawful.
- Required fields, duplicates, freshness and parser drift can fail a run.
- Browser rendering is a targeted fallback, not the default path.
- Every recurring run is idempotent, observable and recoverable.
Frequently Asked Questions
Should I store raw HTML as well as cleaned rows?
Store raw responses or snapshots when lawful and useful for replay, then write cleaned records separately. Apply an explicit retention policy and protect access because raw pages may contain sensitive data.
How do I decide between a hosted API and Scrapy?
Use Scrapy when you need direct control over crawling, code and data handling. Consider a hosted API when the team prefers API-key calls, asynchronous runs, exports and schedules without operating crawler infrastructure; evaluate rendering, residency and lock-in for the specific provider.
Recommended Free Tools
Is Airflow a web scraper?
No. Airflow orchestrates scraper tasks and downstream ETL/ELT, while the crawler still handles requests, parsing, politeness and persistence.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




