Driver FixRecommendedSound, Wi-Fi or graphics acting up? Check drivers firstFind missing or outdated drivers fast.Check DriversOctober DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsSlow PC?RecommendedPC slow today? Run a repair scan before it gets worseResolve common Windows issues and optimize system performance.Scan Now×
Skip to content
airflow

How to Build a Web Scraping Data Pipeline That Scales

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A reliable web-scraping data pipeline is a sequence of separable stages: define what you may collect, schedule URLs, download politely, parse fields, clean and validate items, persist raw and structured data, and orchestrate and monitor every run. Start with a self-hosted Scrapy pipeline for control; add browser rendering only for pages that create their data in JavaScript; use Airflow when scrapes must trigger downstream transformations or analytics.

Start with a data contract and source policy

Before writing a spider, document the output and the boundaries of the crawl. This prevents a parser from quietly changing your dataset when a website redesigns its HTML.

  • Allowed scope: domains, URL patterns, robots.txt directives, terms, authentication boundaries and any pages that must never be requested.
  • Fields: names, data types, required versus optional values, units, timezone and an example of a valid record.
  • Freshness: how often each source needs updating and how long raw responses and cleaned records are retained.
  • Provenance: source URL, retrieval timestamp, parser version and (where useful) a content hash.
  • Failure policy: retry limits, acceptable null rates, duplicate rules and what should page an operator.

Robots.txt is a signal to honor alongside site terms and applicable law; it is not authorization by itself. Translate Crawl-delay and Request-rate into your own downloader settings because Scrapy does not apply those directives automatically.

Use a pipeline with explicit stages

Scrapy’s documented flow is engine, scheduler, downloader, spider and item pipeline. Keep those responsibilities separate so you can change storage or scheduling without rewriting extraction code.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  1. Discovery and scheduling: seed URLs, assign priorities, deduplicate requests and enforce per-domain concurrency.
  2. Downloading: apply timeouts, retries, caching and measured concurrency. Record status codes and latency.
  3. Parsing: select fields with CSS or XPath and yield typed items rather than writing directly to a database.
  4. Item processing: normalize, validate, deduplicate and attach provenance.
  5. Persistence: retain lawful raw responses for replay, then write cleaned rows to a database, warehouse or object store.
  6. Orchestration: schedule runs, trigger dependent jobs and alert on drift or freshness failures.

Build the first Scrapy project

Create an item contract

Define the fields your spider promises. A small dataclass makes type and validation rules visible to reviewers.

from dataclasses import dataclass
from datetime import datetime
from typing import Optional

@dataclass
class Product:
    url: str
    name: str
    price: Optional[float]
    retrieved_at: datetime
    parser_version: str = "products-v1"

Configure polite downloading

In settings.py, begin conservatively and tune from measurements rather than guessing:

ROBOTSTXT_OBEY = True
DOWNLOAD_DELAY = 2.0
CONCURRENT_REQUESTS_PER_DOMAIN = 2
DOWNLOAD_TIMEOUT = 30
RETRY_ENABLED = True
RETRY_TIMES = 2
HTTPCACHE_ENABLED = True
FEEDS = {
    "s3://your-bucket/products/%(time)s.json": {
        "format": "json",
        "encoding": "utf8",
        "store_empty": False,
        "overwrite": False,
    }
}

Map each site’s published delay or request-rate guidance to DOWNLOAD_DELAY and concurrency. If 429 or 503 responses, latency, or ban-page signals increase, lower concurrency and add delay before considering more retries.

Parse and yield typed items

import scrapy
from datetime import datetime, timezone
from .items import Product

class ProductSpider(scrapy.Spider):
    name = "products"
    allowed_domains = ["example.com"]
    start_urls = ["https://example.com/catalog"]

    def parse(self, response):
        for card in response.css("article.product"):
            raw_price = card.css(".price::text").get()
            price = float(raw_price.replace("$", "").replace(",", "")) if raw_price else None
            yield Product(
                url=response.urljoin(card.css("a::attr(href)").get()),
                name=card.css("h2::text").get(default="").strip(),
                price=price,
                retrieved_at=datetime.now(timezone.utc),
            )
        yield from response.follow_all(response.css("a.next::attr(href)"), self.parse)

Use selectors that tolerate harmless whitespace but fail loudly when required fields disappear. Treat parser changes as schema changes: version extractors and monitor field-level null rates.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Clean, validate and deduplicate in item pipelines

Scrapy’s item pipeline is the natural place to cleanse fields, validate records, drop duplicates and persist items. Keep these operations deterministic and ordered.

from scrapy.exceptions import DropItem

class ValidateProductPipeline:
    seen = set()

    def process_item(self, item, spider):
        item["name"] = " ".join(item["name"].split())
        if not item["url"] or not item["name"]:
            raise DropItem("missing required field")
        if item["price"] is not None and item["price"] < 0:
            raise DropItem("negative price")
        key = (item["url"], item["name"])
        if key in self.seen:
            raise DropItem("duplicate")
        self.seen.add(key)
        return item

For multi-process or multi-day jobs, replace the in-memory set with a database uniqueness constraint or a durable key-value store. Store retrieval time, source URL and parser version with every cleaned record so an analyst can explain where a value came from.

Choose an export and storage boundary

Feed exports can write JSON, CSV or XML directly to storage backends such as Amazon S3. A practical pattern is immutable raw snapshots in object storage plus cleaned, queryable rows in a database or warehouse. Keep the raw layer only where collection and retention are lawful, and encrypt credentials and sensitive fields.

Handle JavaScript-heavy pages selectively

First determine whether the required data exists in the initial HTML or an XHR/JSON response. Prefer the response when possible: it is faster and easier to retry. When content is genuinely rendered in a browser, integrate a renderer such as scrapy-playwright for only those requests.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Route browser work by URL or an explicit request flag instead of enabling it globally.
  • Wait for a meaningful selector or network-idle condition, not an arbitrary long sleep.
  • Set a browser timeout and cap concurrent pages; browsers consume substantially more memory than HTTP requests.
  • Capture the final URL, status and a screenshot or HTML snapshot for debugging when a selector fails.

Do not use browser automation for every page by default. A mixed queue—ordinary HTTP first, browser fallback only when a page requires it—usually gives better throughput and lower operating cost than an all-browser crawler.

Schedule recurring runs with Airflow

When a scrape must trigger transformations, quality checks or analytics, put it in an Airflow DAG. Airflow documents ETL/ELT as a core use case and supports datasets, object storage and extensible providers. Its 2023 survey reports that 90% of respondents use Airflow for ETL/ELT powering analytics use cases; that figure describes survey respondents, not every Airflow installation.

from datetime import datetime
from airflow import DAG
from airflow.operators.bash import BashOperator

with DAG(
    dag_id="products_scrape",
    start_date=datetime(2024, 1, 1),
    schedule="0 3 * * *",
    catchup=False,
    tags=["scraping"],
) as dag:
    scrape = BashOperator(
        task_id="scrape",
        bash_command="scrapy crawl products -s JOBDIR=/tmp/products_job",
    )
    transform = BashOperator(
        task_id="transform",
        bash_command="python jobs/transform_products.py",
    )
    quality = BashOperator(
        task_id="quality_check",
        bash_command="python jobs/check_products.py",
    )
    scrape >> transform >> quality

For production, use durable worker storage rather than a local temporary directory, make each task idempotent, and publish a dataset or object-store completion event for downstream consumers. Separate backfills from regular schedules so a historical replay does not overload the source.

Make reliability measurable

Emit metrics per run and per domain:

  • requests attempted, completed, retried and abandoned;
  • HTTP status distribution, latency percentiles and timeout count;
  • items yielded, validation drops, duplicate rate and field-level null rates;
  • freshness of the newest accepted record and parser version;
  • browser-page count, render timeout rate and memory use.

Alert on sustained 429/503 responses, a sudden rise in blank or ban pages, a drop in item yield, or a freshness deadline breach. Keep retries bounded: retrying a blocked request indefinitely increases pressure without improving data.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Choose an operating model

Approach JavaScript rendering Rate and retry control Scheduling and dependencies Operating trade-off
Self-hosted Scrapy HTTP by default; add a browser integration when needed Direct control of delays, concurrency and retries Pair with a scheduler or Airflow Maximum control; you operate workers, storage and monitoring
Browser-augmented Scrapy Suitable for client-rendered pages Control remains in your code, with higher resource demand Same orchestration options Useful fallback, but slower and more memory-intensive
Hosted scraping API Depends on the provider’s rendering and extraction features Provider exposes API-level controls Often includes asynchronous runs or schedules; verify each service Less infrastructure to operate, with provider dependency and data-residency questions to assess
Airflow-based orchestration Not a renderer; it schedules your scraper Delegated to the scraper task Strong dependency handling and dataset-triggered workflows An orchestrator, not a replacement for a crawler

Common failures and fixes

429 or 503 responses rise

Cause: concurrency or retry pressure is too high. Lower per-domain concurrency, increase delay, honor published request limits, and reduce retry attempts. Check whether a ban page is being returned with a 200 status.

Selectors suddenly return null

Cause: a template or parser contract changed. Save a failing response, compare its structure with a known-good snapshot, update and version the extractor, then run a null-rate quality check before deploying.

Only part of the page is present

Cause: the data is client-rendered or loaded after the initial response. Inspect network responses first; if no usable endpoint exists, route that URL to a bounded Playwright-rendered request and wait for a specific selector.

Duplicate rows appear across runs

Cause: the deduplication key is only in memory or is not stable. Define a canonical key, enforce uniqueness at the persistence layer, and make upserts idempotent.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Airflow task succeeds but data is missing

Cause: the task checked process exit status but not output quality. Make the task fail when freshness, minimum row count or required-field checks fail, and publish an explicit completion marker only after validation.

Credentials or private pages leak into logs

Cause: verbose request logging or unredacted URLs. Store secrets in the orchestrator’s connection manager, redact authorization headers and query parameters, and restrict raw-response access.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup

When your job needs a rendered visual snapshot rather than a maintained browser worker, ScreenshotNeo provides a website screenshot API and MCP server. One GET request returns PNG, JPEG, WebP or PDF; it can load lazy images, target a CSS-selected element, set dark mode, viewport and retina scale, wait for a selector, delay or network idle, run custom CSS or JavaScript, click an element, hide selectors, set headers, cookies, user agent, authorization, timezone or geolocation, block ads, trackers, requests or resource types, resize images, cache with a TTL, create signed links, run asynchronous jobs with signed webhooks, capture up to 100 URLs per bulk call, expose usage and OpenAPI endpoints, and accept parameter names used by other screenshot APIs.

Before capture it accepts cookie or consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups and chat widgets; each step can be disabled. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads and cache hits are not billed, and response headers report the page verdict and whether the request was billed.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

See the ScreenshotNeo API documentation for parameters and response formats. The same request in Python:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)

And in Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`HTTP ${res.status}`);
const fs = await import('node:fs/promises');
await fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer()));

ScreenshotNeo’s MCP server exposes take_screenshot, get_page_info and capture_pdf to Claude, Cursor and other MCP clients, so an AI agent can request captures without custom browser glue. Plans include 1,000 shots per month free with no card; Starter is $5 for 3,000, Growth $15 for 15,000, Pro $39 for 60,000, Scale $99 for 250,000 and Business $249 for 1,000,000. Yearly billing gives two months free, and every feature is on every plan. Create a free ScreenshotNeo account to start with the 1,000 monthly shots.

Final implementation checklist

  • Policy, authorization, robots.txt and retention decisions are recorded.
  • Scheduler, downloader, parser, pipeline, storage and orchestration have separate responsibilities.
  • Delays, concurrency and retries reflect observed server responses.
  • Raw evidence and provenance are retained where lawful.
  • Required fields, duplicates, freshness and parser drift can fail a run.
  • Browser rendering is a targeted fallback, not the default path.
  • Every recurring run is idempotent, observable and recoverable.

Frequently Asked Questions

Should I store raw HTML as well as cleaned rows?

Store raw responses or snapshots when lawful and useful for replay, then write cleaned records separately. Apply an explicit retention policy and protect access because raw pages may contain sensitive data.

How do I decide between a hosted API and Scrapy?

Use Scrapy when you need direct control over crawling, code and data handling. Consider a hosted API when the team prefers API-key calls, asynchronous runs, exports and schedules without operating crawler infrastructure; evaluate rendering, residency and lock-in for the specific provider.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Is Airflow a web scraper?

No. Airflow orchestrates scraper tasks and downstream ETL/ELT, while the crawler still handles requests, parsing, politeness and persistence.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Read next

Recommended PC Tool
Recommended PC Tool
PC Slower Than It Used to Be?Free scan - under a minute
Outdated Drivers Are Slowing You DownFree scan - exact matches

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.