Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The most dependable web-scraping template is a small pipeline you adapt to one site: configure a URL and selectors, check the site’s rules, fetch HTML, parse fields, validate records, and save structured output. A template is a starting structure—not a universal scraper. Markup, permissions, cookies, rendering, and rate limits differ from site to site.

A practical scraping workflow

Use this sequence for a one-off extraction or as the foundation of a larger crawler:

  1. Configure: define the target URL, headers, selectors, output path, and conservative pacing.
  2. Check the site: inspect the correct origin’s robots.txt, terms, and any official API or developer documentation. Prefer an official API when it is available and appropriate.
  3. Fetch: handle network errors, redirects, and HTTP status codes explicitly.
  4. Parse: extract named fields with a parser and selectors.
  5. Validate: flag missing fields, malformed values, duplicates, and unexpected markup changes.
  6. Save and log: write JSON or CSV and retain enough context to diagnose a failed run.

Before automating, confirm that your use is permitted under the target site’s terms and applicable rules. Stop or request permission when access is restricted. This article explains implementation; it does not determine whether a particular use is lawful in your jurisdiction.

How do I scrape a website with Python?

For content present in the initial HTML response, Python’s requests and Beautiful Soup are a clear starting point. Install them with python -m pip install requests beautifulsoup4.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A reusable template

from __future__ import annotations

import csv
import logging
import time
from dataclasses import dataclass, asdict
from pathlib import Path
from typing import Iterable
from urllib.parse import urljoin

import requests
from bs4 import BeautifulSoup

@dataclass
class Record:
    title: str
    url: str
    price: str

TARGET_URL = "https://example.com/products"
OUTPUT = Path("products.csv")
SELECTORS = {
    "item": ".product-card",
    "title": ".product-title",
    "url": "a.product-link",
    "price": ".price",
}
HEADERS = {"User-Agent": "ResearchBot/1.0 (contact: [email protected])"}
TIMEOUT = 30
DELAY_SECONDS = 2.0

logging.basicConfig(level=logging.INFO, format="%(levelname)s %(message)s")


def fetch(url: str) -> str:
    try:
        response = requests.get(url, headers=HEADERS, timeout=TIMEOUT,
                                allow_redirects=True)
        response.raise_for_status()
    except requests.RequestException as exc:
        raise RuntimeError(f"request failed for {url}: {exc}") from exc
    return response.text


def text_or_empty(node) -> str:
    return node.get_text(" ", strip=True) if node else ""


def parse(html: str, base_url: str) -> list[Record]:
    soup = BeautifulSoup(html, "html.parser")
    rows: list[Record] = []
    for card in soup.select(SELECTORS["item"]):
        link = card.select_one(SELECTORS["url"])
        item = Record(
            title=text_or_empty(card.select_one(SELECTORS["title"])),
            url=urljoin(base_url, link.get("href", "")) if link else "",
            price=text_or_empty(card.select_one(SELECTORS["price"])),
        )
        rows.append(item)
    return rows


def validate(rows: Iterable[Record]) -> list[Record]:
    valid: list[Record] = []
    seen: set[str] = set()
    for row in rows:
        if not row.title or not row.url:
            logging.warning("missing required field: %s", row)
            continue
        if row.url in seen:
            logging.warning("duplicate URL skipped: %s", row.url)
            continue
        seen.add(row.url)
        valid.append(row)
    return valid


def save(rows: list[Record], path: Path) -> None:
    with path.open("w", newline="", encoding="utf-8") as handle:
        writer = csv.DictWriter(handle, fieldnames=["title", "url", "price"])
        writer.writeheader()
        writer.writerows(asdict(row) for row in rows)


def main() -> None:
    html = fetch(TARGET_URL)
    rows = validate(parse(html, TARGET_URL))
    save(rows, OUTPUT)
    logging.info("saved %d records to %s", len(rows), OUTPUT)


if __name__ == "__main__":
    main()

Adapt the selectors, not the whole program

Replace TARGET_URL and the four CSS selectors after inspecting the page. Keep selectors anchored to stable attributes such as a documented class, data-* attribute, or semantic element. Avoid selectors based on a fragile chain of generated class names. Test with a saved HTML fixture so a site redesign produces a visible failure instead of silently corrupting data.

Pagination and pacing

For pages numbered with a predictable query parameter, fetch one page, parse it, validate it, then sleep before the next request. Stop when the next link is absent or when a maximum page count is reached. Do not use concurrency to evade restrictions. Keep a request log containing URL, timestamp, status, record count, and parser version.

How do I make a web-scraper template?

Separate site-specific configuration from reusable control flow. A useful configuration records:

  • the exact origin and URL pattern;
  • selectors for each field and the item container;
  • pagination rules;
  • request headers that are truthful and appropriate;
  • timeout, retry, and delay policy;
  • validation rules and an output schema.

Write a small fixture test that checks expected fields and a realistic minimum record count. When the page changes, the test should fail. Keep raw HTML for failed pages where the site’s terms permit retention, and redact credentials or personal data from logs.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Robots.txt is guidance, not access control

Google says crawler instructions in robots.txt cannot enforce crawler behavior, and a disallowed URL may still be indexed when linked elsewhere (Google’s robots.txt introduction). Do not use the file to protect private data.

Google’s documented scope is the host, protocol, and port where the file is served; a subdomain’s file does not automatically govern the parent domain. Google documents UTF-8 plain text, a 500 KiB size limit, and no support for crawl-delay in its crawler behavior (robots.txt specification). Read the file for technical guidance, then check terms and other instructions separately.

Honoring robots rules in Scrapy

For repeated crawling, Scrapy provides scheduling, retries, item pipelines, and downloader middleware. Its robots middleware filters forbidden requests when enabled with ROBOTSTXT_OBEY = True; the documentation identifies Protego as the default parser (Scrapy downloader middleware).

# settings.py
ROBOTSTXT_OBEY = True
DOWNLOAD_DELAY = 2
CONCURRENT_REQUESTS_PER_DOMAIN = 1

These settings help implement a policy; they do not grant permission or make a restricted use acceptable.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Should I use Scrapy or Playwright?

Choose according to the page and job, not a blanket “best scraper” claim. No head-to-head speed, cost, or reliability result establishes a universal winner.

Situation Better starting point Reason
Data is in the initial response; one or a few pages Requests plus a parser Small dependency footprint and direct control
Many URLs, scheduling, retries, pipelines, policy middleware Scrapy Built-in crawling architecture; robots filtering can be enabled
Content appears after JavaScript, clicks, scrolling, or browser-only requests Playwright Runs a browser and exposes interaction and network events

When browser automation is necessary

Use Playwright when the required data is not present in the initial HTML or the workflow depends on rendered interaction. Its Python Request API exposes request, response, completion, and failure events (Playwright Request API). A completed HTTP request is not automatically a successful application response: statuses such as 404 and 503 still complete as HTTP responses, so inspect the status.

from playwright.sync_api import sync_playwright

with sync_playwright() as p:
    browser = p.chromium.launch()
    page = browser.new_page()
    response = page.goto("https://example.com", wait_until="networkidle")
    if response is None or response.status >= 400:
        raise RuntimeError(f"page load failed: {response.status if response else 'no response'}")
    title = page.locator("h1").inner_text()
    print(title)
    browser.close()

Browser runs add installation, memory, startup time, and rendering failure modes. Keep selectors and waits explicit, and capture diagnostics such as the final URL and status.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Fetching, validation, and failure handling

Transport and HTTP errors

  • DNS, TLS, or timeout: verify the URL, network, and timeout; retry only transient failures with bounded backoff.
  • 3xx redirect: record the final URL and confirm it remains an allowed target.
  • 401 or 403: do not attempt to bypass access controls; use an official API or request access.
  • 404: remove stale URLs or update discovery logic.
  • 429: slow down, respect any stated limit, and stop if instructed.
  • 5xx: treat as a server failure; retry sparingly and preserve the response context.

Empty or suspicious output

A 200 response can contain a consent wall, bot check, login page, or an empty shell. Validate content, not just status: require key fields, compare counts with a known fixture, and flag sudden zero-record runs. Never treat a successful fetch as proof that extraction is correct or permitted.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Markup drift

When selectors return nothing, save the response, inspect whether the content moved into an iframe or JavaScript-rendered request, and update the configuration. Add a parser version to output metadata so downstream users can identify when a schema changed.

Or skip the browser setup

If you need a clean screenshot or rendered capture while documenting a scraping workflow, ScreenshotNeo provides a GET-based website screenshot API and MCP server. It accepts cookie or consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be turned off. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing status. Its MCP tools—take_screenshot, get_page_info, and capture_pdf—work with Claude, Cursor, and other MCP clients.

Use the API with the documentation:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

The Free plan includes 1,000 shots per month with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.

Performance, reliability, and cost decisions

  • Measure records per request, error rate, and validation failures rather than assuming a faster run is better.
  • Cache responses where permitted, use conditional requests when supported, and avoid re-fetching unchanged pages.
  • Use a queue and bounded concurrency for large jobs; keep per-domain pacing conservative.
  • Store raw inputs, parser version, timestamps, and status so results can be reproduced.
  • Estimate browser capacity separately from HTTP capacity because browser processes consume substantially more resources.

Frequently Asked Questions

Can I scrape any page listed in robots.txt?

No. robots.txt is crawler guidance, not a grant of permission. Check the site’s terms, technical instructions, applicable rules, and whether an official API is available.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Why did my scraper get a 200 response but no records?

The response may be a consent wall, bot check, login page, JavaScript shell, or changed markup. Save and inspect the HTML, validate required fields, and use browser automation only when the content genuinely requires rendering.

When should I move from a script to Scrapy?

Move when you need repeatable multi-URL scheduling, retries, downloader middleware, item pipelines, and centralized policy settings. Keep a simple script for small jobs whose data is already in HTML.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.