Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Web crawling discovers and retrieves pages; web scraping extracts specific data from them. A crawler may follow links across a site for indexing or inventory. A scraper may request one page, feed or API endpoint and save only fields such as a product name, price or publication date. One system can do both, but the permission, engineering and legal questions are different.

This guide explains robots.txt, responsible collection, legal boundaries, implementation choices, failure handling and when an API or managed service is safer than brittle HTML extraction.

What is the difference between web scraping and web crawling?

Web crawling is discovery and retrieval

A crawler is an automated client that discovers URLs and requests resources, commonly by following links recursively. Search engines use crawling to find pages for indexing. A crawler’s primary questions are: Which URLs exist? Which links should be visited next? How often can this host be requested?

Web scraping is focused extraction

A scraper takes a known page, feed or API response and extracts selected fields. It might collect a page title, article body, table rows or structured metadata without traversing the rest of the site. Its primary questions are: Which fields are necessary? How should they be normalized? How will changes to the page be detected?

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

They often appear together

A news-monitoring job can crawl category pages to discover article URLs, then scrape each article’s title and timestamp. A product catalog can crawl pagination and scrape the SKU and price from every product page. Keep the discovery queue, extraction rules and data-retention policy separate so that a change in one does not silently expand the other.

What does robots.txt do?

Under the Robots Exclusion Protocol described in RFC 9309 (IETF, 2022), a site’s top-level /robots.txt publishes requested behavior for automated clients. Your crawler should select the matching user-agent group and apply the most specific Allow or Disallow path rule. The standard’s important qualification is that “These rules are not a form of access authorization.”

Scope matters

Rules apply to the relevant host, protocol and port. A file on https://example.com/robots.txt does not automatically govern https://shop.example.com/, another protocol, or a different port. Fetch and evaluate each origin you intend to visit.

Fetch failures need a conservative policy

Record the response status and retrieval time. Treat an unavailable or unreachable robots file as a reason to slow down and avoid expanding the crawl until you can make a deliberate decision. RFC 9309 recommends conservative caching, generally no more than 24 hours unless the server is unreachable; do not keep using an old policy indefinitely.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Robots.txt is not a privacy or indexing control

Google describes robots.txt as a way to manage crawler access and traffic, not as a reliable method for keeping a URL out of search results. Site owners who need exclusion should use noindex or authentication. For a collector, a disallow rule is still a strong signal to honor, but it does not grant permission to access a private area.

Is web scraping legal?

There is no universal yes-or-no answer. The result depends on jurisdiction, the date, whether the material is public or authenticated, the site’s terms and notices, the data collected, and what your system actually does. Public visibility does not erase contract, copyright, privacy, trespass, misappropriation, unjust-enrichment or conversion issues.

The Ninth Circuit’s 2022 hiQ Labs v. LinkedIn opinion concerned a preliminary injunction and public LinkedIn profiles. On that record it treated access to publicly available pages as unlikely to be “without authorization” under the CFAA, but it did not create a general scraping license. That decision should not be generalized to password-protected pages, bypassed controls, every type of data or every country.

Questions to answer before collecting

  • What is the purpose, and which exact fields are necessary?
  • Which countries’ privacy and data-protection rules apply to you and to the people represented in the data?
  • Is there an API, export or permissioned feed you can use instead?
  • Are you crossing a login, paywall, CAPTCHA, bot check or other technical control? Do not bypass it.
  • Do the terms, notices or an explicit owner request prohibit or limit automated collection?
  • How long will you retain the data, who can access it, and how will deletion or correction requests be handled?

For a commercial or sensitive project, obtain advice for the relevant jurisdiction rather than treating a public URL as a legal conclusion.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How do I scrape a website responsibly?

  1. Define the job. Write down the purpose, URL scope, fields, geography, retention period and lawful basis. A narrow field list reduces load and risk.
  2. Prefer permissioned sources. Check for an official API, data export or feed. These interfaces are normally more stable and easier to govern than HTML selectors.
  3. Inspect robots.txt. Fetch the target host’s file, record the exact content and timestamp, parse the user-agent group, and enforce the most specific rule before queuing a URL.
  4. Read site boundaries. Review terms, notices, authentication requirements and opt-out instructions. Never bypass a login, paywall, CAPTCHA or technical access control.
  5. Identify your client. Use a stable user-agent string and, where appropriate, a contact address so an operator can reach you.
  6. Control traffic. Use low concurrency, a delay or token bucket, exponential backoff, caching, conditional requests and a kill switch. Stop or reduce scope after repeated 403, 429 or 5xx responses.
  7. Minimize and protect data. Extract only necessary fields, restrict access, encrypt sensitive stores, retain source URLs and timestamps, and support deletion or correction where applicable.
  8. Validate and audit. Test parsers against layout changes, monitor status and extraction error rates, keep representative fixtures, and log permissions, robots decisions, requests, responses and shutdown events.

Which collection approach fits the job?

Decision axis Prefer this Why and what to watch
API versus HTML extraction Official API or export Documented fields and versioning reduce selector breakage. Check quotas, authentication, license terms and pagination.
Public versus authenticated data Public, permissioned endpoints Authenticated data carries additional access, contract and privacy obligations. Do not share credentials or defeat controls.
One-off research versus recurring production crawl One-off: small script; recurring: scheduled, observable pipeline Production jobs need queues, retries, rate limits, alerting, retention rules and a kill switch.
Static HTML versus JavaScript-rendered pages Static request when content is in the response; a permitted browser renderer otherwise Rendering costs more time and resources. Wait for a specific selector rather than an arbitrary long delay, and do not use rendering to evade bot checks.
Self-hosted versus managed infrastructure Self-hosted for maximum control; managed for operational features Compare permission controls, stability, cost, observability, rate control, data protection and maintenance burden.

A small, responsible Python crawler and scraper

The following example is intentionally narrow: it checks robots.txt, identifies itself, limits requests to one host, waits between requests, stops on repeated server errors, and extracts the page title and first heading. Install the dependencies with python -m pip install requests beautifulsoup4. Replace the example URL only with a page you are permitted to access.

import time
from urllib.parse import urljoin, urldefrag, urlparse
from urllib.robotparser import RobotFileParser

import requests
from bs4 import BeautifulSoup

START_URL = "https://example.com/"
USER_AGENT = "Freedom251ResearchBot/1.0 (+https://example.com/contact)"
DELAY_SECONDS = 2.0
MAX_PAGES = 20

session = requests.Session()
session.headers.update({"User-Agent": USER_AGENT, "Accept": "text/html"})
origin = f"{urlparse(START_URL).scheme}://{urlparse(START_URL).netloc}"
robots_url = urljoin(origin, "/robots.txt")
robots = RobotFileParser()
robots.set_url(robots_url)
try:
    robots.read()
except Exception as exc:
    raise RuntimeError(f"Cannot establish a current robots policy: {exc}")

queue = [START_URL]
seen = set()
records = []
server_errors = 0

while queue and len(records) < MAX_PAGES:
    url, _ = urldefrag(queue.pop(0))
    parsed = urlparse(url)
    if url in seen or parsed.netloc != urlparse(START_URL).netloc:
        continue
    seen.add(url)
    if not robots.can_fetch(USER_AGENT, url):
        continue

    time.sleep(DELAY_SECONDS)
    try:
        response = session.get(url, timeout=20)
    except requests.RequestException:
        continue

    if response.status_code in (403, 429) or response.status_code >= 500:
        server_errors += 1
        if server_errors >= 3:
            break
        continue
    if response.status_code != 200 or "text/html" not in response.headers.get("Content-Type", ""):
        continue
    server_errors = 0

    soup = BeautifulSoup(response.text, "html.parser")
    records.append({
        "url": url,
        "title": soup.title.get_text(" ", strip=True) if soup.title else None,
        "h1": soup.find("h1").get_text(" ", strip=True) if soup.find("h1") else None,
        "retrieved_at": time.strftime("%Y-%m-%dT%H:%M:%SZ", time.gmtime()),
    })

    for link in soup.select("a[href]"):
        candidate = urljoin(url, link["href"])
        if urlparse(candidate).netloc == parsed.netloc:
            queue.append(candidate)

print(records)

What this example deliberately does not do

  • It does not log in, solve a CAPTCHA, circumvent a paywall or spoof a browser to defeat a control.
  • It does not assume that a robots fetch error means permission to continue.
  • It does not download every resource, retain full HTML forever or collect unrelated personal data.
  • It is not a production scheduler. Add durable queues, conditional requests, structured logs, secret management and monitoring before operating at recurring scale.

Common failures and how to recover

403 Forbidden

The owner may prohibit automated access, require authentication or have detected an unusual rate. Stop, review permission and terms, lower scope only if permitted, and contact the operator. Do not rotate identities or bypass the restriction.

429 Too Many Requests

Honor any Retry-After value, reduce concurrency, increase delay and cache successful responses. Repeated 429 responses are a stop signal, not an invitation to add more proxies.

5xx responses or timeouts

Use bounded retries with exponential backoff and jitter, then stop after a threshold. Cache what succeeded and preserve the failing URL, status and timestamp for diagnosis.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Empty content from a JavaScript page

The requested HTML may contain only an application shell. First look for a documented API or embedded structured data. If browser rendering is permitted, wait for a meaningful selector or network-idle condition, set a maximum wait, and record that the page required rendering.

Parser suddenly returns null fields

Assume the layout changed or an alternate template was served. Keep raw samples under an appropriate retention policy, validate required selectors, alert on extraction-rate drops and version your parser. Do not silently publish empty values.

Robots rules appear contradictory

Confirm the host, protocol, port and user-agent group. Apply the most specific matching path rule, record the decision and re-fetch when the policy is stale. A subdomain’s file may differ from the parent domain’s file.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server for developers. It accepts a URL and returns a PNG, JPEG, WebP or PDF. Before capture, it can accept the cookie or consent banner like a visitor and remove more than 60 known consent platforms, newsletter popups and chat widgets; each cleanup step can be turned off. Only clean shots are billed: bot checks or CAPTCHAs, blank pages, timeouts, failed loads and cache hits cost nothing, and every response identifies the result with X-Page-Verdict and X-Billed headers.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use the same request from a shell:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

See the ScreenshotNeo API documentation for request parameters. Options include full-page capture with lazy images loaded, a CSS-selector element capture, dark mode, 12 device presets or a custom viewport, retina scale, PDF paper size/margins/landscape/page ranges, HTML/CSS-to-image, custom CSS and JavaScript, pre-capture clicks, hidden selectors, waits for a selector/delay/network idle, blocking ads/trackers/requests/resource types, custom headers/cookies/user agent/Authorization, timezone and geolocation, transparent backgrounds, image resizing, selectable-TTL caching, signed public image links, asynchronous jobs with signed webhooks, bulk capture of up to 100 URLs per call, a usage API and an OpenAPI specification. Parameter names used by other screenshot APIs also work for easier migration. An MCP server exposes take_screenshot, get_page_info and capture_pdf to Claude, Cursor and other MCP clients.

Plan Included shots Price
Free 1,000 per month $0, no card
Starter 3,000 $5
Growth 15,000 $15
Pro 60,000 $39
Scale 250,000 $99
Business 1,000,000 $249

Yearly billing gives two months free, and every feature is available on every plan. If you want clean screenshots without maintaining browser orchestration, start with 1,000 free screenshots a month—no card required.

Performance, reliability and cost controls

  • Bound the crawl: cap URLs, depth, response size and total runtime; use a kill switch.
  • Reduce duplicate work: canonicalize URLs, remove fragments, cache responses and use conditional requests where the server supports them.
  • Protect the origin: keep concurrency low, honor backoff signals and schedule large jobs during an agreed window.
  • Measure useful outcomes: track allowed versus skipped URLs, status classes, latency, bytes, parser success and policy decisions.
  • Budget rendering: JavaScript browsers consume more CPU, memory and time than direct HTTP requests. Render only pages that need it.
  • Plan for change: version selectors and schemas, retain enough evidence to debug a failure, and alert before bad data reaches downstream users.

FAQ

Can one project crawl and scrape at the same time?

Yes. Treat link discovery as a bounded queue and extraction as a separate stage with its own fields, validation and retention rules. This prevents a newly discovered link pattern from silently changing what you store.

Should I keep the complete HTML response?

Only when you have a documented operational or legal reason and an appropriate retention policy. For many jobs, storing the extracted fields, source URL, retrieval time and a small diagnostic sample is enough.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What is the safest default when a site owner objects?

Pause the affected job, preserve the request and policy logs, review the scope and permission, and honor a clear opt-out or stop request while you resolve the issue.

Frequently Asked Questions

Can one project crawl and scrape at the same time?

Yes. Treat link discovery as a bounded queue and extraction as a separate stage with its own fields, validation and retention rules.

Should I keep the complete HTML response?

Only when you have a documented operational or legal reason and an appropriate retention policy. Often the extracted fields, source URL, retrieval time and a diagnostic sample are sufficient.

What is the safest default when a site owner objects?

Pause the affected job, preserve request and policy logs, review permission and scope, and honor a clear opt-out or stop request while resolving the issue.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.