Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

For a small, bounded crawl, use Python’s requests library to fetch pages and Beautiful Soup to parse them; keep a queue of in-scope URLs and stop at a page or depth limit. For a larger spider that needs request scheduling, callbacks and retries, use Scrapy. Before either approach, check for an official API or export, read the site’s crawl instructions, and set a conservative request rate.

Choose the right Python crawling approach

A crawler fetches a starting page, extracts information, and may follow selected links to discover more pages. That differs from scraping a single known page: crawling adds URL discovery, scope rules, deduplication and stopping conditions.

Need Good starting point Why
A small, one-off crawl with a few fields requests and Beautiful Soup You can see and control the fetch, parsing, queue and limits directly.
A multi-page spider with scheduling, callbacks and project settings Scrapy Spiders yield requests; Scrapy’s downloader fetches them and sends responses to callbacks for extraction or further requests. See Scrapy’s request and response documentation.
Pages whose useful content appears only after JavaScript runs A browser-rendering component, if direct HTTP fetching is insufficient Ordinary HTTP clients receive server responses; they do not execute page JavaScript. Scrapy’s ecosystem lists browser-rendering integrations, but not every site needs one. See the Scrapy project overview.

Before crawling HTML, look for an official API, bulk export or search endpoint. Scrapy’s optimization guide notes that documented interfaces can be faster for your crawler and cheaper for the site than fetching pages individually: Scrapy optimization guidance.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Plan the crawl before sending requests

  1. Define the purpose and fields. Decide what you need to collect, such as page titles and canonical URLs, and avoid fetching or retaining unrelated data.
  2. Set scope. List the allowed hostnames and paths. Choose a maximum page count or link depth; do not treat every discovered link as permission to crawl another site.
  3. Check for a supported interface. Prefer a published API or export where it supplies the information you need.
  4. Read robots.txt and relevant site terms. Use the published crawl guidance to shape the run. Robots instructions are not access authorization: RFC 9309 says, “These rules are not a form of access authorization.” Read the IETF Robots Exclusion Protocol specification.
  5. Test a small sample. Inspect status codes and content types, verify the extracted fields, and only then decide whether to expand the crawl.
  6. Set a conservative rate. Limit concurrency and add per-domain delay; monitor failures, retries, response times and signs of throttling.
  7. Record progress. Keep enough metadata, such as visited URLs and outcomes, to diagnose failures and resume without needlessly refetching everything.

Robots.txt is also not a way to keep a URL out of Google’s index. Google explains that a blocked URL may still be indexed if discovered elsewhere; use appropriate access controls or a noindex directive for indexing goals rather than relying on a crawl block. See Google’s robots.txt guide.

Build a small bounded crawler with Python

This example crawls only pages on one hostname, visits at most 20 pages, follows links only one level from the start page, and waits between requests. It extracts titles and links; change the extraction logic to match the site and data you are authorized to collect. The limits are deliberately modest, not a claim that a particular rate is appropriate for every website.

Install the dependencies

Use Python 3 and install the two packages in your active environment:

python -m pip install requests beautifulsoup4

Save and run the crawler

Save as crawl.py, replace the example URL with a site you are permitted to crawl, and run python crawl.py.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
from collections import deque
from time import sleep
from urllib.parse import urljoin, urldefrag, urlparse

import requests
from bs4 import BeautifulSoup

START_URL = "https://example.com/"
MAX_PAGES = 20
MAX_DEPTH = 1
DELAY_SECONDS = 2

start = urldefrag(START_URL)[0]
allowed_host = urlparse(start).netloc.lower()
queue = deque([(start, 0)])
seen = {start}
session = requests.Session()
session.headers.update({"User-Agent": "ExampleResearchCrawler/1.0 (contact: [email protected])"})

while queue and len(seen) <= MAX_PAGES:
    url, depth = queue.popleft()
    try:
        response = session.get(url, timeout=(5, 20))
        print(response.status_code, response.headers.get("Content-Type", ""), url)
        response.raise_for_status()

        content_type = response.headers.get("Content-Type", "").lower()
        if "text/html" not in content_type:
            continue

        soup = BeautifulSoup(response.text, "html.parser")
        title = soup.title.get_text(" ", strip=True) if soup.title else ""
        print({"url": response.url, "title": title})

        if depth >= MAX_DEPTH:
            continue

        for link in soup.select("a[href]"):
            absolute = urldefrag(urljoin(response.url, link["href"]))[0]
            parsed = urlparse(absolute)
            if parsed.scheme not in ("http", "https"):
                continue
            if parsed.netloc.lower() != allowed_host:
                continue
            if absolute not in seen and len(seen) < MAX_PAGES:
                seen.add(absolute)
                queue.append((absolute, depth + 1))
    except requests.RequestException as exc:
        print(f"Request failed for {url}: {exc}")
    finally:
        sleep(DELAY_SECONDS)

session.close()

Understand and adjust the controls

  • allowed_host prevents following external links. If a site uses a separate subdomain for relevant pages, add that hostname explicitly rather than accepting every host.
  • MAX_PAGES bounds the number of unique URLs added to the queue. The start URL counts toward the bound.
  • MAX_DEPTH defines how many link-following steps are allowed from the start page. With a value of 1, the crawler fetches the start page and its directly linked pages, but does not follow links from those pages.
  • DELAY_SECONDS inserts a pause after each loop iteration. For a single-worker example this is a simple throttle, not a universal safe rate; follow site guidance and lower the rate or stop if the server signals stress.
  • The request timeout is a connect/read timeout pair. It prevents a single slow request from hanging the run indefinitely; it does not guarantee that the page is complete.
  • The content-type check avoids trying to parse PDFs or images as HTML. Add explicit handling only for content types you actually need.

This is an instructional skeleton, not a production crawler. It keeps its queue and results in memory, does not implement persistent resume, and does not itself parse robots.txt or retry transient failures. For a one-off small task, those constraints may be acceptable; for recurring or larger crawls, use a framework and persist crawl state deliberately.

Use Scrapy when you need a spider framework

Scrapy formalizes the request/response cycle and lets a spider yield requests for other pages while callbacks process returned responses. Install it with python -m pip install scrapy, create a project with scrapy startproject sitecrawler, then add a spider file such as sitecrawler/spiders/pages.py.

import scrapy

class PagesSpider(scrapy.Spider):
    name = "pages"
    allowed_domains = ["example.com"]
    start_urls = ["https://example.com/"]

    custom_settings = {
        "ROBOTSTXT_OBEY": True,
        "DOWNLOAD_DELAY": 2,
        "CONCURRENT_REQUESTS_PER_DOMAIN": 1,
        "CLOSESPIDER_PAGECOUNT": 20,
    }

    def parse(self, response):
        yield {
            "url": response.url,
            "title": response.css("title::text").get(default="").strip(),
        }
        for href in response.css("a::attr(href)").getall():
            yield response.follow(href, callback=self.parse)

Run it from the project directory with scrapy crawl pages -O pages.jsonl. The output option writes scraped items to a JSON Lines file. Review the spider’s selectors, allowed domains, page-count bound and settings before running against a real site. The example follows in-domain links until the page-count limit closes the spider; for a narrower path or depth policy, add explicit URL checks and depth handling rather than assuming every in-domain route belongs in scope.

Robots directives and Scrapy settings

ROBOTSTXT_OBEY enables Scrapy’s robots middleware behavior, but do not assume this enforces every rate directive. Scrapy’s optimization documentation says it does not automatically act on robots.txt Crawl-delay and Request-rate directives. Translate relevant guidance into settings such as DOWNLOAD_DELAY and per-domain concurrency, and monitor the target’s responses: Scrapy optimization documentation.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Handle JavaScript-rendered pages selectively

First inspect the HTML response from a direct request. If the required content is already in the response, a browser is unnecessary. If the response contains only a shell and the desired data is loaded later, check whether the site exposes an API intended for that data. Only add browser rendering when direct retrieval of an appropriate endpoint is not viable and the crawl is within the site’s rules.

Browser rendering costs more resources than fetching HTML directly and can add complexity around waits, cookies and dynamic state. Scrapy’s project ecosystem lists browser integrations, but their existence does not establish that any one integration is required or appropriate for a particular site. Keep the crawl bounded and validate that the browser is capturing the intended content rather than merely waiting longer.

Common failures and practical fixes

Symptom Likely cause What to do
401 or 403 response The resource requires authorization or the server denies the request. Stop and confirm that you have permission and the intended authentication method. Do not try to bypass access controls.
429 response or increasing 5xx errors The server is throttling requests or under load. Pause or stop; reduce concurrency and rate, and follow any published retry guidance. Do not respond by increasing request volume.
200 response but empty extraction The selector does not match, markup changed, or content is populated by JavaScript. Inspect the returned HTML and content type; test selectors on a small sample and determine whether an official endpoint or browser rendering is appropriate.
Repeated URLs or a crawl that grows unexpectedly Query-string variants, calendars, filters or URL fragments create many apparent links. Deduplicate normalized URLs, remove fragments, restrict paths, and enforce page/depth limits. Strip query parameters only when you know they do not distinguish content.
Timeouts Slow server, unstable connection or an overly broad workload. Keep timeouts finite, reduce load, record failed URLs and retry selectively rather than looping immediately on every failure.
Scrapy keeps visiting more pages than expected The spider follows all matching in-domain links, including routes outside the intended section. Tighten allowed_domains and add path, depth or page-count constraints; test on a small sample.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Make runs easier to resume and maintain

For repeated crawls, store each result with its URL and a crawl timestamp, and record status or error details separately from extracted fields. Persist visited URLs or completed work so a restarted run does not begin blindly. Keep a small representative sample of expected fields and check it after site changes; a successful HTTP response does not prove an extraction is still correct.

Keep extraction narrow, set a maximum workload, and watch response statuses, retries and latency as you change delay or concurrency. Scrapy’s optimization guidance discusses measuring crawl performance and tuning carefully rather than treating maximum throughput as the goal. A faster crawl is not automatically a better one if it increases server load or causes throttling.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Or skip the browser setup

ScreenshotNeo is a screenshot API, not a link-following web crawler: use it when your goal is a visual capture of a page, rather than extracting records across a site. It can return an image or PDF from one request. For a Python screenshot request, see the ScreenshotNeo documentation:

import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
open("shot.webp", "wb").write(r.content)

ScreenshotNeo says it accepts cookie or consent banners as a visitor and removes more than 60 known consent platforms, newsletter popups and chat widgets before capture; each step can be turned off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads and cache hits cost nothing, and response headers indicate the page verdict and billing status. Its MCP server provides take_screenshot, get_page_info and capture_pdf for AI agents using Claude, Cursor or another MCP client. The Free plan includes 1,000 shots per month with no card; paid plans start at $5 for 3,000 shots. See ScreenshotNeo for the product details. Sign up for 1,000 free screenshots a month, with no card required.

Further reading

For a longer book-length treatment, O’Reilly lists Ryan Mitchell’s Web Scraping with Python, 3rd Edition as published in February 2024, with coverage including requests, HTML parsing, Scrapy, JavaScript pages, APIs and data handling: publisher page. It is optional; a bounded crawl can be built with the basic workflow above.

Frequently Asked Questions

Does reading robots.txt mean I have permission to crawl a site?

No. RFC 9309 explicitly distinguishes robots instructions from access authorization. Permission and applicable terms depend on the site and circumstances.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Can the simple Python example crawl JavaScript-rendered content?

Not when the needed content is absent from the server-returned HTML. Check for an appropriate official endpoint first; otherwise a browser-rendering approach may be necessary.

Should I use Scrapy Cloud to run a crawler?

Only after you have a working, permitted crawl and a reason to deploy or schedule it. Hosting is optional; the local Scrapy example can be run without it.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.