Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

A scalable Python crawler is a pipeline, not just a fast HTTP loop: it needs a scoped URL frontier, bounded fetching, respectful per-host scheduling, parsing, duplicate control, durable storage, and operational visibility. For a maintainable crawl, Scrapy is a practical starting point; a small asyncio client is useful when you deliberately want to own those pieces. Neither framework choice nor more concurrency guarantees more useful throughput.

How a crawler works

A crawler repeatedly fetches pages, extracts useful records and candidate links, then schedules eligible links for later requests. Keeping those jobs distinct makes the system easier to reason about as the URL set grows.

  1. Scope and seeds: Decide which hosts, paths, content types, and crawl depths are in bounds. Start from known URLs rather than an unbounded search.
  2. Frontier: Queue URLs with scheduling state such as queued, fetched, failed, or deferred. Normalize and deduplicate before enqueueing, and persist the frontier if a crawl must survive a process restart.
  3. Fetcher: Reuse connections, use explicit timeouts, cap response sizes, validate schemes and redirects, and bound concurrency. Apply request limits by host, not just globally.
  4. Politeness and robots: Identify the crawler, read and obey robots.txt rules, pace requests, and back off on errors or signs of blocking.
  5. Parser and link policy: Extract records and links, normalize links cautiously, then filter candidates against the crawl scope and content-type policy.
  6. Storage and observability: Persist extracted records and enough crawl state to resume, troubleshoot, and avoid reprocessing work.

This pipeline works in one process first. Scaling should change a measured bottleneck, not merely add workers.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Start with a bounded Scrapy spider

Scrapy provides a crawler framework, scheduler, request deduplication, settings, and spider conventions. It is a credible default when you need structured extraction and a crawl that can be maintained beyond a short script. The example below starts at example.org, follows links on that host, and exports extracted page data as JSON Lines.

Install and save the spider

Install Scrapy in a virtual environment with python -m pip install scrapy. Save this as crawler.py:

import scrapy


class ExampleSpider(scrapy.Spider):
    name = "example"
    allowed_domains = ["example.org"]
    start_urls = ["https://example.org/"]

    custom_settings = {
        # Keep robots compliance on; use a crawler identity you control.
        "ROBOTSTXT_OBEY": True,
        "USER_AGENT": "ExampleResearchBot/1.0 (+https://example.org/bot)",
        # These limits apply to this crawler instance.
        "CONCURRENT_REQUESTS": 8,
        "CONCURRENT_REQUESTS_PER_DOMAIN": 2,
        "DOWNLOAD_DELAY": 1.0,
        "AUTOTHROTTLE_ENABLED": True,
        "AUTOTHROTTLE_START_DELAY": 1.0,
        "AUTOTHROTTLE_MAX_DELAY": 30.0,
        "AUTOTHROTTLE_TARGET_CONCURRENCY": 1.0,
        "DOWNLOAD_TIMEOUT": 20,
        "DOWNLOAD_MAXSIZE": 10 * 1024 * 1024,
        "DEPTH_LIMIT": 3,
        "RETRY_TIMES": 2,
    }

    def parse(self, response):
        title = response.xpath("normalize-space(string(//title))").get()
        description = response.xpath(
            "normalize-space(string(//meta[translate(@name,"
            " 'ABCDEFGHIJKLMNOPQRSTUVWXYZ', 'abcdefghijklmnopqrstuvwxyz')='description']"
            "/@content))"
        ).get()

        yield {
            "url": response.url,
            "status": response.status,
            "title": title,
            "description": description,
        }

        for href in response.css("a::attr(href)").getall():
            # response.follow resolves relative URLs; allowed_domains and the
            # scheduler keep the crawl constrained and suppress duplicate requests.
            yield response.follow(href, callback=self.parse)

Run it and check the output

From the directory containing crawler.py, run scrapy runspider crawler.py -O pages.jsonl. The spider fetches the start URL and eligible linked pages, then writes one JSON object per line. Change both allowed_domains and start_urls to your intended host before crawling it. Replace the sample User-Agent with a truthful, contactable crawler identity appropriate to your use.

The example deliberately has conservative limits and a shallow depth ceiling. They are starting controls, not universal safe values: a site’s capacity, published policy, response latency, and the crawl’s purpose determine what is appropriate. Scrapy’s documented settings are per crawler instance, so two running instances can together exceed the limits shown here.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Choosing Scrapy or a small asyncio crawler

Use a small custom asyncio client when the task is deliberately narrow—for example, a teaching exercise or a fixed list of URLs—and your team is prepared to implement and maintain the frontier, retries, robots policy, deduplication, storage, and monitoring. Async libraries such as aiohttp can make concurrent network operations convenient, but the loop and all safeguards remain your responsibility.

Use Scrapy when you want an integrated crawling project structure, request scheduling machinery, spider callbacks, and operational settings. Scrapy documents AsyncCrawlerProcess and AsyncCrawlerRunner for running spiders from scripts or integrating with an existing event loop. Coroutine callbacks are supported; to use asyncio-based libraries such as aiohttp, enable asyncio support as described in Scrapy’s coroutines documentation.

There is no workload-independent speed winner. Achievable crawl rate depends on target response behavior, network, parsing cost, storage, and the request policy you are willing and permitted to use. No comparative throughput benchmark is established here, so measure your own workload rather than treating async syntax as a performance promise.

Make the crawler polite by design

Follow robots.txt rules deliberately

RFC 9309 places the Robots Exclusion Protocol file at the host’s top-level /robots.txt path and specifies UTF-8 text. After successfully fetching the file, a crawler must parse and follow its parseable rules. The RFC says crawlers should follow at least five consecutive redirects. If the file is unreachable because of a server or network error, the crawler must assume complete disallow; if it is unavailable through a 4xx response, the RFC says access may be allowed. Rule matching uses the most specific matching path rule, and when matching Allow and Disallow rules are equivalent, Allow should be used. RFC 9309 recommends not using a cached file for more than 24 hours unless it is unreachable. See the RFC 9309 specification.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Robots.txt is a crawler protocol, not an access-control system or permission grant. As RFC 9309 states, “The Robots Exclusion Protocol is not a substitute for valid content security measures.” Do not infer that a disallowed path is private, or that a permitted path is authorized for your purpose.

Control aggregate host load

Global concurrency caps total simultaneous requests; per-domain concurrency and download delay constrain pressure on an individual host. AutoThrottle can adjust delay based on observed response latency, but it does not replace scope rules or an explicit conservative policy. Identify your crawler in its User-Agent and make the identity and purpose contactable where appropriate; Scrapy recommends a documented, contactable identity when crawling is allowed.

Back off when a host returns errors, rate-limit responses, or blocking pages. Retry transient failures only a bounded number of times, with delay; endless retries turn an outage into more load. Also decide how to handle redirects: validate schemes and destinations, avoid following redirects out of scope, and do not blindly fetch arbitrary protocols. Scrapy’s project documentation covers crawler runners, settings, and common practices including distributed crawls.

Build the frontier and storage for recovery

For a short, disposable crawl, an in-memory scheduler may be enough. If losing progress on process failure is unacceptable, persist queued URLs and their state. A durable frontier should record at least the normalized request key, discovered URL, status, retry count, and scheduling information needed by your policy. Mark completion only after the fetch and record write reach the state you consider successful, so a crash does not silently discard work.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Deduplicate before enqueueing and again at the storage boundary where practical. URL normalization needs restraint: lowercasing a host and removing a fragment are usually safe normalization steps, but query parameters can change page content or identify distinct records. Do not strip tracking-looking parameters or reorder query values without evidence that the target treats them as equivalent. Keep the original URL alongside any canonical key you use.

Filter candidates before scheduling. Check scheme, host, path scope, depth, and likely content type; otherwise calendars, faceted search, session IDs, and infinite URL patterns can create an effectively unbounded crawl. A depth limit is a useful safety rail, but it is not a substitute for a URL policy.

What changes when a crawl must scale out?

First measure a single crawler: queue depth, fetch success and error counts, response latency, retry volume, duplicate rate, memory, storage write delays, and requests per host are useful engineering signals. They are diagnostic metrics, not target benchmarks. If the queue grows while the fetcher is waiting on remote responses, network policy may constrain throughput. If downloads complete but CPU or storage is saturated, adding network concurrency may make the bottleneck worse.

Scrapy does not provide built-in multi-server distributed crawling. Its documentation says, “Scrapy doesn’t provide any built-in facility for running crawls in a distributed (multi-server) manner.” For independent spiders, schedule separate spider runs. For one large spider spread across machines, partition the URL input and design the coordination around it: who owns each partition, how duplicates crossing partitions are detected, where retries and durable state live, and how results are aggregated.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Adding crawler processes does not automatically increase useful throughput. Each crawler has its own concurrency and throttle settings; running several can multiply both resource use and requests to the same host. Partition work by host or another explicit rule, coordinate per-host budgets centrally when multiple workers can reach one host, and keep host impact separate from internal throughput. Raise limits only after observing their effect on both your system and the destination.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup

A crawler is the right tool for traversing links and extracting records. If the task is instead to capture a single page as an image or PDF, ScreenshotNeo’s screenshot API is a separate one-request option—not a replacement for a link-following crawler. It can remove cookie/consent banners, newsletter popups, and chat widgets before capture. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed; response headers report the page verdict and billing outcome. It also provides an MCP server with take_screenshot, get_page_info, and capture_pdf tools for AI agents.

Install requests with python -m pip install requests, set an API key, and use the documented call below. See the ScreenshotNeo API documentation for request options.

import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
open("shot.webp", "wb").write(r.content)

The same request can be made from cURL:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Or from Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

ScreenshotNeo offers 1,000 screenshots a month free without a card; paid plans start at $5 for 3,000. Sign up for the free plan.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Common crawler failures and fixes

  • The crawl queue keeps growing: Inspect URL patterns and duplicate rate. Add explicit path, query, depth, and content-type filters; do not assume every discovered link belongs in scope.
  • Many requests fail or repeat: Check response codes, timeout rates, redirects, and retry counts. Reduce per-host load and use bounded retries with backoff rather than replaying failures immediately.
  • A crawl produces little data: Inspect response status, response body, and selectors on representative pages. The example extracts HTML title and description; it does not execute page JavaScript, so pages whose content is rendered only client-side may require a different rendering strategy.
  • The target returns blocking or rate-limit responses: Stop or slow the affected host, review its robots rules and access policy, and contact the site operator if the crawl is legitimate. Do not try to bypass access controls.
  • Progress disappears after interruption: An exported result file alone is not a durable frontier. Persist scheduling state and define how in-flight URLs are recovered before relying on restartable crawls.
  • Adding workers overloads a host: Reduce the combined per-host request budget. Per-crawler settings do not coordinate across processes or machines.

Further reading

For a broader treatment of crawler models, Scrapy, storage, and parallel scraping, O’Reilly lists Web Scraping with Python, 3rd Edition by Ryan Mitchell, published in February 2024. The publisher describes it as an intermediate-to-advanced, 352-page book.

Frequently Asked Questions

Does a depth limit mean the crawler will visit every page up to that level?

No. It caps link distance from the start requests; scope rules, robots rules, response failures, and duplicate suppression can all reduce the pages that are actually fetched.

Can an HTML crawler read content that appears only after JavaScript runs?

Not with the Scrapy example as written: it parses the returned response body. If the desired content is absent from that HTML, choose a rendering approach suited to the site and keep its request policy within the same scope and politeness controls.

Is robots.txt enough to establish that a crawl is allowed?

No. It communicates crawler preferences, but it is not authorization, access control, or a statement of the site’s legal terms.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.