Driver FixRecommendedSound, Wi-Fi or graphics acting up? Check drivers firstFind missing or outdated drivers fast.Check DriversOctober DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsPC HealthRecommendedCrashes, freezes, slowdowns? Check your PC nowSpot repairable issues before they interrupt work.Check PC×
Skip to content
aiohttp

Asynchronous Web Crawling at Scale: Architecture, Concurrency, Scrapy, aiohttp, and Compliance

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The reliable way to crawl asynchronously at scale is to separate orchestration from HTTP transport, then bound work twice: a global limit for the whole crawler and a per-domain limit with an explicit delay. Scrapy gives you a scheduler, retries, throttling, parsing, and exports; aiohttp gives you an asyncio transport layer and connection pool that you assemble into your own crawler. In both designs, durable URL state, robots.txt policy, cancellation, and metrics matter as much as raw concurrency.

What “asynchronous at scale” actually requires

Async I/O lets one process keep many requests in flight while other requests wait on DNS, connection establishment, or response bytes. It does not remove limits imposed by target sites, your file descriptors, DNS resolver, CPU, memory, parser, or storage. A crawler that simply raises a semaphore often gets more 429 responses, timeouts, bans, and retries—and fewer useful pages per minute.

Use this pipeline as a starting architecture:

  1. Seed ingestion: accept URLs from files, APIs, sitemaps, or databases.
  2. Normalization: canonicalize schemes, hosts, ports, fragments, and tracking parameters according to your application’s rules.
  3. Durable frontier: keep queued, in-progress, succeeded, failed, and deferred URLs in persistent storage so a process restart does not erase work.
  4. Deduplication: use a durable unique key for the canonical URL (and, when needed, a content fingerprint).
  5. Policy state: maintain robots.txt, delay, concurrency, retry, and backoff state separately for each host or domain.
  6. Fetch workers: run bounded asynchronous HTTP requests with connection reuse and cancellation.
  7. Parse and persist: move expensive parsing and writes off the fetch path when they could stall network workers.
  8. Metrics and logs: record queue depth, active requests, latency, status codes, bytes, retries, parser lag, and duplicate rates.

Keep queue admission, politeness delays, retries, and cancellation explicit. One slow domain should not occupy every worker needed by unrelated domains.

Choose the orchestration layer

Concern Scrapy-first aiohttp-first
Scheduling and URL frontier Built in through Scrapy’s engine and queues You design queues, ownership, priorities, and persistence
Retries and throttling Settings, retry middleware, and AutoThrottle Explicit retry budgets, token buckets, and backoff code
Transport control Controlled through downloader settings and middleware Direct asyncio session, connector, timeout, and streaming controls
Parsing and exports Selectors, item pipelines, and feed exporters Any parser and storage stack you choose
Multi-machine operation No built-in shared queue for one spider; partition inputs or assign queue ownership Same limitation unless you build shared ownership and deduplication
Operational effort Lower for conventional crawls Higher, but useful when transport and event-loop behavior must be customized

Scrapy exposes asynchronous runners such as AsyncCrawlerProcess and AsyncCrawlerRunner, works with the asyncio reactor, and provides crawl-level controls including CONCURRENT_REQUESTS, CONCURRENT_REQUESTS_PER_DOMAIN, DOWNLOAD_DELAY, and AutoThrottle. Choose aiohttp when you need a smaller transport layer inside an existing asyncio service or need to control every queue and connection detail yourself.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Set safe concurrency and rate limits

Use two independent limits

Set a global maximum to protect your process and a per-domain maximum to protect each target. A global value of 200 with a per-domain value of 4 means at most 200 active requests overall, but no more than four to one domain. Add a per-domain delay or token bucket so a fast response does not turn into a burst.

There is no universal pages-per-second number. Useful throughput depends on target latency and tolerance, DNS, response size, parser cost, storage, and retry behavior. Start conservatively, observe status codes and latency, and increase domain parallelism only while CPU, memory, DNS, file descriptors, and downstream systems remain healthy.

Scrapy settings

CONCURRENT_REQUESTS = 64
CONCURRENT_REQUESTS_PER_DOMAIN = 4
DOWNLOAD_DELAY = 0.5
AUTOTHROTTLE_ENABLED = True
AUTOTHROTTLE_START_DELAY = 0.5
AUTOTHROTTLE_MAX_DELAY = 30
AUTOTHROTTLE_TARGET_CONCURRENCY = 2.0
RETRY_ENABLED = True
RETRY_TIMES = 2
COOKIES_ENABLED = False

For a broad crawl spanning many domains, use Scrapy’s DownloaderAwarePriorityQueue. Scrapy documents it for keeping many domains active while allowing each site to be crawled slowly; the default priority queue is optimized for a single-domain crawl.

Aiohttp controls

Reuse one aiohttp.ClientSession (or a deliberately managed pool) instead of opening a connection for every URL. The session connector pools connections. Calling session.get() obtains response headers; reading the body is a separate awaited operation, so always consume or close the response.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
import asyncio
import aiohttp
from urllib.parse import urlparse

GLOBAL_LIMIT = 100
PER_HOST_LIMIT = 3
REQUEST_TIMEOUT = aiohttp.ClientTimeout(total=45, connect=10, sock_read=30)

class HostLimiter:
    def __init__(self):
        self.locks = {}

    def for_url(self, url):
        host = urlparse(url).netloc.lower()
        return self.locks.setdefault(host, asyncio.Semaphore(PER_HOST_LIMIT))

async def fetch(session, url, global_sem, host_limiter):
    async with global_sem, host_limiter.for_url(url):
        await asyncio.sleep(0.25)  # replace with a per-host token bucket in production
        try:
            async with session.get(url, allow_redirects=True) as response:
                body = await response.read()
                return {
                    "url": str(response.url),
                    "status": response.status,
                    "headers": dict(response.headers),
                    "body": body,
                }
        except (asyncio.TimeoutError, aiohttp.ClientError) as exc:
            return {"url": url, "error": repr(exc)}

async def main(urls):
    global_sem = asyncio.Semaphore(GLOBAL_LIMIT)
    limiter = HostLimiter()
    connector = aiohttp.TCPConnector(limit=GLOBAL_LIMIT, limit_per_host=PER_HOST_LIMIT)
    async with aiohttp.ClientSession(connector=connector,
                                     timeout=REQUEST_TIMEOUT,
                                     headers={"User-Agent": "ExampleCrawler/1.0"}) as session:
        tasks = [fetch(session, url, global_sem, limiter) for url in urls]
        for result in await asyncio.gather(*tasks, return_exceptions=False):
            print(result["url"], result.get("status"), result.get("error"))

if __name__ == "__main__":
    asyncio.run(main(["https://example.com/", "https://www.iana.org/"]))

This example limits simultaneous work but does not implement robots.txt or durable queues; add those before using it on sites you do not control. In a long crawl, use bounded queue capacity so producers cannot accumulate unlimited URLs, apply response-size limits, and propagate cancellation when a job is stopped.

Build a Scrapy broad crawl

For many domains, partition seeds by URL or domain and let each worker run a normal spider. Set CONCURRENT_REQUESTS_PER_DOMAIN and delays rather than one enormous global value. Keep cookies disabled unless the target requires a session, and enable HTTP caching during development to avoid repeatedly downloading the same resource.

Scrapy does not provide built-in multi-server distribution for one large spider. A documented approach is to partition URL lists and run the partitions on separate Scrapyd servers. For a stronger production design, assign frontier ownership explicitly (for example by a stable hash of host), store leases with expirations, and write deduplication and checkpoints to durable storage. A worker must be able to resume after a crash without issuing an unbounded duplicate wave.

Robots.txt is a scheduler input, not an afterthought

Fetch and parse robots.txt before admitting URLs for a host. Apply the most specific matching rule, record the policy version and fetch time, and handle redirects and status failures. Translate Crawl-delay and Request-rate into your delay and concurrency settings; Scrapy does not apply those directives automatically.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

RFC 9309 (September 2022) defines the Robots Exclusion Protocol. A crawler that successfully downloads a file must follow its parseable rules. Its handling of redirects distinguishes unavailable from unreachable responses, and it calls for conservative caching. If the file is unreachable under the applicable semantics, have the scheduler fail closed rather than continue blindly. Robots rules are not authentication or a security boundary: the RFC explicitly says, “These rules are not a form of access authorization.”

Prefer an API, bulk export, search endpoint, or sitemap when it can replace page-by-page crawling. This reduces load and usually produces cleaner, more stable data.

Retries, timeouts, and cancellation

Budget retries as capacity

Retry only transient failures such as connection resets, selected 5xx responses, and 429 responses after honoring any server-provided delay. Use exponential backoff with jitter and a small maximum attempt count. A retry of a slow response consumes a slot that could serve a new URL; Scrapy’s optimization guidance warns that retries of slow or failing responses can substantially reduce crawl capacity.

Use separate timeouts

Set connect, total, and read timeouts separately when your client supports them. A stuck read should release its semaphore slot. Cap response bytes before buffering large bodies, and stream to disk when full documents are expected.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Make shutdown resumable

On cancellation, stop admitting new URLs, cancel pending tasks, let active responses close, and persist their URLs back to the frontier with a retry state. Checkpoint after batches rather than only at process exit.

Scaling across machines

Scale domain parallelism first, not requests to one host. A practical partition key is normalized hostname; it keeps per-host rate state local and reduces coordination. If URLs from one domain must be processed by several workers, place that domain’s token bucket and lease ownership in a shared store.

  • Queue: durable ready, delayed, in-flight, succeeded, and dead-letter states.
  • Deduplication: an atomic insert or compare-and-set operation on the canonical URL.
  • Leases: ownership expiry so a crashed worker’s URLs return to the queue.
  • Policy cache: robots body, status, fetched time, expiry, and parser result.
  • Observability: per-domain latency, status distribution, retry rate, bytes, queue age, and parser lag.

Raise global concurrency only when those measurements remain healthy. Improve DNS resolution, reduce unnecessary retries, and lower download timeouts for stuck requests before adding machines. When memory is constrained, use disk-backed job state; breadth-first scheduling usually spreads work across domains, while depth-first scheduling can concentrate it.

Performance and cost notes

Benchmark representative domains with explicit safety limits instead of publishing a generic pages-per-second target. Include redirects, robots fetches, realistic body sizes, parser work, writes, and failure rates in the test. Measure useful documents per minute, not just request completions.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Connection reuse reduces handshake overhead. Bounded queues prevent memory growth. Disabling unnecessary cookies lowers state. Caching during development prevents accidental load, but production cache policy must respect freshness and site rules. Storage and parsing can become the bottleneck after network concurrency rises; profile them independently.

Troubleshooting common failures

Many 429 or 503 responses

Lower per-domain concurrency, increase delay, honor Retry-After, and reduce retry attempts. A higher global limit will usually make this worse.

Requests hang until the process runs out of slots

Set connect and read timeouts, ensure every response body is consumed or closed, cap body size, and verify cancellation reaches waiting tasks.

Duplicate pages after a restart

Move deduplication and frontier state out of memory. Use atomic URL claims and leases, and checkpoint in-flight work so recovery is deterministic.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

One domain stalls the whole crawl

Use per-domain queues and semaphores rather than one shared lock. Apply a circuit breaker or delayed queue for repeated failures so other domains continue.

Robots decisions are inconsistent

Centralize parsing, store the fetched status and timestamp, follow redirects, and apply the most specific rule before enqueueing. Do not let individual workers invent their own policy.

CPU or memory spikes when concurrency rises

Reduce body buffering, stream large responses, move parsing to bounded workers, and inspect queue depth and file-descriptor use. Increase concurrency only after the bottleneck is removed.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup

If your task is to collect rendered page images rather than crawl links and extract documents, ScreenshotNeo provides a one-request website screenshot API. It accepts consent banners as a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be disabled. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing result. It also offers an MCP server with take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use the API documentation at screenshotneo.com/docs/ for all options, including full-page and selector captures, device presets, dark mode, custom CSS and JavaScript, waits, request blocking, headers and cookies, geolocation, PDFs, resizing, caching, signed links, asynchronous webhooks, bulk capture of up to 100 URLs per call, usage reporting, and the OpenAPI specification.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

The Free plan includes 1,000 shots per month with no card. Paid plans are Starter $5 for 3,000, Growth $15 for 15,000, Pro $39 for 60,000, Scale $99 for 250,000, and Business $249 for 1,000,000; yearly billing gives two months free, and every feature is on every plan. Create a free ScreenshotNeo account to start without a card.

Frequently Asked Questions

Should I use threads instead of asyncio for a crawler?

For I/O-heavy HTTP work, asyncio provides explicit connection pooling and cancellation. Threads can work, but mixing blocking DNS, parsers, or libraries into an event loop requires careful isolation and bounded executors.

Is a single shared queue enough for a distributed crawl?

Only if it also provides atomic claims, leases, durable deduplication, delayed retries, and recovery for abandoned work. Otherwise workers will duplicate URLs or lose them during failures.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Can robots.txt tell me whether a page is legal to access?

No. Robots Exclusion rules communicate crawler preferences; RFC 9309 states that they are not access authorization. Authentication, contracts, copyright, and applicable law remain separate questions.

What should I monitor first after deployment?

Monitor queue age and depth, active requests, per-domain latency, status-code distribution, retry rate, bytes, parser lag, duplicate rate, and resource limits such as memory, CPU, DNS, and file descriptors.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Read next

Recommended PC Tool
Recommended PC Tool
PC Slower Than It Used to Be?Free scan - under a minute
Crashes, No Sound, or Screen Glitches?Free driver scan

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.