The reliable way to crawl asynchronously at scale is to separate orchestration from HTTP transport, then bound work twice: a global limit for the whole crawler and a per-domain limit with an explicit delay. Scrapy gives you a scheduler, retries, throttling, parsing, and exports; aiohttp gives you an asyncio transport layer and connection pool that you assemble into your own crawler. In both designs, durable URL state, robots.txt policy, cancellation, and metrics matter as much as raw concurrency.
What “asynchronous at scale” actually requires
Async I/O lets one process keep many requests in flight while other requests wait on DNS, connection establishment, or response bytes. It does not remove limits imposed by target sites, your file descriptors, DNS resolver, CPU, memory, parser, or storage. A crawler that simply raises a semaphore often gets more 429 responses, timeouts, bans, and retries—and fewer useful pages per minute.
Use this pipeline as a starting architecture:
- Seed ingestion: accept URLs from files, APIs, sitemaps, or databases.
- Normalization: canonicalize schemes, hosts, ports, fragments, and tracking parameters according to your application’s rules.
- Durable frontier: keep queued, in-progress, succeeded, failed, and deferred URLs in persistent storage so a process restart does not erase work.
- Deduplication: use a durable unique key for the canonical URL (and, when needed, a content fingerprint).
- Policy state: maintain robots.txt, delay, concurrency, retry, and backoff state separately for each host or domain.
- Fetch workers: run bounded asynchronous HTTP requests with connection reuse and cancellation.
- Parse and persist: move expensive parsing and writes off the fetch path when they could stall network workers.
- Metrics and logs: record queue depth, active requests, latency, status codes, bytes, retries, parser lag, and duplicate rates.
Keep queue admission, politeness delays, retries, and cancellation explicit. One slow domain should not occupy every worker needed by unrelated domains.
Choose the orchestration layer
| Concern | Scrapy-first | aiohttp-first |
|---|---|---|
| Scheduling and URL frontier | Built in through Scrapy’s engine and queues | You design queues, ownership, priorities, and persistence |
| Retries and throttling | Settings, retry middleware, and AutoThrottle | Explicit retry budgets, token buckets, and backoff code |
| Transport control | Controlled through downloader settings and middleware | Direct asyncio session, connector, timeout, and streaming controls |
| Parsing and exports | Selectors, item pipelines, and feed exporters | Any parser and storage stack you choose |
| Multi-machine operation | No built-in shared queue for one spider; partition inputs or assign queue ownership | Same limitation unless you build shared ownership and deduplication |
| Operational effort | Lower for conventional crawls | Higher, but useful when transport and event-loop behavior must be customized |
Scrapy exposes asynchronous runners such as AsyncCrawlerProcess and AsyncCrawlerRunner, works with the asyncio reactor, and provides crawl-level controls including CONCURRENT_REQUESTS, CONCURRENT_REQUESTS_PER_DOMAIN, DOWNLOAD_DELAY, and AutoThrottle. Choose aiohttp when you need a smaller transport layer inside an existing asyncio service or need to control every queue and connection detail yourself.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
#1 Best Overall
Set safe concurrency and rate limits
Use two independent limits
Set a global maximum to protect your process and a per-domain maximum to protect each target. A global value of 200 with a per-domain value of 4 means at most 200 active requests overall, but no more than four to one domain. Add a per-domain delay or token bucket so a fast response does not turn into a burst.
There is no universal pages-per-second number. Useful throughput depends on target latency and tolerance, DNS, response size, parser cost, storage, and retry behavior. Start conservatively, observe status codes and latency, and increase domain parallelism only while CPU, memory, DNS, file descriptors, and downstream systems remain healthy.
Scrapy settings
CONCURRENT_REQUESTS = 64
CONCURRENT_REQUESTS_PER_DOMAIN = 4
DOWNLOAD_DELAY = 0.5
AUTOTHROTTLE_ENABLED = True
AUTOTHROTTLE_START_DELAY = 0.5
AUTOTHROTTLE_MAX_DELAY = 30
AUTOTHROTTLE_TARGET_CONCURRENCY = 2.0
RETRY_ENABLED = True
RETRY_TIMES = 2
COOKIES_ENABLED = False
For a broad crawl spanning many domains, use Scrapy’s DownloaderAwarePriorityQueue. Scrapy documents it for keeping many domains active while allowing each site to be crawled slowly; the default priority queue is optimized for a single-domain crawl.
Aiohttp controls
Reuse one aiohttp.ClientSession (or a deliberately managed pool) instead of opening a connection for every URL. The session connector pools connections. Calling session.get() obtains response headers; reading the body is a separate awaited operation, so always consume or close the response.
The Tool Desk
Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →import asyncio
import aiohttp
from urllib.parse import urlparse
GLOBAL_LIMIT = 100
PER_HOST_LIMIT = 3
REQUEST_TIMEOUT = aiohttp.ClientTimeout(total=45, connect=10, sock_read=30)
class HostLimiter:
def __init__(self):
self.locks = {}
def for_url(self, url):
host = urlparse(url).netloc.lower()
return self.locks.setdefault(host, asyncio.Semaphore(PER_HOST_LIMIT))
async def fetch(session, url, global_sem, host_limiter):
async with global_sem, host_limiter.for_url(url):
await asyncio.sleep(0.25) # replace with a per-host token bucket in production
try:
async with session.get(url, allow_redirects=True) as response:
body = await response.read()
return {
"url": str(response.url),
"status": response.status,
"headers": dict(response.headers),
"body": body,
}
except (asyncio.TimeoutError, aiohttp.ClientError) as exc:
return {"url": url, "error": repr(exc)}
async def main(urls):
global_sem = asyncio.Semaphore(GLOBAL_LIMIT)
limiter = HostLimiter()
connector = aiohttp.TCPConnector(limit=GLOBAL_LIMIT, limit_per_host=PER_HOST_LIMIT)
async with aiohttp.ClientSession(connector=connector,
timeout=REQUEST_TIMEOUT,
headers={"User-Agent": "ExampleCrawler/1.0"}) as session:
tasks = [fetch(session, url, global_sem, limiter) for url in urls]
for result in await asyncio.gather(*tasks, return_exceptions=False):
print(result["url"], result.get("status"), result.get("error"))
if __name__ == "__main__":
asyncio.run(main(["https://example.com/", "https://www.iana.org/"]))
This example limits simultaneous work but does not implement robots.txt or durable queues; add those before using it on sites you do not control. In a long crawl, use bounded queue capacity so producers cannot accumulate unlimited URLs, apply response-size limits, and propagate cancellation when a job is stopped.
Build a Scrapy broad crawl
For many domains, partition seeds by URL or domain and let each worker run a normal spider. Set CONCURRENT_REQUESTS_PER_DOMAIN and delays rather than one enormous global value. Keep cookies disabled unless the target requires a session, and enable HTTP caching during development to avoid repeatedly downloading the same resource.
Rank #2
Scrapy does not provide built-in multi-server distribution for one large spider. A documented approach is to partition URL lists and run the partitions on separate Scrapyd servers. For a stronger production design, assign frontier ownership explicitly (for example by a stable hash of host), store leases with expirations, and write deduplication and checkpoints to durable storage. A worker must be able to resume after a crash without issuing an unbounded duplicate wave.
Robots.txt is a scheduler input, not an afterthought
Fetch and parse robots.txt before admitting URLs for a host. Apply the most specific matching rule, record the policy version and fetch time, and handle redirects and status failures. Translate Crawl-delay and Request-rate into your delay and concurrency settings; Scrapy does not apply those directives automatically.
RFC 9309 (September 2022) defines the Robots Exclusion Protocol. A crawler that successfully downloads a file must follow its parseable rules. Its handling of redirects distinguishes unavailable from unreachable responses, and it calls for conservative caching. If the file is unreachable under the applicable semantics, have the scheduler fail closed rather than continue blindly. Robots rules are not authentication or a security boundary: the RFC explicitly says, “These rules are not a form of access authorization.”
Prefer an API, bulk export, search endpoint, or sitemap when it can replace page-by-page crawling. This reduces load and usually produces cleaner, more stable data.
Retries, timeouts, and cancellation
Budget retries as capacity
Retry only transient failures such as connection resets, selected 5xx responses, and 429 responses after honoring any server-provided delay. Use exponential backoff with jitter and a small maximum attempt count. A retry of a slow response consumes a slot that could serve a new URL; Scrapy’s optimization guidance warns that retries of slow or failing responses can substantially reduce crawl capacity.
Use separate timeouts
Set connect, total, and read timeouts separately when your client supports them. A stuck read should release its semaphore slot. Cap response bytes before buffering large bodies, and stream to disk when full documents are expected.
Make shutdown resumable
On cancellation, stop admitting new URLs, cancel pending tasks, let active responses close, and persist their URLs back to the frontier with a retry state. Checkpoint after batches rather than only at process exit.
Scaling across machines
Scale domain parallelism first, not requests to one host. A practical partition key is normalized hostname; it keeps per-host rate state local and reduces coordination. If URLs from one domain must be processed by several workers, place that domain’s token bucket and lease ownership in a shared store.
- Queue: durable ready, delayed, in-flight, succeeded, and dead-letter states.
- Deduplication: an atomic insert or compare-and-set operation on the canonical URL.
- Leases: ownership expiry so a crashed worker’s URLs return to the queue.
- Policy cache: robots body, status, fetched time, expiry, and parser result.
- Observability: per-domain latency, status distribution, retry rate, bytes, queue age, and parser lag.
Raise global concurrency only when those measurements remain healthy. Improve DNS resolution, reduce unnecessary retries, and lower download timeouts for stuck requests before adding machines. When memory is constrained, use disk-backed job state; breadth-first scheduling usually spreads work across domains, while depth-first scheduling can concentrate it.
Performance and cost notes
Benchmark representative domains with explicit safety limits instead of publishing a generic pages-per-second target. Include redirects, robots fetches, realistic body sizes, parser work, writes, and failure rates in the test. Measure useful documents per minute, not just request completions.
Connection reuse reduces handshake overhead. Bounded queues prevent memory growth. Disabling unnecessary cookies lowers state. Caching during development prevents accidental load, but production cache policy must respect freshness and site rules. Storage and parsing can become the bottleneck after network concurrency rises; profile them independently.
Troubleshooting common failures
Many 429 or 503 responses
Lower per-domain concurrency, increase delay, honor Retry-After, and reduce retry attempts. A higher global limit will usually make this worse.
Requests hang until the process runs out of slots
Set connect and read timeouts, ensure every response body is consumed or closed, cap body size, and verify cancellation reaches waiting tasks.
Duplicate pages after a restart
Move deduplication and frontier state out of memory. Use atomic URL claims and leases, and checkpoint in-flight work so recovery is deterministic.
PC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minuteOne domain stalls the whole crawl
Use per-domain queues and semaphores rather than one shared lock. Apply a circuit breaker or delayed queue for repeated failures so other domains continue.
Robots decisions are inconsistent
Centralize parsing, store the fetched status and timestamp, follow redirects, and apply the most specific rule before enqueueing. Do not let individual workers invent their own policy.
CPU or memory spikes when concurrency rises
Reduce body buffering, stream large responses, move parsing to bounded workers, and inspect queue depth and file-descriptor use. Increase concurrency only after the bottleneck is removed.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Or skip the browser setup
If your task is to collect rendered page images rather than crawl links and extract documents, ScreenshotNeo provides a one-request website screenshot API. It accepts consent banners as a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be disabled. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing result. It also offers an MCP server with take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients.
Use the API documentation at screenshotneo.com/docs/ for all options, including full-page and selector captures, device presets, dark mode, custom CSS and JavaScript, waits, request blocking, headers and cookies, geolocation, PDFs, resizing, caching, signed links, asynchronous webhooks, bulk capture of up to 100 URLs per call, usage reporting, and the OpenAPI specification.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
The Free plan includes 1,000 shots per month with no card. Paid plans are Starter $5 for 3,000, Growth $15 for 15,000, Pro $39 for 60,000, Scale $99 for 250,000, and Business $249 for 1,000,000; yearly billing gives two months free, and every feature is on every plan. Create a free ScreenshotNeo account to start without a card.
Frequently Asked Questions
Should I use threads instead of asyncio for a crawler?
For I/O-heavy HTTP work, asyncio provides explicit connection pooling and cancellation. Threads can work, but mixing blocking DNS, parsers, or libraries into an event loop requires careful isolation and bounded executors.
Is a single shared queue enough for a distributed crawl?
Only if it also provides atomic claims, leases, durable deduplication, delayed retries, and recovery for abandoned work. Otherwise workers will duplicate URLs or lose them during failures.
Do these 3 things before closing this tab:
1Clear out junk files and repair common Windows errors2Scan for outdated or missing drivers - takes under a minute3Repair Windows errors before they cause bigger problemsCan robots.txt tell me whether a page is legal to access?
No. Robots Exclusion rules communicate crawler preferences; RFC 9309 states that they are not access authorization. Authentication, contracts, copyright, and applicable law remain separate questions.
What should I monitor first after deployment?
Monitor queue age and depth, active requests, per-domain latency, status-code distribution, retry rate, bytes, parser lag, duplicate rate, and resource limits such as memory, CPU, DNS, and file descriptors.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




