Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

Use several controls together: obey the target site’s robots.txt and published limits, prefer an API or export, cap both total and per-domain concurrency, enforce a minimum delay, and back off when latency or errors rise. Start at one request at a time per domain, increase load in small steps, and reduce it immediately when you see 429 or 503 responses, ban pages, rising retries, or worsening latency.

Start with permission and the least expensive interface

Before tuning a crawler, identify every hostname it will contact and fetch that host’s robots.txt using the crawler’s actual user-agent. Treat disallowed paths as out of scope. A robots file is not a universal speed limit, but its Crawl-delay and Request-rate directives are explicit signals that should be translated into your scheduler’s settings.

Also read the site’s terms, API documentation and export options. A documented API, search endpoint or bulk export normally creates less load than downloading and parsing many HTML pages. If the site publishes a rate limit, make that limit the ceiling for your client rather than trying to discover a higher one experimentally. Schedule large crawls during the site’s stated or likely idle period when possible.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The controls that make throttling reliable

Control What it limits Why it matters
Global concurrency Simultaneous requests across all hosts Prevents your process from creating a burst or exhausting local sockets and bandwidth.
Per-domain concurrency Simultaneous requests to one host Stops a single site from receiving your full crawler capacity.
Minimum delay Wait between consecutive requests to a domain Controls request frequency even when responses are fast.
Adaptive delay Delay based on observed latency and status Responds to changing server load instead of using one fixed guess.
Backoff and bounded retries What happens after transient failures or throttling Prevents a failure from becoming a retry storm.

Apply global and per-domain limits at the same time. A global limit of 20 does not make a crawl polite if all 20 requests target one host. Conversely, a per-domain limit without a global cap can overload your own machine or several smaller sites at once.

A conservative rollout procedure

  1. Inventory hosts. Separate the scheduler’s queue by effective domain (including port where relevant), and record each host’s robots rules and documented limit.
  2. Choose a starting point. Use one concurrent request per domain and a visible delay. For an undocumented site, begin conservatively rather than treating a successful first request as permission to accelerate.
  3. Bound the whole crawler. Set a global concurrency cap as well as a per-domain cap. Keep the per-domain value low enough that a short response burst cannot overwhelm the server.
  4. Measure every request. Log host, start and finish time, status, retry count, and whether the response was served from a cache. Compute request rate and latency per domain, not only aggregate averages.
  5. Increase gradually. Raise concurrency or lower delay in small steps, allowing enough requests at each step to observe the trend. Stop increasing when 429/503 responses, ban pages, retry counts, or latency begin to rise.
  6. Back off on signals. Reduce concurrency and lengthen the delay. Honor a server-provided Retry-After value when present; do not repeatedly retry a rate-limited response while ignoring it.
  7. Keep retries finite. Retry timeouts and other transient failures with exponential backoff and jitter, but do not retry robots-denied URLs. Put permanently failing URLs on a review queue instead of looping forever.

Scrapy settings for a fixed, predictable throttle

Scrapy exposes the three basic controls directly. CONCURRENT_REQUESTS is the global cap, CONCURRENT_REQUESTS_PER_DOMAIN limits one domain, and DOWNLOAD_DELAY sets the minimum wait between consecutive requests to that domain. A project generated by scrapy startproject uses one request per second per domain by default; that is a Scrapy default, not a safe universal limit for every site.

# settings.py
ROBOTSTXT_OBEY = True

# Conservative starting values; tune per target and its published rules.
CONCURRENT_REQUESTS = 8
CONCURRENT_REQUESTS_PER_DOMAIN = 1
DOWNLOAD_DELAY = 1.0

# Keep retries bounded and limited to transient conditions.
RETRY_ENABLED = True
RETRY_TIMES = 2
RETRY_HTTP_CODES = [408, 500, 502, 503, 504]

ROBOTSTXT_OBEY enables Scrapy’s robots middleware, which filters forbidden requests. Scrapy’s RetryMiddleware is intended for transient failures such as timeouts and HTTP 500 responses; it is not a reason to retry a disallowed URL or to hammer a host returning 429.

If a site’s robots file states a crawl delay, set DOWNLOAD_DELAY to at least that interval. Translate a request-rate instruction into both a frequency and a concurrency policy; when the wording is ambiguous, choose the more conservative interpretation and ask the site owner if the crawl is important.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use AutoThrottle when response time changes

A fixed delay is easy to reason about, but it cannot tell whether a host is becoming busy. Scrapy’s AutoThrottle calculates a target delay from observed response latency divided by target concurrency, averages that with the previous delay, and never lets a non-200 response shorten the delay. The result is clamped between DOWNLOAD_DELAY and AUTOTHROTTLE_MAX_DELAY.

# settings.py
ROBOTSTXT_OBEY = True
CONCURRENT_REQUESTS = 16
CONCURRENT_REQUESTS_PER_DOMAIN = 2
DOWNLOAD_DELAY = 0.5

AUTOTHROTTLE_ENABLED = True
AUTOTHROTTLE_START_DELAY = 5.0
AUTOTHROTTLE_MAX_DELAY = 60.0
AUTOTHROTTLE_TARGET_CONCURRENCY = 1.0

The current Scrapy documentation lists 5.0 seconds as the default start delay, 60.0 seconds as the default maximum delay, and 1.0 as the default target concurrency. These are documented settings, not guarantees about a particular server. A lower target, such as 0.5, makes the crawler more conservative and polite. Keep the per-domain concurrency cap in place: AutoThrottle adjusts waiting time, while the cap prevents simultaneous bursts.

AutoThrottle should be evaluated with status and latency logs. If errors climb while average latency appears normal, inspect the actual responses for ban pages or a rate-limit body; a successful TCP connection does not mean the site accepts the crawl.

Fixed delay, concurrency, adaptive delay and backoff: choosing a mix

Approach Politeness Throughput Response to changing load Operational simplicity
Fixed delay Predictable when based on a published rate Can waste capacity when responses are slow or fast Low High
Concurrency cap Prevents bursts Good for independent URLs, subject to server limits Low by itself High
AutoThrottle Slows as latency and errors indicate pressure Usually better than an overly large fixed delay High Moderate
Backoff Protects a host after failures Temporarily lower High after an error Moderate

Most production crawlers combine all four: a conservative floor, global and per-domain caps, AutoThrottle for normal variation, and backoff for explicit throttling or transient outages.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Implementing backoff without a retry storm

For a custom client, use a per-domain scheduler rather than sleeping in every worker independently. Each domain should have a next-allowed timestamp and a concurrency counter. A worker acquires a domain slot, waits until the timestamp, sends one request, records the result, then releases the slot and advances the timestamp by the current delay.

  1. On a successful response, retain or cautiously reduce the delay only after a sustained healthy window.
  2. On 429 or 503, increase the delay, reduce that domain’s concurrency, and honor Retry-After if supplied.
  3. On a timeout or 5xx transient error, retry only a bounded number of times with exponential backoff and random jitter.
  4. On a robots denial, authentication failure, or other permanent condition, stop retrying and record the reason.

Jitter prevents many workers from waking at exactly the same instant. Persist the scheduler state if a crawl can restart; otherwise a process restart may erase the backoff and immediately recreate the burst that caused the problem.

How to tell that your rate is too high

  • HTTP 429 or 503 responses increase, especially in consecutive batches.
  • The site returns a ban, challenge, or “too many requests” page with an otherwise successful HTTP status.
  • Retry counts rise or the same URLs repeatedly time out.
  • Median and tail latency trend upward as concurrency increases.
  • The server publishes a new rate-limit response or asks you to slow down.

Do not infer safety from your own CPU utilization or from a short run with no errors. A host can tolerate a small sample and reject a longer crawl, or apply limits per IP, account, path, or time window. Keep per-domain metrics so one overloaded host is not hidden by healthy traffic elsewhere.

Troubleshooting common failures

429 Too Many Requests

Cause: the request rate or concurrency exceeded a server or intermediary limit. Fix: stop increasing load, honor Retry-After, lower per-domain concurrency, lengthen the delay, and resume with a gradual ramp. Do not solve it by multiplying retries.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

503 responses or timeouts

Cause: temporary server pressure, an overloaded path, or a network timeout. Fix: use bounded exponential backoff, reduce concurrency, and compare latency by host and path. Retry only the transient status codes you have explicitly selected.

Robots-denied requests still appear in logs

Cause: robots middleware is disabled, the wrong user-agent was evaluated, or requests were generated before filtering. Fix: enable ROBOTSTXT_OBEY, verify the user-agent and robots URL, and remove denied URLs from the queue rather than retrying them.

AutoThrottle is still too aggressive

Cause: the target concurrency is too high for that host, or a low floor permits a burst. Fix: lower AUTOTHROTTLE_TARGET_CONCURRENCY (for example, to 0.5), lower CONCURRENT_REQUESTS_PER_DOMAIN, and raise DOWNLOAD_DELAY. Check non-200 responses; AutoThrottle will not use them as a reason to speed up.

Throughput is unexpectedly low

Cause: a delay inherited from robots rules, a high observed latency, a small global cap, or repeated retries. Fix: inspect per-domain delay, active slots, status counts and retry logs before raising limits. If a documented API or export exists, switch to it instead of optimizing page requests.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Performance, caching and data quality

Throttle the work that reaches the origin. Cache responses when the site’s terms permit it, avoid refetching unchanged URLs, and deduplicate the queue before requests are scheduled. Conditional requests can reduce transferred bytes, but they still count as requests for many rate policies. Keep parsing and storage work separate from the network scheduler so slow downstream processing does not accidentally create uncontrolled request bursts.

Record at least request rate, active concurrency, status code, retry count, response size and latency percentiles per domain. A daily average can hide a burst; one-minute or shorter windows are useful during ramp-up. Retain enough history to correlate a settings change with the first 429, 503 or latency increase.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup

If your actual requirement is to obtain clean visual captures rather than crawl and parse pages, ScreenshotNeo provides a single website-screenshot request. It accepts consent banners before capture and removes more than 60 known consent platforms, newsletter popups and chat widgets; each step can be turned off. Bot checks, blank pages, timeouts, failed loads and cache hits are not billed, and the response identifies the result with X-Page-Verdict and X-Billed headers. This is a screenshot service, not a substitute for permission to scrape or for a site’s API terms.

Use the documented endpoint and options at https://screenshotneo.com/docs/. A basic call is:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

The same request in Python:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

And Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

ScreenshotNeo also supports full-page captures with lazy images loaded, CSS-selector element capture, dark mode, 12 device presets and custom viewports, retina scale, PDF settings, custom CSS and JavaScript, clicks, selector or network-idle waits, request and resource blocking, headers, cookies, user agents, authorization, timezone and geolocation, transparent backgrounds, resizing, chosen cache TTLs, signed image links, asynchronous jobs with signed webhooks, bulk capture of up to 100 URLs per call, a usage API and an OpenAPI specification. Its parameter names are compatible with those used by other screenshot APIs, which can simplify a migration.

Every plan includes every feature: 1,000 screenshots per month free with no card; Starter is $5 for 3,000, Growth $15 for 15,000, Pro $39 for 60,000, Scale $99 for 250,000, and Business $249 for 1,000,000. Yearly billing gives two months free. An MCP server exposes take_screenshot, get_page_info and capture_pdf to Claude, Cursor and other MCP clients, so an AI agent can request captures without you maintaining browser automation. Start with the free ScreenshotNeo account.

Further reading

For a longer treatment of crawler architecture and Scrapy, Ryan Mitchell’s Web Scraping with Python, 2nd Edition is a 306-page O’Reilly book published in April 2018. Its contents include “Writing Web Crawlers” and “Scrapy.” Scrapy’s current documentation remains the authority for the settings and defaults described above, which can change in future releases.

Frequently Asked Questions

Is one request per second always safe?

No. One request per second is the default per-domain behavior of a Scrapy project generated by startproject, not a universal allowance. The target site’s rules, infrastructure and account limits determine an appropriate rate.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Should I use delay or concurrency to prevent 429 errors?

Use both. A delay limits frequency, while a per-domain concurrency cap prevents simultaneous bursts; either control alone can leave gaps in protection.

Can retries bypass a rate limit?

No. Unbounded retries usually intensify the problem. Honor server-provided wait instructions, back off, reduce concurrency and keep retry counts finite.

When should I choose AutoThrottle?

Choose it when response latency varies or you crawl more than a small, stable set of pages. Keep explicit global and per-domain caps and tune its target conservatively.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.