Do these 3 things before closing this tab:
1Clear out junk files and repair common Windows errors2Scan for outdated or missing drivers - takes under a minute3Repair Windows errors before they cause bigger problemsScale a web scraper by measuring a representative crawl, finding its actual bottleneck, and increasing work only while the target site and your own system remain healthy. Set conservative per-domain limits, watch response and resource signals as you tune, and distribute independent work only when a measured constraint justifies the extra coordination. More workers do not automatically mean more useful data: they can also multiply traffic, retries, memory use, and operational failure modes.
What scalable scraping means in practice
A scalable scraper processes more useful work without losing control of target-site load, data quality, or its own resource use. Useful work might mean successfully extracted records per unit of time—not simply requests sent. The right design depends on what limits a representative crawl: request production, downloads, parsing, CPU, memory, DNS, bandwidth, disk, or the target site’s tolerance.
The implementation examples here use Scrapy. Its settings and AutoThrottle behavior are framework-specific; do not assume another crawler applies the same controls in the same way. Scrapy’s optimization guidance is documented in its Optimization documentation.
Measure before increasing concurrency
Run a crawl representative of the sites, page types, and extraction work you intend to handle. Record a baseline, then change one limiting factor at a time and compare useful output as well as error and resource signals. This measurement-and-feedback loop is an engineering approach based on Scrapy’s diagnostic guidance, not a claim that a particular setting will produce a particular speedup.
#1 Best Overall
Signals worth recording
- Pages downloaded and items extracted per unit of time.
- Response status counts, retry counts, and response latency.
- Active downloader requests and scheduler queue depth.
- CPU, memory, network bandwidth, and disk activity.
Interpret the signals together. A flat crawl rate after raising concurrency suggests that another constraint may be limiting progress. An empty scheduler can indicate that the spider is not producing requests fast enough. A queue that grows continually means discovery is outpacing downloads and can drive memory use upward. If responses arrive faster than callbacks or item pipelines can handle them, response processing may be the bottleneck. These are diagnostic possibilities, not proof of a cause; inspect the crawl before changing architecture.
Choose an access path before crawling pages
Before building a broad page crawl, check whether the site offers a documented API, bulk export, or search endpoint that provides the data you need. Scrapy’s optimization guidance notes that such access paths can be faster for the scraper and cheaper for the site; the available fields, freshness, terms, and rate limits depend on the target.
Check the site’s terms and documentation, and inspect its robots.txt for crawler guidance. The Robots Exclusion Protocol is specified in RFC 9309. A robots file is not authorization to disregard site terms, access controls, or applicable law. Scrapy’s ROBOTSTXT_OBEY setting concerns robots rules; Scrapy does not automatically translate robots.txt Crawl-delay or Request-rate directives into download settings. If relevant, map those instructions and any stricter documented API limits into your own configuration.
Set request limits per target, not just globally
Scrapy has separate controls for total active downloads, per-domain concurrency, and spacing requests to a domain. A global ceiling by itself does not keep traffic to one host within a sensible bound, particularly when a crawl spans only a few domains—or when several spiders or workers are running.
Recommended Free Tools
| Scrapy setting | What it controls | How to use it |
|---|---|---|
CONCURRENT_REQUESTS |
Global cap on active downloads. | Set a total limit that fits your resources and intended aggregate workload. |
CONCURRENT_REQUESTS_PER_DOMAIN |
Concurrent requests for a domain. | Use it to cap simultaneous work against an individual target. |
DOWNLOAD_DELAY |
Spacing between requests to a domain. | Use a delay when the target’s guidance or observed behavior calls for slower requests. |
There is no universal safe requests-per-second value established here. Site tolerance, terms, and documented service limits differ. Increase limits gradually; keep an eye on 429 and 503 responses, retries, and latency, and back off when signals or the target’s guidance warrant it.
Let AutoThrottle adapt to response behavior
Scrapy’s AutoThrottle adjusts per-slot download delay using observed response latency, while respecting configured delay and concurrency bounds. AUTOTHROTTLE_TARGET_CONCURRENCY is an average target the extension tries to approach, not a hard instantaneous cap. Non-200 response latencies may increase the delay, but do not cause AutoThrottle to reduce it. Read the AutoThrottle documentation for the extension’s algorithm and settings.
AutoThrottle is a control, not a guarantee that a particular site will accept a particular rate. Keep per-domain limits, inspect responses, and respect site-specific terms and instructions.
A small Scrapy starting point
The following illustrates conservative, explicit settings and a spider that follows links from a seed page and extracts matching elements. Replace the example domain, start URL, selectors, and item fields with ones that the target site permits you to access. It assumes Scrapy is installed and a Scrapy project has been created; place the spider in that project’s spiders directory and edit its settings.py.
Quick wins for a faster PC:
Clear out junk files and repair common Windows errorsFree Scan →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Repair Windows errors before they cause bigger problemsFix Now →Rank #3
Project settings
ROBOTSTXT_OBEY = True
CONCURRENT_REQUESTS = 16
CONCURRENT_REQUESTS_PER_DOMAIN = 2
DOWNLOAD_DELAY = 1.0
AUTOTHROTTLE_ENABLED = True
AUTOTHROTTLE_START_DELAY = 1.0
AUTOTHROTTLE_MAX_DELAY = 60.0
AUTOTHROTTLE_TARGET_CONCURRENCY = 1.0
These are example starting values, not a universal recommendation or a guarantee of acceptable load. Confirm the target’s rules and tune against your measured crawl. Scrapy settings can also be supplied for an individual run using -s; use that for controlled experiments rather than changing several variables at once.
Spider example
import scrapy
class CatalogSpider(scrapy.Spider):
name = "catalog"
allowed_domains = ["example.com"]
start_urls = ["https://example.com/catalog/"]
def parse(self, response):
for card in response.css(".product-card"):
yield {
"name": card.css(".product-name::text").get(default="").strip(),
"url": response.urljoin(
card.css("a::attr(href)").get(default="")
),
}
for href in response.css("a.next::attr(href)").getall():
yield response.follow(href, callback=self.parse)
Run it from the project directory with scrapy crawl catalog -O items.jsonl. The spider is deliberately simple: confirm that the selectors match real pages, that pagination does not create loops, and that output is correct before raising limits. The sample’s illustrative domain is not a target recommendation.
Or skip the browser setup
If the job is to capture a page as an image or PDF rather than extract structured records, ScreenshotNeo can return a screenshot or PDF with one GET request. It is not a replacement for a crawler that discovers URLs or extracts records. The request below saves a WebP capture of the example URL; put your API key in place of YOUR_API_KEY. See the ScreenshotNeo API documentation for request options.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Equivalent Python:
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
Equivalent Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
- Cookie and consent banners, newsletter popups, and chat widgets are removed before the shot; each cleanup step can be turned off.
- Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed; response headers identify the page verdict and billing status.
- An MCP server provides
take_screenshot,get_page_info, andcapture_pdftools for AI agents and MCP clients. - The free plan includes 1,000 screenshots per month with no card. Paid plans start at $5 for 3,000 shots; other monthly tiers are Growth at $15 for 15,000, Pro at $39 for 60,000, Scale at $99 for 250,000, and Business at $249 for 1,000,000. Yearly billing gives two months free, and every feature is on every plan.
Sign up for 1,000 free screenshots a month with no card.
Free tools Windows power users keep installed
One-click scans. No signup required.
When and how to distribute a crawl
Scrapy does not provide built-in multi-server crawling. Its documented patterns are to distribute separate spider runs across Scrapyd instances, or divide one large spider’s URLs into partitions and schedule those partitions on separate servers. In either pattern, define how work is assigned and tracked at the application level; do not assume that adding instances automatically coordinates URL ownership.
Make partitions recoverable
For a production system, persist task state and outputs so a worker restart does not silently lose uncompleted work. Make writes idempotent or deduplicate where repeated tasks are possible, and make retries bounded and observable. These are engineering recommendations for implementing partitioned work; Scrapy’s distribution documentation does not prescribe a particular queue, database, or exactly-once processing design.
Track aggregate target load across every worker. Multiple spiders in one process each have their own concurrency and politeness settings, so their combined activity can exceed what any one spider’s setting suggests. The same principle applies when you add processes or hosts: evaluate total requests to each domain, not only per-worker settings. See Scrapy’s Common Practices documentation for its multi-spider and multi-server examples.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Scale the resource that is actually limiting you
More concurrency mainly adds network work; it does not make every other part of a crawl faster. Scrapy’s optimization guide notes that most work in a process runs in one thread, so CPU-bound crawling can hit a one-core ceiling unless work is moved elsewhere. It also identifies bandwidth, DNS lookups across many domains, scheduler queues, memory, disk writes, callbacks, and item pipelines as possible limits.
Best Value
| Measured constraint | Potential response | Watch for |
|---|---|---|
| CPU-bound callbacks or extraction in one process | Consider splitting work across processes to use more than one CPU core. | Whether CPU, rather than target limits or downloads, is actually saturated. |
| Memory rising alongside a growing scheduler queue | Check whether URL discovery is outpacing downloads; control work creation and make partitioning deliberate if needed. | Queue depth and memory after each change. |
| Response handling or item pipelines falling behind | Investigate callback and pipeline capacity before raising downloader concurrency. | Downloaded responses versus extracted and persisted items. |
| Network or DNS constraints in a broad crawl | Measure bandwidth and DNS behavior across the domain mix. | Whether the resource limit moves when concurrency changes. |
For broad crawls across many domains, a higher global concurrency can coexist with conservative per-domain caps, provided CPU and memory capacity support it. Scrapy’s documentation describes that relationship but does not make any sample configuration a universal setting. Workers can increase throughput only when they address the measured bottleneck; otherwise they add coordination and resource use without necessarily improving useful output.
Troubleshooting common scaling symptoms
- Throughput stays flat after increasing concurrency: Recheck the measured signals. Look for CPU, bandwidth, DNS, request production, callback or pipeline capacity, scheduler growth, or target responses that point to a different constraint.
- 429 or 503 responses rise, or latency worsens: Reduce pressure and inspect the target’s terms, documentation, and robots guidance. Review retries as well as response status; do not keep raising limits to compensate for a target that is signaling strain.
- The scheduler queue keeps growing: Discovery may be producing work faster than the downloader can complete it. Check link-following behavior and memory use before adding workers, since additional workers can increase aggregate target traffic.
- Responses arrive but item output lags: Inspect callback and pipeline processing rather than assuming the downloader is the bottleneck. Compare downloaded responses with extracted and persisted items.
- Workers repeat or miss URLs: Revisit partition ownership and persisted task state. Add deduplication or idempotent output handling where repeats are possible, and make retries visible and bounded.
- A robots.txt delay appears to be ignored: Scrapy does not automatically apply
Crawl-delayorRequest-rateas download settings. Translate applicable target guidance into settings explicitly, while following any stricter terms or API limits.
A practical decision sequence
- Confirm that the target permits the intended access and check for a suitable documented API or export.
- Run a representative crawl with conservative per-domain limits and collect throughput, response, scheduler, and resource signals.
- Identify the limiting stage, then change one relevant setting or resource at a time.
- Keep the change only if useful output improves without an unacceptable increase in target errors, latency, queue growth, or resource use.
- Move to processes or distributed partitions when measurement supports that choice; define ownership, durable state, retry behavior, and deduplication before relying on multiple workers.
For the trade-off, compare a single process, multiple processes on one host, and multiple hosts by which bottleneck each can address, coordination and recovery complexity, aggregate load per target, and compute, storage, network, and engineering effort. No universal cost or performance figure follows from the Scrapy guidance; those outcomes depend on workload and environment.
Frequently Asked Questions
Should concurrency be adjusted while a production crawl is running?
Treat changes as controlled experiments: record the current signals, change one relevant limit, and monitor the target responses and resource use before retaining the new setting.
Does adding a second worker guarantee that a crawl can resume safely?
No. Recovery depends on how the application records task ownership and output; workers need explicit persisted state and handling for repeated tasks.
Can a robots.txt file authorize access that the site otherwise restricts?
No. Robots guidance does not override site terms, authentication or access controls, or applicable law.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

