A scraper that slows, stalls, or returns worse data around request 10,000 has reached a bottleneck in that particular crawl—not a universal failure threshold. The cause may be target-site throttling, a concurrency or delay setting, a spider that cannot produce requests quickly enough, response-processing backpressure, or local CPU and memory pressure. Before increasing concurrency, compare status codes, latency, queue sizes, downloader activity, retries, CPU, and memory to identify where work is backing up.
What “fails after 10,000 requests” actually means
The request count is a useful point to investigate, not a known limit built into scrapers. Scrapy’s official optimization guidance offers operational signals and tuning advice; it does not establish a general point at which crawlers fail. A crawl can handle far more requests—or fail much earlier—depending on the target’s rules, request pattern, response size, processing work, and machine resources. See the Scrapy optimization guide.
First make the symptom specific. “Fails” could mean throughput falls, the process exits, memory rises until the host runs out, results stop appearing, pagination ends early, HTTP errors increase, or returned records become stale or malformed. Each points to a different part of the system. A rising request count alone does not tell you which one.
Identify where the crawl is slowing down
Use signals together rather than treating one metric as a diagnosis. In particular, compare the scheduler’s queued work with active downloader requests, then inspect response codes, retries, latency, callback and pipeline activity, CPU, and memory over the same period.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
#1 Best Overall
| What you observe | Likely area to inspect | What to check next |
|---|---|---|
| More 429 or 503 responses, ban pages, retries, or rising download latency as concurrency rises | Target-site throttling or blocking | Slow down, inspect response bodies and the site’s published access rules, and compare behavior at a lower request rate. |
| Scheduler has queued requests, but downloader activity is below its global limit | Per-domain concurrency, download delay, or AutoThrottle | Check the effective settings and whether the queued requests target the same domain. |
| Scheduler and downloader are nearly empty | Request generation | Look for slow callbacks, sequential pagination, or logic that is not discovering the next requests. |
| Queues grow without settling, CPU is saturated, or memory continually increases | Response processing or local resource pressure | Profile callbacks and pipelines, examine response sizes and memory trends, and check whether work is accumulating faster than it can be processed. |
These are diagnostic clues, not proof by themselves. A busy downloader may simply be waiting on a slow network, and high CPU may be caused by work other than parsing. Use time-series data and a controlled change to test your leading explanation.
Target-site throttling and blocking
When the target starts returning more 429 or 503 responses, ban pages, or slow responses as concurrency increases, raising concurrency further can worsen throughput and reliability. Scrapy’s optimization guidance identifies those patterns, plus growing retry counts and download latency, as signs that the crawl may have exceeded what the target currently tolerates. Inspect status counts and response bodies over time, not just the final total.
Check the site’s terms and published access options before changing request behavior. Scrapy does not automatically apply robots.txt Crawl-delay and Request-rate directives to its settings; its documentation says to translate applicable directives into download-delay and concurrency settings. Respect those rules and any documented rate limits. IP or proxy rotation is not a general-purpose answer to a target’s limits.
Concurrency, download delay, and AutoThrottle
In Scrapy, several controls can limit how quickly requests reach a domain:
CONCURRENT_REQUESTSlimits simultaneous downloads globally.CONCURRENT_REQUESTS_PER_DOMAINlimits simultaneous requests to the same domain.DOWNLOAD_DELAYsets a minimum interval between requests to a domain.
If work is queued while downloader activity stays below the global cap, inspect the per-domain limit, delay, and whether AutoThrottle is pacing requests. A global concurrency increase will not remove a stricter per-domain limit or delay.
What AutoThrottle does—and does not do
AutoThrottle adjusts per-site delays using response latency, aiming for a configured average concurrency. That target is a goal, not a hard limit, and the regular concurrency and delay settings still apply. Its design avoids reducing delay based on fast non-200 responses, because an error response can result from an excessive request rate. Read the AutoThrottle documentation before changing its settings.
Tune one control at a time and watch the target’s responses. If errors or latency rise, back off. Do not treat a higher concurrency number as success if it produces fewer useful records per hour or more failed requests.
When the spider cannot produce work quickly enough
If both the scheduler and downloader are nearly empty, the crawler may not be generating requests fast enough to use its available download capacity. A spider that discovers each page only after downloading and processing the previous one—for example, strictly sequential pagination—cannot gain much from adding concurrent downloader slots.
The Tool Desk
Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Rank #3
Review whether independent pages can be discovered earlier without violating the site’s allowed rate or breaking required ordering. Check callback logic for slow parsing, unnecessary serial steps, and paths that stop following pagination. Do not parallelize requests if the next request depends on data from the previous response or if doing so would exceed the target’s rules.
Response processing, queues, CPU, and memory
Downloads are only one part of a crawl. If responses arrive faster than callbacks or item pipelines can process them, processing becomes the bottleneck and the framework can apply backpressure. A scheduler queue that grows without settling means requests are being discovered faster than they are downloaded; over a long crawl, that backlog can contribute to memory exhaustion.
Scrapy’s optimization guidance notes that it runs in one process and that, apart from DNS and work explicitly moved to a thread, most work runs in one thread. One CPU core can therefore be the ceiling for CPU-bound work. Profile CPU use and watch memory for leaks or unbounded growth. Raising downloader concurrency does not fix a CPU-bound selector or a slow pipeline; it can instead increase queued responses and memory pressure. See the optimization guidance for its profiling and resource recommendations.
Check whether response bodies are unexpectedly large, whether callbacks retain objects, and whether pipeline work blocks progress. Compare queue size, CPU, and memory over time: a temporary increase during a burst differs from a queue or memory line that keeps climbing throughout the run.
Retry amplification can make a crawl look stuck
Retries consume downloader capacity, especially when a site is slow or failing. Scrapy’s version 2.7.1 documentation on broad crawls warns that repeated timeout retries can substantially slow broad crawls and prevent capacity from being reused for other domains. The point applies to broad crawls; it is not a universal prescription for a particular retry setting.
Inspect which failures are being retried and how often. Set retry behavior to fit the failure type and crawl shape. Increasing retries indiscriminately may extend a stall, while disabling retries entirely may discard transiently recoverable work. Keep enough logging to distinguish an original failure from a retry that later succeeded.
A practical diagnostic sequence
- Define the failure. Record whether you are seeing lower throughput, process exit, rising memory, empty output, incomplete pagination, HTTP errors, ban pages, or malformed data.
- Graph status, retries, and latency. Look for changes that coincide with rising concurrency or crawl volume. A rise in 429/503 responses, ban pages, retries, or latency is a reason to reduce pressure and re-check the site’s rules.
- Compare scheduler and downloader activity. Queued work with underused downloader slots suggests a per-domain cap, delay, or AutoThrottle. Nearly empty queues suggest request generation may be limiting throughput.
- Inspect processing and host resources. Check callback and pipeline workload, response sizes, CPU use, and memory trends. Look for queues that grow without settling.
- Change one setting gradually. Make a small, reversible adjustment, then observe latency, errors, useful output, and resource use. If the site’s responses worsen, back off rather than pushing the rate higher.
- Check published ways to access the data. If the site offers an authorized API, bulk export, or documented search endpoint, review its terms and stated rate. Scrapy’s optimization guidance notes that these methods can be faster for the crawler and cheaper for the target than crawling pages.
Or skip the browser setup
If the task is to capture pages as screenshots or PDFs rather than extract structured records, ScreenshotNeo offers a website screenshot API and MCP server. It accepts a URL in one GET request and can return PNG, JPEG, WebP, or PDF. Its clean-shot steps accept cookie or consent banners like a visitor and remove 60+ known consent platforms, newsletter popups, and chat widgets; each step can be turned off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers identify the page verdict and billing status. Its MCP server provides take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients.
For example, install the Python requests package, put your key in place of YOUR_API_KEY, and run:
Do these 3 things before closing this tab:
1Clear out junk files and repair common Windows errors2Fix the driver behind crashes, sound loss and screen glitches3Repair Windows errors before they cause bigger problemsimport requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
open("shot.webp", "wb").write(r.content)
See the ScreenshotNeo API documentation for request options. The same request as cURL:
Best Value
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Or in Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
Use an API key you control and avoid putting it in public client-side code. The service also supports full-page captures with lazy images loaded, CSS-selector element capture, dark mode, 12 device presets and custom viewports, retina scale, PDF paper size/margins/landscape/page ranges, HTML/CSS capture, custom CSS and JavaScript, pre-capture clicks, hidden selectors, selector/delay/network-idle waits, request and resource blocking, custom headers/cookies/user agent/Authorization, timezone and geolocation, transparent backgrounds, resizing, configurable-TTL caching, signed image links, asynchronous jobs with signed webhooks, bulk capture of 100 URLs per call, a usage API, and an OpenAPI specification. Screenshot API parameter names used by other providers also work to make migration easier.
Free includes 1,000 screenshots per month with no card. Paid plans are Starter at $5 for 3,000, Growth at $15 for 15,000, Pro at $39 for 60,000, Scale at $99 for 250,000, and Business at $249 for 1,000,000; yearly billing gives two months free. Every feature is available on every plan. Sign up for ScreenshotNeo’s free plan to get 1,000 screenshots a month with no card.
Common troubleshooting cases
- 429 or 503 counts rise after a concurrency change: Reduce request pressure, check the target’s published limits and terms, and monitor whether latency and error rates recover.
- Requests wait while downloader slots are idle: Check per-domain concurrency, download delay, and AutoThrottle; the global cap may not be the limiting setting.
- The downloader goes idle and output stops: Inspect pagination and callback logic, and confirm that the spider is still producing requests rather than waiting on a serial dependency.
- The queue or memory grows throughout the crawl: Look for discovery outpacing downloads, responses outpacing processing, retained objects, or slow callbacks and pipelines. Do not add concurrency until the backlog’s cause is understood.
- Broad crawls spend time retrying timeouts: Inspect retry counts and the affected domains. Repeated retries can tie up capacity; tune retry behavior to the failure and crawl shape.
- Records are missing even though requests complete: Distinguish transport success from extraction success. Check response bodies, parsing errors, pagination coverage, and whether a target changed its page structure.
Frequently asked questions
Does Scrapy stop after 10,000 requests?
No general 10,000-request cutoff is established in the cited Scrapy guidance. Treat that count as the point where a bottleneck became visible in your workload, then diagnose its signals.
Should I raise concurrency to make a slow crawl faster?
Only if measurements indicate downloader capacity is the limiting factor and the target continues to respond acceptably. If latency, 429/503 responses, retries, or local resource pressure are increasing, more concurrency can make the crawl less reliable.
Is an empty downloader proof that a site blocked me?
No. Nearly empty scheduler and downloader queues can mean the spider is not producing requests quickly enough. Check request-generation and pagination logic alongside status codes and response bodies.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

