Scalable automated data collection starts with the least costly source that meets your needs: use a supported API, export, or search endpoint when available; crawl pages only when necessary. For a crawl, scale by partitioning known URLs, bounding concurrency per host, reacting to access signals, and saving durable outputs so collection and downstream processing can recover independently.
Choose the right way to access the data
Before building a crawler, look for a documented API, bulk export, or search endpoint. These interfaces can be faster for your collector and cheaper for the website than retrieving pages one by one. Read the interface documentation, terms, and rate limits before scheduling requests. Scrapy’s optimization guidance discusses this source-selection principle: Scrapy: best practices.
If page crawling is required, find a sitemap or another permitted source of URL lists. A list of known URLs reduces serial discovery and gives the scheduler enough work to distribute early. Choose based on what the source actually exposes, how fresh the data must be, whether important content depends on JavaScript, the crawl’s size and cadence, and what access limits apply. There is no universally best interface or crawl design.
Design the work so it can be distributed
A scalable collector needs more than additional workers. It needs an explicit source of work, a way to avoid duplicate URLs, bounded request concurrency, retry and recovery behavior, and durable outputs. Separate URL discovery, fetching, parsing, and storage where that makes retries or downstream processing easier.
#1 Best Overall
Partition URLs, not just processes
For a large crawl spread across machines, prepare URL partitions and assign each partition to a separate spider run or worker group. Scrapy’s documented approach is to divide URLs among runs; Scrapy itself does not provide a built-in distributed, multi-server crawling facility. You must provide the coordination and partitioning layer around the framework. See the Scrapy documentation on running multiple spiders and distributed crawling (Scrapy 2.19.0 documentation).
Make partitions non-overlapping where practical, but keep a deduplication mechanism because URLs can recur through redirects, links, or retries. Record each URL’s state—queued, fetched, failed, or intentionally skipped—so a worker failure does not require restarting the whole crawl. Treat a partition as recoverable work, not as an all-or-nothing job.
Account for aggregate load
Running multiple spiders in one process applies concurrency and politeness settings separately to each crawler. Scrapy advises dividing those settings by the number of simultaneous crawlers when the goal is to keep combined load unchanged. Re-running the same spider with the same per-spider concurrency does not create free capacity: it multiplies aggregate requests to the site. Confirm the applicable settings and behavior in the Scrapy guidance.
Rank #2
- Wiley
- Language: english
- Book - storytelling with data: a data visualization guide for business professionals
Set a crawl rate from site signals, not guesswork
There is no single safe request rate for every website. AWS Prescriptive Guidance gives context-dependent examples: one request every 10–15 seconds may suit small or medium websites, while 1–2 requests per second may suit larger sites or crawls with explicit permission. Those are operational recommendations, not measured universal thresholds. Start conservatively, increase gradually only when justified, and monitor the site’s responses as well as your own queue and retry metrics. See AWS Prescriptive Guidance: web crawling at scale.
Free tools Windows power users keep installed
One-click scans. No signup required.
- Watch for 429 and 503 responses, rising retry counts, ban pages, and increasing download latency. Scrapy recommends gradually raising concurrency while monitoring these signals.
- Pause after a 429 (“Too many requests”). If 403 (“Forbidden”) responses continue, consider stopping rather than attempting to work around the restriction.
- Identify the collector in its User-Agent and honor a site owner’s request to stop.
- Do not assume a crawler automatically applies every robots.txt directive. Scrapy documents that it does not automatically translate robots.txt
Crawl-delayandRequest-ratedirectives into downloader delay and concurrency settings; configure the applicable controls yourself.
Scrapy’s crawler settings are described at Scrapy settings. Confirm current setting names and effects against the version you deploy.
Make collection responsible and recoverable
Check the site’s robots.txt rules for your crawler’s user agent, terms of service, privacy policy, and applicable legal restrictions. Robots.txt is an important operational signal, not a complete legal determination. AWS recommends polite rates, identifying your crawler, using sitemaps where available, batching work, pausing after 429 responses, considering stopping after repeated 403 responses, and stopping if the site owner asks. These are operational practices, not legal advice. AWS’s guidance states: “Always check and respect the rules in the robots.txt file.” Read its full crawling guidance.
Rank #3
Batch work so that a problem is contained to a manageable slice. Persist parsed records and, where useful, raw documents or response metadata. Durable outputs let ingestion and analysis proceed separately from fetching, and let you retry failed processing without fetching the same pages again. Apply access controls to collected data that reflect its sensitivity and the permissions under which it was obtained.
Connect the crawler to storage and processing
One AWS reference architecture uses EventBridge Scheduler to start jobs, AWS Batch to orchestrate them, crawler jobs in ECS containers on Fargate, and Amazon S3 for retrieved records and raw documents. Downstream applications then ingest or process the stored files. This is one implementation example, not a requirement; workload size, latency, budget, existing infrastructure, and operational skills should determine your design. See AWS’s architecture description.
A managed connector may fit a narrower use case. AWS’s Bedrock web-crawler connector documents seed URL scope, per-host crawl-rate limits, page-count limits, include and exclude patterns, and incremental synchronization. Its documentation says it should be used only for websites you own or are authorized to crawl. It supports static web pages, so verify that limitation against your need for dynamic, JavaScript-dependent content before choosing it. Details are in AWS Bedrock web-crawler connector documentation.
Rank #4
Capture rendered pages when collection needs screenshots
Some pipelines need a visual record of a page rather than only extracted text or structured records. A browser-based capture can be one component of that workflow, but it does not replace source permission, sensible request pacing, URL partitioning, or durable storage. If a page’s relevant content is rendered dynamically, check that the capture method waits for the content you need and test representative pages before scaling.
ScreenshotNeo is a website screenshot API and MCP server for developers. Its one-request API returns an image or PDF, and its available capture options include full-page screenshots, CSS selector capture, viewport and device settings, waiting for a selector or network idle, and custom CSS or JavaScript. For a collection pipeline, use a queue and controlled workers around capture requests rather than firing an unbounded burst. ScreenshotNeo does not remove the need to respect each target site’s rules.
Or skip the browser setup
A single GET request can return a screenshot; the code below uses Stripe as the example target. See the ScreenshotNeo API documentation for request options.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Best Value
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
ScreenshotNeo accepts cookie or consent banners as a visitor and removes 60+ known consent platforms, newsletter popups, and chat widgets before capture; each cleanup step can be turned off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers report the page verdict and whether the request was billed. Its MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients. The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000. Sign up for 1,000 free screenshots a month, with no card required.
Troubleshoot a crawl that is not scaling cleanly
- More workers produce more errors: Check aggregate per-host concurrency across every spider and machine. Reduce the combined rate, observe latency and status codes, and partition work instead of blindly multiplying identical workers.
- The crawl repeatedly rediscovers URLs: Seed from a sitemap or permitted URL list, deduplicate before enqueueing, and persist URL state across worker restarts.
- Retries keep increasing: Inspect 429/503 responses, timeouts, and latency. Slow down or pause; do not treat retries as proof that more concurrency is needed.
- 403 responses continue: Consider stopping and check whether the site allows the activity. Do not attempt to evade access controls.
- Some expected content is missing: Determine whether it is static or JavaScript-rendered, whether your crawl waits for it, and whether a documented API or export provides it more reliably.
- A worker failure loses progress: Persist queue and result state and make each partition resumable; keep fetched data separate from later parsing or ingestion work.
Estimate performance and cost without assuming a universal formula
Throughput depends on the source’s permitted rate, response latency, page size, rendering needs, retry frequency, and the work done after fetching. Raising concurrency can shorten elapsed time only while the site and your own infrastructure can handle the resulting load. It can also increase failures, retries, and operational cost. Measure a small permitted batch first, then use observed latency, response codes, and processing time to plan capacity.
For a hosted implementation, account for scheduling, compute, storage, transfer, and downstream processing separately; the AWS architecture is an example, not a price quote. For any source, keep raw and processed data retention aligned with the actual recovery and analysis needs, and avoid collecting fields or pages that the use case does not require.
Quick Recap
Questions to settle before implementation
- Is there a supported API, export, or search endpoint, and what limits and terms apply?
- How many URLs are in scope, how often must they be refreshed, and what latency is acceptable?
- Is content static or browser-rendered, and which readiness condition indicates the data is complete?
- How will work be partitioned, deduplicated, retried, and resumed?
- What per-host rate is permitted, and which responses or latency changes trigger a pause?
- Where will raw and normalized data live, who can access it, and which downstream jobs consume it?
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

