Free tools Windows power users keep installed
One-click scans. No signup required.
Load balancing a web scraper is three separate engineering problems: distributing crawl work among workers, controlling the total request rate each target receives, and routing traffic through proxy endpoints without breaking cookies or authentication. Solve them independently. Scrapy can run partitions across machines and multiple Scrapyd instances, but it has no built-in multi-server scheduler; concurrency and AutoThrottle settings apply to each crawler, not to your whole fleet.
Start with the constraint you actually need
Before adding workers or proxies, write down the target’s published rules, your authorization, and the result you need. “Scale” can mean different things:
- More worker capacity: the crawl queue or parsing work is CPU- or I/O-bound.
- More request concurrency: you need more in-flight requests, subject to the target’s limits.
- More network egress: requests must leave through several permitted IP addresses or regions.
- Session continuity: a sequence must retain cookies, authentication, CSRF tokens, or other state.
Workers do not automatically create a shared frontier. More IPs do not grant permission for a higher rate, and neither prevents a site from returning a block page. Treat identity, pacing, scheduling, and state as explicit design decisions.
What Scrapy provides—and what it does not
Scrapy’s Common Practices documentation states: Scrapy doesn’t provide any built-in facility for running crawls in a distributed (multi-server) manner.
Its documented patterns are operational rather than a cluster-wide scheduler:
The Tool Desk
Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →#1 Best Overall
- Read Before You Buy — No Video Output: These adapters support charging and USB 2.0 data transfer, but cannot transmit video signals. Except for standard USB webcams (which use USB data only), they are not compatible with HDMI/DisplayPort cables, video-capable USB-C hubs, or docking stations with video output.
- Convert USB-A Ports to USB-C: Designed to connect USB-C earphones, cables, flash drives, card readers, and other USB-C accessories to standard USB-A ports. Plug-and-play with no drivers or software required.
- Aluminum Alloy Housing: Built with a sturdy aluminum alloy shell that aids in heat dissipation and protects against daily wear and scratches. Designed to maintain a stable and secure connection.
- Compact & Travel-Friendly: The ultra-compact design allows the adapter to stay plugged into your device without blocking adjacent ports or adding bulk, reducing wear and tear on your original USB ports.
- 12-Month Warranty: Backed by a 12-month manufacturer warranty for peace of mind. Designed to meet strict quality control standards for reliable everyday performance.
- Run spiders on several Scrapyd instances when you have many independent spider runs.
- For one large spider, partition the input URLs and assign disjoint partitions to runs on different machines.
These approaches leave important coordination to you: partition ownership, duplicate prevention, retries, a shared frontier (if required), and aggregate per-domain rate limiting. The documentation does not claim that separate runs share a scheduler or deduplication store. See the Scrapy distributed-crawls guidance.
One crawler with tuned concurrency
A single crawler keeps one scheduler and local concurrency controls, which is often the simplest reliable architecture. Current Scrapy settings documentation lists a default CONCURRENT_REQUESTS of 16 and a fallback per-domain value of 8; these are version- and project-dependent defaults, not performance recommendations. Check your project’s effective settings before changing them: Scrapy settings.
Multiple runs or Scrapyd instances
Separate processes can improve operational capacity, but every process has its own limits. If five crawlers each allow 16 concurrent requests, the fleet can create roughly 80 in-flight requests before accounting for retries and other traffic. You must enforce a fleet-wide budget yourself.
URL-partitioned workers
Partition by a stable key such as hostname, sitemap shard, customer, or hash range. Make partitions disjoint, record ownership, and design retries so a failed worker can resume without another worker processing the same URLs. If several partitions contain the same domain, their rates still add together.
PC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchMake request rate a fleet-wide budget
Concurrency is not a permission level. Set a target-aware budget per origin, then allocate that budget among workers. If your desired aggregate limit is 20 concurrent requests to one domain and four crawlers may reach it, begin by allocating about five per crawler. This mirrors Scrapy’s guidance to divide settings among concurrent crawlers when keeping combined pressure unchanged: Common Practices.
Account for all sources of pressure
- Count every worker, scheduled spider, health check, retry, redirect, and asset request that can hit the origin.
- Separate limits by hostname when policies or capacity differ.
- Use a shared token bucket, gateway, or coordinator when independent processes must obey one limit.
- Measure response status, latency, connection errors, and bytes; lower the budget when the target shows strain or asks you to slow down.
AutoThrottle is local, not global
AutoThrottle adjusts delay from observed response latency while respecting configured minimum delay and maximum concurrency bounds. It operates inside one crawler; the cited documentation does not describe cluster-wide coordination. Run it in each worker, but still enforce a fleet-level limiter and inspect aggregate metrics. Latency is an approximation: a fast response can be a cache or block page, while a slow response can reflect temporary server load. Details are in the AutoThrottle documentation.
Rank #2
- 5-in-1 USB-C Hub: Experience comprehensive connectivity featuring a Power Delivery input, two USB-A 2.0 ports, a USB-A 3.0 port, and an HDMI port. (Note: The USB-C power delivery input port is only for connecting an external wall charger to power your laptop and cannot power peripheral devices.)
- 90W Pass-Through Charging: Achieve optimal charging with 90W pass-through power to your laptop, supported by a total input of 100W, with the hub reserving 10W for operational efficiency. (Note: Wall charger not included.)
- Quick Data Transfers: Accelerate your productivity with rapid data transfers using a high-speed 5Gbps USB 3.0 port and two 480Mbps USB 2.0 ports.
- 4K HDMI Display: Enhance your visual experience with a hub capable of delivering 4K resolution at 30Hz in both mirror and extend modes. Please note that this hub is compatible with MacBook (macOS 12 and newer), Windows 10 and 11, ChromeOS, and laptops equipped with DP Alt Mode and Power Delivery. Note: This device is not compatible with Linux.
- What You Get: Anker USB-C Hub (5-in-1, 4K HDMI), welcome guide, 18-month warranty, and our friendly customer service.
Proxies: routing, not authorization
A proxy pool gives requests different egress endpoints. Scrapy lists an IP pool (including paid services) and a managed API as possible distribution options. Proxy rotation does not set your permitted request rate, identify you to a site, or make a blocked crawl legitimate. Follow robots directives, terms, applicable law, and any written authorization; Scrapy’s practices guide recommends identifying your crawler so site owners can contact its operator and spacing requests: Common Practices.
Compare proxy endpoints on operational properties
- Location and whether that geography is necessary and permitted.
- Stability, connection limits, authentication method, and failure rate.
- How long an endpoint remains available for a multi-request workflow.
- Provider terms, data handling, and total operating cost.
- Per-target controls: an endpoint that works for one origin may be unsuitable for another.
A managed scraping API can move proxy and request-handling operations to a service. Scrapy names Zyte API as an example, and the project site describes its integration as offering automatic proxy rotation and browser fingerprinting. Compare coverage, pacing control, identity control, data handling, fallback behavior, integration effort, and cost rather than assuming any service removes your compliance obligations: Scrapy project site.
Do these 3 things before closing this tab:
1Scan for outdated or missing drivers - takes under a minute2Repair Windows errors before they cause bigger problems3Fix the driver behind crashes, sound loss and screen glitchesSessions, cookies, and proxy affinity
Do not confuse two kinds of stickiness:
| Concept | What it routes or preserves | Typical owner |
|---|---|---|
| Outbound proxy selection | The network path and apparent egress address used to reach the target | Your scraper or proxy provider |
| Application session | Cookies, authentication tokens, CSRF state, carts, and other workflow data | The target application and your session store |
| Reverse-proxy sticky session | A client’s requests to the same backend, often using an application session identifier | The target’s load balancer |
Apache’s mod_proxy_balancer documentation describes backend session stickiness; that feature routes a target’s client session to a backend and is not an instruction to rotate your scraper’s outbound identity: mod_proxy_balancer.
When to keep one proxy for a workflow
Use a stable endpoint for a sequence that depends on cookies, login state, device context, or a target policy that expects continuity. Store cookies per logical session, not in one global jar shared by unrelated accounts. If the target or service documents that identity may change safely, you can relax affinity; otherwise, changing proxies mid-flow can cause reauthentication, fraud checks, or inconsistent state.
When rotation is appropriate
Rotate only for a documented, permitted reason such as distributing independent requests across approved egress locations. Never use rotation to evade a block, rate limit, CAPTCHA, or access control. Keep the session-to-endpoint mapping explicit so a retry does not silently move an authenticated workflow to a different identity.
Designing a restartable crawl
Scrapy jobs can persist scheduler state and resume after a clean stop. The job directory requires the same security care as project source code because it can contain URLs, cookies, and other sensitive state. Scrapy also requires a compatible (in practice, the same) Scrapy version when resuming and warns that cookies may expire while a job is paused: Jobs documentation.
Rank #3
- Sleek 7-in-1 USB-C Hub: Features an HDMI port, two USB-A 3.0 ports, and a USB-C data port, each providing 5Gbps transfer speeds. It also includes a USB-C PD input port for charging up to 100W and dual SD and TF card slots, all in a compact design.
- Flawless 4K@60Hz Video with HDMI: Delivers exceptional clarity and smoothness with its 4K@60Hz HDMI port, making it ideal for high-definition presentations and entertainment. (Note: Only the HDMI port supports video projection; the USB-C port is for data transfer only.)
- Double Up on Efficiency: The two USB-A 3.0 ports and a USB-C port support a fast 5Gbps data rate, significantly boosting your transfer speeds and improving productivity.
- Fast and Reliable 85W Charging: Offers high-capacity, speedy charging for laptops up to 85W, so you spend less time tethered to an outlet and more time being productive.
- What You Get: Anker USB-C Hub (7-in-1), welcome guide, 18-month warranty, and our friendly customer service.
- Stop workers cleanly so pending requests and state are flushed.
- Protect job directories and session material with filesystem permissions and secret management.
- Pin the Scrapy and project versions used to resume.
- On restart, validate authentication and discard expired sessions deliberately rather than replaying stale cookies indefinitely.
- Record partition ownership, retry counts, and last successful checkpoints outside the worker process.
Queue persistence answers “which URLs remain?” It does not guarantee that a login, cookie, proxy lease, or target-side session is still valid.
A practical multi-worker blueprint
- Inventory targets: group URLs by origin and document each site’s limits and permission.
- Choose partition keys: create disjoint, reproducible URL ranges and a durable ownership record.
- Set an aggregate budget: define requests, concurrency, and bandwidth per origin; divide it among active workers.
- Add a shared limiter: use a central service or gateway when workers cannot coordinate locally.
- Map sessions: bind each logical session to its cookie jar and, where required, a stable proxy endpoint.
- Instrument outcomes: collect status classes, latency, retries, proxy errors, block indicators, and queue depth.
- Handle shutdowns: stop cleanly, persist state, verify version compatibility, and revalidate sessions before resuming.
- Review behavior: reduce load when error rates or target signals rise; scaling out is not a substitute for responsible pacing.
Troubleshooting common failures
“Workers duplicate URLs”
Your partitions overlap or retries are not idempotent. Generate partitions from a deterministic key, record ownership centrally, and use a shared deduplication strategy if the crawl requires global uniqueness.
“The target sees too much traffic”
Per-crawler limits were summed incorrectly. Count all processes and retries, then lower each worker’s allocation or enforce a shared token bucket. AutoThrottle in one process cannot see traffic from its peers.
“Login breaks after proxy rotation”
The workflow is stateful. Keep the cookie jar and proxy mapping together, or restart authentication intentionally. Do not assume a new IP is interchangeable with the previous session.
“A resumed job gets unauthorized responses”
Cookies or tokens expired during the pause. Reauthenticate through the documented flow, refresh the session store securely, and avoid replaying stale credentials.
“Proxy failures look like target blocks”
Separate connect timeouts, DNS errors, TLS failures, proxy-authentication errors, and origin HTTP responses in metrics. Test the endpoint independently and remove unstable addresses instead of increasing concurrency.
Rank #4
- Dual Converters, Infinite Potential:Includes 2× USB C male to USB A female adapters and 2× USB A male to USB C female adapters. Perfect for a wide range of uses—tablets with Bluetooth keyboards, expand USB ports on macbook, and more. Two different converters for all your daily needs
- Next-Level 10Gbps & 3A Charging: No more slow 480Mbps, this usb to usb c adapter has a transfer speed of up to 10Gbps, allowing you to do more transferring in less time. This usb adapter fits both USB A and USB C charger, supporting up to 3A fast charging
- Upgraded Exquisite Craftsmanship: With an aluminum alloy housing and metal connector, the usbc to usb adapter is extremely durable and sturdy. Rigorously tested to withstand more than 10,000 times of plugging and unplugging, ensuring long-lasting performance
- Broad Compatible: The usb c to usb adapter widely supports all USB C/ USB A devices like laptops, tablets, cellphones, car chargers, and phone chargers. Such as compatible with MacBook Pro/Air 2023/2022, Thunderbolt 4/3 Devices,Apple MagSafe Watch 9/8/7/SE/Ultra, iPad Pro 2022/2021, Samsung Galaxy S23/S20/S10, and iPhone 17/16/15 Pro. Plug and play
- Please Note: To reach 10Gbps speed, keep the cable under 3.3 ft. For USB A Male to USB C adapters, try flipping the USB C connector. USB C Male to USB A adapters support bidirectional 10Gbps transfer within 3.3 ft
“A worker cannot resume its job”
Check for an unclean shutdown, a changed Scrapy version, inaccessible job-directory permissions, or corrupted state. Restore the compatible environment and restart from a verified checkpoint.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Or skip the browser setup
If your task is collecting page images or PDFs rather than crawling links, ScreenshotNeo provides a website screenshot API and MCP server. It accepts consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, with the result identified by X-Page-Verdict and X-Billed headers. Its MCP tools—take_screenshot, get_page_info, and capture_pdf—work with Claude, Cursor, and other MCP clients.
One request returns PNG, JPEG, WebP, or PDF. The service supports full-page and element capture, device and viewport settings, dark mode, retina scale, PDF paper and page controls, custom CSS and JavaScript, clicks, waits, blocking rules, headers, cookies, user agents, authorization, timezone, geolocation, transparent backgrounds, resizing, chosen cache TTLs, signed image links, asynchronous webhooks, bulk capture of up to 100 URLs per call, usage reporting, and an OpenAPI specification. Parameter names used by other screenshot APIs also work.
Using the ScreenshotNeo documentation:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests; r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90); open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' }); const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.
Best Value
- 5-in-1 Connectivity: Equipped with a 4K HDMI port, a 5 Gbps USB-C data port, two 5 Gbps USB-A ports, and a USB C 100W PD-IN port. Note: The USB C 100W PD-IN port supports only charging and does not support data transfer devices such as headphones or speakers.
- Powerful Pass-Through Charging: Supports up to 85W pass-through charging so you can power up your laptop while you use the hub. Note: Pass-through charging requires a charger (not included). Note: To achieve full power for iPad, we recommend using a 45W wall charger.
- Transfer Files in Seconds: Move files to and from your laptop at speeds of up to 5 Gbps via the USB-C and USB-A data ports. Note: The USB C 5Gbps Data port does not support video output.
- HD Display: Connect to the HDMI port to stream or mirror content to an external monitor in resolutions of up to 4K@30Hz. Note: The USB-C ports do not support video output.
- What You Get: Anker 332 USB-C Hub (5-in-1), welcome guide, our worry-free 18-month warranty, and friendly customer service.
Cost, performance, and reliability trade-offs
| Choice | Benefit | Cost or risk |
|---|---|---|
| Single tuned crawler | Simple scheduler and observability | One process and host may become the bottleneck |
| Several workers | More CPU, memory, and failure isolation | Limits, deduplication, retries, and scheduling must be coordinated |
| Proxy pool | Multiple permitted egress paths | Endpoint fees, instability, authentication, and session complexity |
| Managed API | Less proxy infrastructure to operate | Service cost, integration constraints, data handling, and dependency risk |
Benchmark the complete workflow—fetch, parse, storage, retries, and coordination—not just HTTP throughput. A faster worker that causes more retries or blocks is slower in useful output. Keep a conservative operating mode and a way to drain queues before changing limits.
FAQ
Does adding proxies increase the rate a site permits?
No. Proxies change routing; permission and target-specific pacing remain your responsibility.
Should every worker have its own cookie jar?
Use one jar per logical application session. Separate jars prevent unrelated accounts or workflows from sharing state; a worker may host several jars.
Quick wins for a faster PC:
Scan for outdated or missing drivers - takes under a minuteDriver Scan →Clear out junk files and repair common Windows errorsFree Scan →Is a shared queue mandatory?
No. Disjoint URL partitions can work without one, but a shared frontier, global deduplication, or dynamic reassignment requires an additional coordination layer.
Can AutoThrottle enforce a fleet-wide limit?
No. It adapts within each crawler. Combine it with a limiter that sees all workers when aggregate control matters.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




