What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
Rotate the network route only when your workflow allows it. For independent page requests, choose a proxy from a vetted pool, send the request, record the outcome, and update that proxy’s health. For a login, cart, pagination chain, or any other stateful flow, keep one proxy for the relevant session so the site sees a consistent route. There is no universal “rotate every request” interval. Your legal permission, the target’s published limits, and observed status codes and latency should determine the cadence.
A proxy changes where a request exits the internet; it does not grant permission to collect data or bypass a site’s controls. Check the site’s robots.txt, API, export options, and terms first, identify your crawler honestly, and stop when the target signals that your traffic is unwelcome.
What proxy rotation actually changes
A proxy sits between your scraper and the target. The target normally sees the proxy’s address rather than the machine running your code. Rotation means selecting a different configured proxy for a later request. It does not make an unauthorized crawl acceptable, defeat a CAPTCHA, or guarantee that a request will succeed.
Independent requests
If each URL can be fetched without cookies or earlier page state, different requests may use different proxies. A product catalog where every page is publicly addressable is a typical example. Keep enough information to diagnose the crawl: proxy identifier (not its password), URL, timestamp, status code, elapsed time, and whether the response was a normal page or a block page.
#1 Best Overall
Stateful requests
Login, checkout, consent, multi-step forms, and cursor-based APIs depend on cookies, authentication, or a server-side session. Pin those requests to one proxy for the life of the session unless the site’s documentation explicitly says otherwise. Changing the route mid-flow can invalidate the session or look anomalous. You can start a new session on another proxy after the current unit of work finishes.
Permission and a responsible crawl plan
- Prefer a supported interface. Look for an API, bulk export, feed, or search endpoint before HTML crawling. A supported interface is usually more stable and easier for the site to capacity-plan.
- Read robots.txt and published limits. Scrapy does not automatically enforce robots.txt
Crawl-delayorRequest-rate; translate applicable directives into your own delay and concurrency settings. - Identify yourself. Where crawling is allowed, use a
USER_AGENTthat names your project and provides a contact route. Honest identification gives an owner a way to ask you to adjust the crawler. - Set a stop condition. Define maximum error rates, latency, and daily volume before starting. A rise in 429/503 responses, ban pages, retries, or download time is evidence to slow down or stop—not a signal to add more proxies blindly.
Scrapy’s guidance offers “2 seconds apart or more” as a practice suggestion in its anti-ban discussion. Treat that as context for a particular site and crawl, not a universal quota.
Rotating proxies with Python Requests
Install and represent the pool safely
Proxy URLs include a scheme. Keep credentials in a secret manager or protected environment, never in source control or ordinary logs. Requests documents optional SOCKS support through requests[socks]; socks5 resolves DNS on the client, while socks5h resolves through the proxy.
python -m pip install requests
# Only if you need SOCKS:
python -m pip install "requests[socks]"
Use a mapping with both keys when a request may redirect between schemes:
proxies = {
"http": "http://user:[email protected]:8000",
"https": "http://user:[email protected]:8000",
}
Do not assume an HTTPS proxy URL means the proxy itself speaks HTTPS; the scheme describes how Requests connects to that proxy. Follow the provider’s format.
Choose a proxy per request
import os
import random
import time
import requests
PROXIES = [
os.environ["PROXY_ONE"],
os.environ["PROXY_TWO"],
]
def proxy_mapping(proxy_url):
return {"http": proxy_url, "https": proxy_url}
def fetch(url):
proxy = random.choice(PROXIES)
started = time.monotonic()
try:
# Pass proxies explicitly so environment settings do not silently win.
response = requests.get(
url,
proxies=proxy_mapping(proxy),
headers={"User-Agent": "ExampleResearchBot/1.0 (+mailto:[email protected])"},
timeout=(10, 45),
)
elapsed = time.monotonic() - started
return {
"url": url,
"proxy_id": proxy.split("@")[-1],
"status": response.status_code,
"seconds": round(elapsed, 3),
"body": response.text,
}
except requests.RequestException as exc:
return {"url": url, "proxy_id": proxy.split("@")[-1], "error": type(exc).__name__}
result = fetch("https://example.com/page")
print({k: v for k, v in result.items() if k != "body"})
This is a selection pattern, not a promise that rotation prevents blocking. In production, maintain per-proxy health: temporarily quarantine connection failures, distinguish a target 429 from a dead proxy, and re-test a quarantined route after a delay. Never log the complete credential-bearing URL.
Use a Session when state should persist
A Session reuses cookies and other configuration. Set its proxy mapping once for a related sequence:
import os
import requests
proxy = os.environ["SESSION_PROXY"]
with requests.Session() as session:
session.proxies.update({"http": proxy, "https": proxy})
session.headers["User-Agent"] = "ExampleResearchBot/1.0 (+mailto:[email protected])"
first = session.get("https://example.com/start", timeout=45)
second = session.get("https://example.com/next", timeout=45)
print(first.status_code, second.status_code)
Requests warns that environment proxy variables can override session settings. If route identity matters, pass proxies=... on each call or verify the effective configuration in your deployment. A per-call mapping is also useful when one exceptional request needs a different route.
Rank #3
Rotating proxies with Scrapy
Built-in pacing and concurrency settings
Scrapy separates concurrency from delay. CONCURRENT_REQUESTS_PER_DOMAIN caps simultaneous requests to one domain; DOWNLOAD_DELAY sets a minimum interval between consecutive requests to that domain. CONCURRENT_REQUESTS is the global cap. Raising concurrency can increase throttling, errors, and total crawl time if the target cannot tolerate it.
# settings.py
BOT_NAME = "permitted_crawler"
USER_AGENT = "PermittedCrawler/1.0 (+mailto:[email protected])"
CONCURRENT_REQUESTS = 8
CONCURRENT_REQUESTS_PER_DOMAIN = 2
DOWNLOAD_DELAY = 2
AUTOTHROTTLE_ENABLED = True
AUTOTHROTTLE_START_DELAY = 2
AUTOTHROTTLE_MAX_DELAY = 60
AUTOTHROTTLE_TARGET_CONCURRENCY = 1.0
RETRY_HTTP_CODES = [408, 429, 500, 502, 503, 504]
These values are starting points, not a license or a guaranteed safe profile. Map any applicable robots.txt delay or request-rate directive yourself, then increase load gradually while watching the target’s responses.
Assigning a proxy in middleware
A simple downloader middleware can assign a route to each request. Keep credentials outside settings committed to a repository:
Quick wins for a faster PC:
Scan for outdated or missing drivers - takes under a minuteDriver Scan →Clear out junk files and repair common Windows errorsFree Scan →# middlewares.py
import os
import random
class ProxyPoolMiddleware:
def __init__(self):
self.proxies = [
os.environ["PROXY_ONE"],
os.environ["PROXY_TWO"],
]
@classmethod
def from_crawler(cls, crawler):
return cls()
def process_request(self, request, spider):
# For stateful flows, set request.meta["proxy"] once and copy it onward.
request.meta["proxy"] = random.choice(self.proxies)
# settings.py
DOWNLOADER_MIDDLEWARES = {
"myproject.middlewares.ProxyPoolMiddleware": 543,
}
For a stateful chain, choose the proxy in the first request and propagate request.meta["proxy"] to every follow-up request. For independent pages, the middleware may choose separately, but add health tracking and a site-specific ban test rather than rotating mechanically.
About scrapy-rotating-proxies
The scrapy-rotating-proxies extension tracks working and non-working proxies, periodically checks failed routes, supports a configurable ban policy, and offers per-proxy concurrency controls. Its documentation says you must supply the proxy list and appropriate ban rules. Its default budget is five proxy attempts; that is the package default, not a generally safe retry count. The documentation’s release history lists 0.6.2 from 2019, so verify compatibility with your installed Scrapy version before adopting it.
How to choose a rotation cadence
| Workflow | Route strategy | What to monitor |
|---|---|---|
| Independent public pages | Different vetted proxy when useful; no fixed interval | 429/503 rate, ban-page signatures, latency, connection failures |
| Pagination or API cursor | Keep one proxy for the cursor/session | Cookie validity, authentication errors, duplicate or missing pages |
| Login or checkout test | Pin one route until the flow ends | Session continuity and explicit site limits |
| Target publishes a delay | Honor that delay regardless of pool size | Observed responses after gradual changes |
A larger pool does not justify a higher aggregate request rate. Calculate concurrency and delay across all workers and domains, not just per process. If errors rise, first reduce concurrency, increase delay, and pause retries; only then investigate whether a particular proxy is unhealthy.
What to do after a 429, 503, or block page
- Confirm the signal. Save status, headers, a short body sample, and timing. A 503 can be an overloaded origin, while a branded challenge page may be a deliberate block.
- Stop adding pressure. Pause the affected domain, reduce concurrency, and lengthen the delay. Do not launch more workers or rotate faster.
- Check your policy. Re-read robots.txt, rate-limit documentation, and any API terms. Contact the owner if your identified crawler is permitted but being misclassified.
- Remove bad routes. Quarantine proxies with connection or TLS failures. Do not quarantine every proxy merely because the target returned one transient 503.
- Resume cautiously. Restart at a lower rate and compare status, retry count, and latency. If the target continues to reject traffic, stop the crawl and use an approved API, export, Common Crawl, or managed service instead.
Self-managed pool or managed scraping API?
| Question | Self-managed proxies | Managed API |
|---|---|---|
| Control | Direct control of route, headers, cookies, and raw responses | Provider controls more infrastructure; review supported options |
| Maintenance | You source, test, quarantine, and replace proxies | Less route maintenance, but you depend on provider behavior |
| Output | Best when you need the original response and custom parsing | Useful when you want parsed data or managed retries |
| Sessions and geography | You implement stickiness and select locations | Availability depends on the service’s documented controls |
| Cost and evidence | Compare your engineering time and actual volume | Compare the provider’s current terms at your workload; no universal performance or price winner is established here |
Scrapy’s documentation gives Zyte API (with a Scrapy plugin) and ProxyMesh as examples, not endorsements. Evaluate current compatibility, retention, geography, observability, and retry behavior before committing.
Recommended Free Tools
Or skip the browser setup
If your task is obtaining a clean visual capture rather than crawling raw HTML, ScreenshotNeo makes one HTTP request and returns PNG, JPEG, WebP, or PDF. It accepts consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be disabled. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing result. Its MCP server provides take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients.
See the ScreenshotNeo API documentation for all options, including device presets, full-page lazy-image loading, CSS selectors, JavaScript, waits, blocking rules, cookies, headers, geolocation, PDF page ranges, signed links, asynchronous webhooks, bulk capture, caching TTL, and usage reporting.
Best Value
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
There is a free allowance of 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 shots, and every feature is on every plan. Create a free ScreenshotNeo account to try it.
Operational checklist
- Permission, robots.txt, API/export options, and contactable user agent are documented.
- Proxy credentials are secret-managed and absent from logs.
- Independent and stateful workflows have different route policies.
- Delay and per-domain concurrency reflect the target’s stated limits.
- 429/503 responses, ban pages, retries, latency, and proxy failures are recorded.
- Backoff pauses the crawl instead of increasing rotation.
- Package versions and extension compatibility are checked before deployment.
Frequently Asked Questions
Should I rotate proxies on every request?
No. Rotate only when the request is independent and the target’s limits permit it; keep a stable route for sessions, authentication, and multi-step state.
Does a proxy make scraping legal?
No. It changes the network route, not your permission. Follow the site’s robots.txt, terms, documented limits, and applicable law.
How many proxies do I need?
There is no evidence-based universal number. Size the pool from your permitted workload, required geography, failure rate, and the aggregate delay and concurrency the target tolerates.
What is the difference between a 429 and a 503?
A 429 commonly indicates rate limiting; a 503 can indicate overload or a block response. Inspect headers and body, then back off for either signal rather than assuming a new proxy fixes it.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.
Do these 3 things before closing this tab:
1Repair Windows errors before they cause bigger problems2Fix the driver behind crashes, sound loss and screen glitches3Clear out junk files and repair common Windows errors

