For a scraper that mostly waits for web servers, a modest concurrent.futures.ThreadPoolExecutor can fetch several authorized URLs at once. Give every request a finite timeout, keep a future-to-URL mapping, collect results with as_completed(), and measure throughput and errors before increasing concurrency. Threads do not make unlimited or automatically compliant traffic: the useful worker count depends on your URLs, network, target policies and parser.
When Python threads make a scraper faster
Downloading a page is usually an I/O-bound operation: your program spends much of its time waiting for DNS, a TCP/TLS connection and the server response. While one thread waits, another can process a different request. Python’s concurrency documentation distinguishes this from CPU-bound work, where threads may not provide the same benefit.
Threading is a fit when requests are independent and you have permission to retrieve them. It is not a license to bypass access controls, defeat CAPTCHAs, ignore terms, or send traffic faster than a site can reasonably handle. If parsing, image processing or machine-learning inference dominates the runtime, separate that CPU-heavy stage and benchmark it independently; a process pool or another design may be more appropriate.
What “fast” should mean
Measure more than elapsed seconds. Record total duration, completed pages, HTTP statuses, exceptions, timeouts, retry count, response bytes and the target’s behavior. A faster run that causes many failures or violates a site’s limits is not a successful scraper.
#1 Best Overall
A bounded threaded scraper with the standard library
The following complete example uses urllib.request, which is available with Python. It sets a timeout, closes each response with a context manager, preserves the original URL, and continues when one task fails.
from concurrent.futures import ThreadPoolExecutor, as_completed
from dataclasses import dataclass
from time import perf_counter
from urllib.error import HTTPError, URLError
from urllib.request import Request, urlopen
@dataclass
class FetchResult:
url: str
status: int | None
body: bytes | None
error: str | None
def fetch(url: str, timeout: float = 20.0) -> FetchResult:
request = Request(
url,
headers={"User-Agent": "AuthorizedResearchBot/1.0 (contact: [email protected])"},
)
try:
with urlopen(request, timeout=timeout) as response:
body = response.read()
return FetchResult(url, response.status, body, None)
except HTTPError as exc:
return FetchResult(url, exc.code, None, f"HTTP error: {exc.reason}")
except (URLError, TimeoutError) as exc:
return FetchResult(url, None, None, f"Network error: {exc}")
except Exception as exc:
return FetchResult(url, None, None, f"Unexpected error: {exc}")
def scrape(urls: list[str], max_workers: int = 6) -> list[FetchResult]:
results: list[FetchResult] = []
started = perf_counter()
with ThreadPoolExecutor(max_workers=max_workers) as pool:
future_to_url = {pool.submit(fetch, url): url for url in urls}
for future in as_completed(future_to_url):
url = future_to_url[future]
try:
result = future.result()
except Exception as exc:
# Protect the rest of the batch if fetch() ever leaks an exception.
result = FetchResult(url, None, None, f"Worker error: {exc}")
results.append(result)
if result.error:
print(f"FAIL {url}: {result.error}")
else:
print(f"OK {url}: {result.status} ({len(result.body or b'')} bytes)")
elapsed = perf_counter() - started
ok = sum(1 for result in results if result.error is None)
print(f"{ok}/{len(urls)} succeeded in {elapsed:.2f}s")
return results
if __name__ == "__main__":
urls = [
"https://example.com/",
"https://www.python.org/",
]
scrape(urls, max_workers=6)
max_workers=6 is a conservative starting choice, not a universal optimum. Start lower for a small or sensitive site. The pool bounds the number of simultaneous worker tasks; it does not guarantee a request rate, because response times vary.
Why the future-to-URL mapping matters
as_completed() yields whichever request finishes first, so output is not in input order. Mapping each future to its URL prevents a fast response from being attributed to the wrong page. It also lets one failure be reported while other downloads continue. If you need original ordering, sort the final records by the input list after collection.
Keep downloading separate from parsing
The example returns bytes. Parse HTML after a successful download, or submit parsing to a separate stage. This makes it clear whether a “speedup” came from overlapping network waits or from changing parser work. Decode using the response’s declared charset when available rather than assuming every page is UTF-8.
Respect robots.txt, permissions and site limits
The standard library includes urllib.robotparser for reading a site’s robots.txt. Use it as one technical input to your access decision, not as a replacement for terms, contracts, authentication requirements or applicable law. Crawl only URLs you are authorized to access, identify your client honestly, and provide contact information where appropriate.
Rank #2
Keep concurrency bounded per target host. Avoid a large burst across many domains, and do not use threads to evade rate limits, bot checks or access controls. For transient failures, a small, bounded backoff may help; do not blindly retry permanent 4xx responses or repeat a request indefinitely. Cache data you are allowed to reuse, and honor a site’s stated request guidance.
Choosing urllib or Requests
urllib.request keeps this tutorial dependency-free and supports timeout-enabled, context-managed responses. Requests offers a higher-level API, sessions, automatic keep-alive and connection pooling. Its documentation identifies Python 3.10+ support for its documented 2.34.2 release; verify the current release and supported Python versions when you deploy.
| Consideration | urllib.request |
Requests |
|---|---|---|
| Installation | Included with Python | Third-party package |
| Timeouts | Pass timeout to urlopen |
Pass timeout arguments to request methods |
| Connection reuse | Lower-level handling | Sessions document keep-alive and connection pooling |
| API ergonomics | More explicit request/response objects | Convenient methods, sessions and familiar exceptions |
| Speed | No general winner is established; benchmark equivalent code against your authorized workload. | |
Requests version of the worker
import requests
from concurrent.futures import ThreadPoolExecutor, as_completed
def fetch_with_requests(url: str, timeout: tuple[float, float] = (5, 20)):
try:
response = requests.get(
url,
timeout=timeout,
headers={"User-Agent": "AuthorizedResearchBot/1.0"},
)
return {
"url": url,
"status": response.status_code,
"body": response.content,
"error": None,
}
except requests.RequestException as exc:
return {"url": url, "status": None, "body": None, "error": str(exc)}
def run(urls, workers=6):
with ThreadPoolExecutor(max_workers=workers) as executor:
futures = {executor.submit(fetch_with_requests, url): url for url in urls}
for future in as_completed(futures):
result = future.result()
print(result["url"], result["status"], result["error"])
A shared session can reuse connections, but concurrent access and session configuration should be tested with your Requests version. Keep timeout values explicit even when using a session.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Benchmark your own worker count
No universal thread count or percentage speedup applies to every scraper. Use the same URL list, headers, timeout, parser and compliance limits for each run:
- Run a sequential baseline and save duration, successes, statuses, errors and bytes.
- Run small pools, such as 2, 4 and 6 workers, while watching the target’s response behavior.
- Increase only while elapsed time improves without unacceptable errors, retries, resource use or policy violations.
- Repeat runs at a comparable time and report the environment, date, target set and limits with any numbers you publish.
Do not benchmark against a site you do not control merely to chase a larger number. A local test server or an authorized staging environment gives you reproducible conditions.
Reliability patterns that prevent silent data loss
Use explicit result records
Return URL, status, body (or parsed data), error type and attempt information. Write successful results incrementally if a batch is large, so a process interruption does not discard completed work.
Classify failures
- Timeouts and connection resets: record them and retry only when the operation is safe and the failure is plausibly transient.
- HTTP 429: slow down and follow the server’s stated retry guidance rather than adding workers.
- HTTP 403 or bot checks: do not attempt to evade them; obtain permission or use an approved interface.
- 4xx responses: check the URL, authentication and requested resource before retrying.
- 5xx responses: a bounded backoff can be appropriate, with a cap and a final recorded failure.
- Malformed or unexpected HTML: keep the raw response when permitted and make parsers tolerant of missing fields.
Protect memory and file descriptors
Reading every response fully into memory is simple but may be unsuitable for large files. Stream or impose a maximum size when your application allows it. Always close responses; the context manager in the example does that even when parsing raises an exception.
Recommended Free Tools
Troubleshooting
Everything is slow
Check DNS, TLS, server latency, redirects and your timeout. Confirm that the workload is actually network-bound. If parsing consumes most of the elapsed time, optimize or isolate parsing instead of adding threads.
The program exits with incomplete results
Keep the executor in a with block and consume every submitted future. Persist results as they arrive, and catch exceptions around future.result().
Requests time out after adding workers
Reduce max_workers, inspect target rate limits and verify that your network can sustain the connections. More threads can increase contention and trigger defensive controls.
Character decoding is wrong
Use the response’s declared encoding where reliable, inspect the document’s metadata, and preserve raw bytes for later correction. Do not assume a single encoding for every domain.
Do these 3 things before closing this tab:
1Fix the driver behind crashes, sound loss and screen glitches2Repair Windows errors before they cause bigger problems3Scan for outdated or missing drivers - takes under a minuteRobots rules are unclear
Parse the file, identify the relevant user-agent group and review the site’s terms or contact its operator. A parser can report technical rules; it cannot decide whether your planned collection is lawful or authorized.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Or skip the browser setup
If your goal is a clean image or PDF of a page rather than raw HTML, ScreenshotNeo provides a single website-screenshot API call. It accepts the cookie or consent banner like a visitor, then removes more than 60 known consent platforms, newsletter popups and chat widgets before capture; each cleanup step can be disabled. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads and cache hits are not billed, and response headers identify the page verdict and billing status.
Use the API documentation at https://screenshotneo.com/docs/ for authentication and options. A minimal call is:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python:
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
ScreenshotNeo also offers full-page and selector captures, lazy-image loading, dark mode, device and retina settings, PDF page controls, custom CSS or JavaScript, clicks, waits, request blocking, headers, cookies, user agents, timezone and geolocation, transparent backgrounds, resizing, configurable caching, signed image links, asynchronous webhooks, bulk capture of up to 100 URLs per call, a usage API and an OpenAPI specification. Its MCP server exposes take_screenshot, get_page_info and capture_pdf to Claude, Cursor and other MCP clients.
Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchWindows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallEvery plan includes every feature. The Free plan provides 1,000 shots per month without a card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account to try it.
Best Value
FAQ
Can threads bypass a site’s robots.txt?
No. Threads are an execution technique, not permission. Check robots guidance, terms, authorization and applicable rules before collecting pages.
Should I use one thread per URL?
No. Submit tasks to a bounded executor. An unbounded thread-per-URL design can exhaust local resources and overwhelm a target.
Do I need a browser for every scraper?
No. For pages whose content is delivered in the initial HTTP response, an HTTP client is simpler. Browser automation is needed only when authorized collection genuinely requires client-side rendering or interaction.
Frequently Asked Questions
Can I preserve input order while using as_completed()?
Yes. Store each result by its original index or URL, then rebuild the output in the order of the input list after all futures finish.
What should I log for a production crawl?
Log the URL, timestamp, attempt, status, elapsed time, response size, exception class, retry decision and final outcome, while excluding credentials and sensitive response data.
When should I stop increasing concurrency?
Stop when latency or success rate stops improving, target errors or rate-limit responses rise, local resources become constrained, or the site’s permitted request behavior would be exceeded.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.
The Tool Desk
Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →




