The Tool Desk
Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →To fetch many pages without waiting for each one sequentially, run the blocking fetch function in a ThreadPoolExecutor, or use asyncio with an async client such as aiohttp. In both designs, cap concurrency, reuse a session, set finite timeouts, keep every result tied to its URL, and handle failures per task.
The right choice depends on your program: threads integrate with existing Requests code, while asyncio fits an async application and large I/O-bound batches. Neither is universally faster, and neither removes the need to respect a site’s robots rules, terms, and rate limits.
Choose the concurrency model first
| Situation | Use | Why |
|---|---|---|
| Your fetch function uses Requests or another blocking client | ThreadPoolExecutor |
Minimal changes to synchronous code; waiting happens in worker threads. |
Your application already uses async/await |
asyncio with aiohttp |
Native coroutine scheduling, connector limits, and structured cancellation. |
| The site publishes an API or bulk export | Use that interface | An official endpoint is usually more predictable and can be cheaper for the site than HTML crawling. |
Concurrency is a cap, not a target. A value that works for one domain can trigger throttling or bans on another. Start conservatively, observe responses, and tune per host.
Before you send requests
Check access rules
Read the destination’s robots.txt, terms, and any documented API policy. Robots directives are useful guidance but are not a complete legal determination; terms and applicable law may impose additional limits.
#1 Best Overall
Identify the work you actually need
Prefer a documented API, sitemap, RSS feed, or bulk export when available. Fetch only required URLs, cache repeat work, and use a descriptive user agent where appropriate.
Plan the output contract
Concurrent tasks finish out of order. Decide whether callers need completion order, input order, or a record containing URL, status, timing, and error. The examples below preserve URL association and restore input order when requested.
Option 1: ThreadPoolExecutor with Requests
This pattern is suitable when your existing scraper is synchronous. A requests.Session persists cookies and reuses pooled connections, so create one session per worker thread rather than sharing a mutable session across threads.
- Define a blocking fetch function with a finite timeout.
- Submit URLs to a bounded
ThreadPoolExecutor. - Map each
Futureback to its URL. - Consume futures with
as_completedand catch errors per URL.
from concurrent.futures import ThreadPoolExecutor, as_completed
from threading import local
import requests
_thread_state = local()
def session_for_thread():
if not hasattr(_thread_state, "session"):
s = requests.Session()
s.headers.update({
"User-Agent": "ExampleResearchBot/1.0 (+https://example.com/bot-info)"
})
_thread_state.session = s
return _thread_state.session
def fetch(url, timeout=(5, 30)):
"""Return a result record; never hide which URL failed."""
try:
response = session_for_thread().get(url, timeout=timeout)
response.raise_for_status()
return {
"url": url,
"status": response.status_code,
"text": response.text,
"error": None,
}
except requests.RequestException as exc:
return {"url": url, "status": None, "text": None, "error": str(exc)}
def scrape(urls, max_workers=8):
results = []
with ThreadPoolExecutor(max_workers=max_workers) as pool:
future_to_url = {pool.submit(fetch, url): url for url in urls}
for future in as_completed(future_to_url):
url = future_to_url[future]
try:
result = future.result()
except Exception as exc: # protects the batch from unexpected worker errors
result = {"url": url, "status": None, "text": None, "error": repr(exc)}
results.append(result)
return results
urls = [
"https://example.com/one",
"https://example.com/two",
"https://example.org/three",
]
completed_order = scrape(urls, max_workers=8)
input_order = sorted(completed_order, key=lambda item: urls.index(item["url"]))
for item in input_order:
print(item["url"], item["status"], item["error"])
max_workers limits simultaneous calls; it does not guarantee that eight requests are safe for a target. For many hosts, group URLs by hostname and apply a separate, lower limit or delay to each group. If your fetch function is pure and stateless, a shared immutable configuration is fine; avoid mutating one Requests session from multiple threads.
Rank #2
Preserving order without an expensive lookup
For large batches, attach an index before submitting:
indexed = list(enumerate(urls))
ordered = [None] * len(indexed)
with ThreadPoolExecutor(max_workers=8) as pool:
futures = {pool.submit(fetch, url): i for i, url in indexed}
for future in as_completed(futures):
ordered[futures[future]] = future.result()
Option 2: asyncio and aiohttp
Python describes asyncio as “often a perfect fit for IO-bound and high-level structured network code.” Use a truly asynchronous HTTP client; calling blocking Requests inside a coroutine stalls the event loop.
import asyncio
import aiohttp
URLS = [
"https://example.com/one",
"https://example.com/two",
"https://example.org/three",
]
async def fetch(session, url):
try:
async with session.get(url) as response:
response.raise_for_status()
text = await response.text()
return {"url": url, "status": response.status, "text": text, "error": None}
except (aiohttp.ClientError, asyncio.TimeoutError) as exc:
return {"url": url, "status": None, "text": None, "error": str(exc)}
async def scrape(urls):
timeout = aiohttp.ClientTimeout(total=30, connect=5, sock_read=25)
connector = aiohttp.TCPConnector(limit=20, limit_per_host=4)
headers = {"User-Agent": "ExampleResearchBot/1.0 (+https://example.com/bot-info)"}
async with aiohttp.ClientSession(
connector=connector, timeout=timeout, headers=headers
) as session:
tasks = [asyncio.create_task(fetch(session, url)) for url in urls]
# gather returns results in the same order as tasks, even if completion differs.
return await asyncio.gather(*tasks)
results = asyncio.run(scrape(URLS))
for item in results:
print(item["url"], item["status"], item["error"])
Why the connector settings matter
TCPConnector(limit=20) caps total open connections and limit_per_host=4 caps connections to one host. aiohttp’s reference currently documents a total default of 100 and no per-host limit; treat those as library defaults, not a safe scraping rate. A reusable ClientSession owns the pool and keep-alive connections, and the async context manager closes it reliably.
Adding a semaphore for application-level control
A connector limits connections, while a semaphore can limit active application tasks (including parsing or retries):
sem = asyncio.Semaphore(10)
async def limited_fetch(session, url):
async with sem:
return await fetch(session, url)
Do not create an unbounded task for millions of URLs. Process an iterator in batches or use a queue so memory use remains bounded.
Timeouts, retries, and failure handling
Use separate connect and read budgets
A connect timeout catches unreachable hosts; a read timeout prevents a server that accepts a connection but never finishes from occupying a worker indefinitely. Choose values based on your workload and record them with results.
Retry only transient failures
Retries can multiply load. Consider retrying a small number of times for connection resets, 502, 503, or 504 responses, with exponential backoff and jitter. Do not blindly retry 401, 403, 404, or a robots-denied URL. Honor Retry-After when supplied.
Never lose the URL
Catch exceptions inside each task or immediately after retrieving its future. Store status, exception text, and attempt count so one bad page does not erase diagnostic context or abort the entire batch.
Do these 3 things before closing this tab:
1Scan for outdated or missing drivers - takes under a minute2Repair Windows errors before they cause bigger problems3Fix the driver behind crashes, sound loss and screen glitchesRate limits, robots, and polite crawling
Python’s urllib.robotparser can evaluate can_fetch and expose a site’s crawl_delay or request_rate when declared. Apply those signals before scheduling work. Scrapy’s guidance also warns that exceeding a site’s tolerated rate can cause throttling, errors, or bans.
- Use per-domain concurrency and delays rather than one global number.
- Cache responses and avoid fetching unchanged pages.
- Identify your bot and provide contact information where suitable.
- Stop or slow down after repeated 429 or 503 responses.
- Prefer the site’s official API or export for recurring jobs.
Performance and reliability trade-offs
Threads are usually the simplest migration path for blocking code. Asyncio avoids one thread per request and gives explicit scheduling, but requires async-compatible libraries throughout the request path. Reusing sessions reduces connection setup and preserves cookies. Actual throughput depends on latency, server throttling, DNS and TLS costs, task count, parsing time, and your network; the documented mechanisms do not establish a universal speed winner.
Measure the right things
- Record start and end times, status, response bytes, and exception type.
- Watch timeout, 429, 5xx, and connection-error rates separately.
- Compare a small concurrency sweep (for example, 2, 4, and 8) against the target’s guidance.
- Monitor memory if response bodies are large; stream or discard content you do not need.
Common problems and fixes
| Symptom | Likely cause | Fix |
|---|---|---|
| Everything is still sequential | Blocking Requests call runs directly in the event loop, or the pool has one worker. | Use aiohttp in asyncio, or increase the pool cap modestly. |
| Many timeouts | Concurrency exceeds target capacity, or timeout is too short. | Lower global/per-host limits, inspect connect versus read timing, then adjust finite budgets. |
| 429 or 403 responses | Rate or access policy was violated. | Stop, read robots and terms, honor Retry-After, reduce rate, or use the official API. |
| Results appear scrambled | Completion order differs from input order. | Keep the future-to-URL map and sort by saved input index, or use ordered asyncio.gather results. |
| Connections accumulate | Sessions are created per request or not closed. | Reuse one session per batch (or per worker thread) and close it with a context manager. |
| One exception stops the batch | Future or task errors are awaited without per-item handling. | Catch exceptions per item and return an error record. |
Or skip the browser setup
If your goal is reliable page images rather than HTML extraction, ScreenshotNeo provides a one-request screenshot API and an MCP server for Claude, Cursor, and other MCP clients. It accepts consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets; bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, with the verdict exposed in response headers. Every feature is available on every plan, including full-page lazy-image loading, CSS-selector element capture, device presets, custom headers and cookies, waits, blocking rules, PDFs, signed links, async webhooks, bulk capture of up to 100 URLs per call, and a usage API.
cURL
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
Node.js
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
See the ScreenshotNeo API documentation for options and response headers. The Free plan includes 1,000 screenshots each month with no card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.
Frequently asked questions
Can I share one Requests Session across all threads?
Use a session per worker thread as shown, or otherwise verify thread-safety for your exact client and adapters. A single mutable session shared by many threads can create subtle state problems.
Best Value
Should I use processes instead of threads?
Processes are generally unnecessary for network waiting and add serialization and startup overhead. Consider them when CPU-heavy parsing dominates after downloads, not merely because the URL count is large.
How many URLs can one batch contain?
There is no universal limit. Bound in-flight tasks, memory, and per-host pressure; for very large jobs, stream URLs through a queue and checkpoint results.
Frequently Asked Questions
Can I share one Requests Session across all threads?
Use a session per worker thread as shown, or otherwise verify thread-safety for your exact client and adapters. A single mutable session shared by many threads can create subtle state problems.
Should I use processes instead of threads?
Processes are generally unnecessary for network waiting and add serialization and startup overhead. Consider them when CPU-heavy parsing dominates after downloads, not merely because the URL count is large.
How many URLs can one batch contain?
There is no universal limit. Bound in-flight tasks, memory, and per-host pressure; for very large jobs, stream URLs through a queue and checkpoint results.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

