Free tools Windows power users keep installed
One-click scans. No signup required.
For a scraper that mostly waits for websites to respond, use concurrent I/O: choose asyncio with an async HTTP client when your application is already asynchronous or must coordinate many requests; choose a thread pool when synchronous request and parsing code is simpler. Use processes when CPU-heavy parsing or transformation—not network waiting—is the bottleneck. None is a universal speed winner: measure the same workload under the same conditions before changing architectures.
First identify what is making the scraper slow
A scraper’s elapsed time can come from distinct kinds of work. It may wait for DNS, connection setup, server responses, or downloads; it may spend time parsing HTML and transforming data; or it may be slowed by retries, rate limits, or an inefficient pipeline. Concurrency helps most when independent tasks spend substantial time waiting. It cannot make a slow server respond faster, and it will not automatically accelerate CPU-heavy Python code.
- Mostly network waiting: overlap requests using async I/O or threads.
- Mostly Python CPU work: profile the parsing or transformation stage; a process pool may help parallelize that stage.
- Mixed workload: keep network fetching concurrent, then send only expensive CPU work to processes if measurement shows that the added complexity is worthwhile.
The Python Software Foundation’s Concurrent Execution documentation for Python 3.14.7 frames the choice around whether work is CPU-bound or I/O-bound and whether the preferred style is cooperative event-driven or preemptive multitasking. That is a model-selection guide, not a claim that one approach always wins.
Processes, threads, and async compared
| Approach | Best fit | Main trade-off | Implementation cue |
|---|---|---|---|
Async / asyncio |
Many network waits with an async-capable client, especially in an async application. | Requires non-blocking calls and cooperative code; blocking work stalls the event loop. | Use an async HTTP client such as HTTPX’s AsyncClient and await its request methods. |
| Threads | Blocking synchronous network libraries or adding concurrency to existing synchronous code. | Thread coordination and shared-state concerns; ordinary CPython’s GIL limits parallel execution of Python bytecode for CPU-bound work. | Use a thread pool for blocking functions, or an executor to move blocking work off an event loop. |
| Processes | CPU-heavy parsing or transformations that need parallel Python execution. | More operational and data-sharing complexity; process-pool inputs and outputs must meet pickling constraints. | Isolate the CPU-heavy function and pass serializable inputs and results. |
This is a practical guide, not a benchmark. The Python documentation describes asyncio as an event-loop scheduler that switches tasks to support non-blocking I/O. A thread pool can overlap blocking I/O, but ordinary CPython’s GIL means threads are not a general way to run Python bytecode in parallel for CPU-bound work. A process pool uses multiple processes to sidestep that limitation, at the cost of transferring work and data between processes.
#1 Best Overall
Use async when the HTTP client and the work are actually asynchronous
Async works well when a task can yield while waiting—for example, while an HTTP response is in flight—so the event loop can run another task. Writing async def around a synchronous request does not make that request non-blocking. Likewise, a long CPU-bound function executed directly in a coroutine prevents the loop from servicing other tasks until that function returns.
HTTPX provides both synchronous and asynchronous interfaces. This illustrative pattern uses its async interface, a shared client, and a semaphore to cap simultaneous requests. Install HTTPX with python -m pip install httpx. Set the target URLs to pages you are authorized to access and adapt response handling to the site’s terms and your extraction needs.
import asyncio
import httpx
URLS = [
"https://example.com/",
"https://www.python.org/",
]
MAX_IN_FLIGHT = 5
async def main():
limits = httpx.Limits(max_connections=MAX_IN_FLIGHT)
timeout = httpx.Timeout(20.0)
semaphore = asyncio.Semaphore(MAX_IN_FLIGHT)
async with httpx.AsyncClient(
limits=limits,
timeout=timeout,
follow_redirects=True,
) as client:
async def fetch(url):
async with semaphore:
response = await client.get(url)
response.raise_for_status()
return url, response.status_code, len(response.content)
results = await asyncio.gather(
*(fetch(url) for url in URLS),
return_exceptions=True,
)
for result in results:
if isinstance(result, Exception):
print("Request failed:", repr(result))
else:
print(result)
if __name__ == "__main__":
asyncio.run(main())
The semaphore and connection limit make the concurrency cap explicit; five is only an example, not a universally safe or optimal setting. gather with return_exceptions=True lets other outcomes be reported when an individual request fails. For production work, add deliberate retry rules for transient errors, logging, and a policy for handling non-success status codes. Do not retry indiscriminately: retries add load and can make an overloaded target worse.
Keep blocking work off the event loop
If a legacy blocking function must run alongside async tasks, Python’s event-loop documentation shows executor-based offloading. For example, await asyncio.to_thread(blocking_function, argument) can move a blocking call to a worker thread. This prevents that call from occupying the event-loop thread, but does not turn CPU-bound Python bytecode into parallel work under ordinary CPython.
Quick wins for a faster PC:
Repair Windows errors before they cause bigger problemsFix Now →Scan for outdated or missing drivers - takes under a minuteDriver Scan →Rank #3
Use threads when synchronous code is the simplest fit
If the scraper already uses a synchronous HTTP library and the expensive part is waiting for responses, a thread pool can add concurrency without rewriting the whole program around coroutines. Keep worker functions independent where practical, bound the number of workers, and avoid unsafe shared mutable state. The example below uses Python’s standard-library executor; replace fetch with your blocking request-and-parse function.
from concurrent.futures import ThreadPoolExecutor, as_completed
from urllib.request import urlopen
URLS = ["https://example.com/", "https://www.python.org/"]
def fetch(url):
with urlopen(url, timeout=20) as response:
body = response.read()
return url, response.status, len(body)
if __name__ == "__main__":
with ThreadPoolExecutor(max_workers=5) as pool:
futures = [pool.submit(fetch, url) for url in URLS]
for future in as_completed(futures):
try:
print(future.result())
except Exception as exc:
print("Request failed:", repr(exc))
The worker count of five is illustrative, not a recommendation for every site or workload. A larger pool can raise memory use, connection pressure, error rates, or the chance of triggering rate limits. Match concurrency to the target’s published rules and observed behavior.
Use processes for measured CPU-heavy stages
When profiling shows that parsing or transformation consumes substantial CPU time, isolate that function and consider a process pool. Do not send open clients, sockets, or unnecessarily large objects between workers; pass simple serializable inputs and return compact results. Python’s ProcessPoolExecutor uses multiprocessing, so submitted callables, arguments, and results must be picklable, and the main module must be importable by worker subprocesses.
from concurrent.futures import ProcessPoolExecutor
def transform(html):
# Replace with CPU-heavy parsing or transformation.
return html.count("<a")
if __name__ == "__main__":
pages = ["<a href='x'>one</a>", "<p>text</p>"]
with ProcessPoolExecutor() as pool:
counts = list(pool.map(transform, pages))
print(counts)
The if __name__ == "__main__": guard is important for worker-process startup on platforms that import the main module. Creating processes and serializing data have costs, so a process pool can be a poor fit for tiny tasks or work that is mostly waiting on HTTP.
Windows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallCrashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minuteBest Value
How to decide with a fair measurement
- Record a sequential baseline. Use a representative URL set and capture total elapsed time, successful pages, failures, retries, memory, and CPU use.
- Separate stages. Time network fetches and parsing/transformation independently where possible. This reveals whether waiting or CPU work dominates.
- Change one variable. Compare a bounded async implementation or thread pool for I/O, and test processes only for a CPU-heavy stage. Keep the URL set, parsing behavior, timeouts, and output work constant.
- Repeat under comparable conditions. Remote response times fluctuate; note Python and library versions, concurrency limits, machine, target conditions, and the actual request set. Do not infer a general winner from a single run.
- Check quality and impact, not just elapsed time. Compare successful pages per second alongside errors, retries, memory and CPU use. Respect the site’s rate limits and avoid increasing request pressure simply to improve a local timing number.
No independently validated end-to-end benchmark establishes a universal speedup or request threshold for these approaches. Treat any speed claim without the workload, versions, limits, target conditions, and measurements as a rule of thumb rather than a result that will necessarily transfer to your scraper.
Common failure modes and fixes
- Async appears no faster than sequential: check that the HTTP calls are truly awaited through an async client, rather than synchronous calls inside coroutines. Confirm that independent requests overlap and that the target is not limiting or serializing them.
- Other async tasks freeze during a request or parse: find synchronous blocking calls or long CPU work running on the event-loop thread. Use an async-capable operation or offload blocking work to an executor; consider processes for CPU-heavy Python work.
- More threads or tasks increase errors: reduce the concurrency cap, inspect status codes and timeout behavior, and follow the site’s limits. More simultaneous requests can increase pressure without improving useful throughput.
- Process-pool submission fails: make sure worker functions and data are picklable, define worker functions at module scope, and use the main-module guard so subprocesses can import the program correctly.
- The program is faster but returns fewer usable pages: compare successful pages and retries as well as elapsed time. A high request rate that causes failures is not a throughput improvement.
Or skip the browser setup
If your task is to capture rendered webpages as images or PDFs—not extract structured data from pages—a screenshot API can avoid building and maintaining browser-capture infrastructure. ScreenshotNeo is a website screenshot API and MCP server; it is not a replacement for a scraper that needs page text or records. For screenshot work, its one-request example is:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
See the ScreenshotNeo API documentation for request options. It accepts cookie or consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; those steps can be turned off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers report the page verdict and billing status. Its MCP server offers take_screenshot, get_page_info, and capture_pdf tools for AI agents and MCP clients. The free plan includes 1,000 shots per month with no card; paid plans start at $5 for 3,000 shots. Sign up free for ScreenshotNeo.
Frequently Asked Questions
Does asyncio make Python code run on multiple CPU cores?
No. Asyncio coordinates cooperative tasks through an event loop; it is useful for overlapping non-blocking waits, not by itself for parallel execution of CPU-bound Python code.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Can I combine async fetching and process-based parsing?
Yes, when measurements justify the added complexity: fetch with bounded async I/O, then submit only the CPU-heavy parsing or transformation inputs to a process pool.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




