Do these 3 things before closing this tab:
1Scan for outdated or missing drivers - takes under a minute2Clear out junk files and repair common Windows errors3Fix the driver behind crashes, sound loss and screen glitchesSome links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
To fetch multiple web pages without waiting for each response in sequence, use Python’s asyncio to coordinate concurrent tasks and aiohttp to make asynchronous HTTP requests. Parse each returned HTML document separately with the parser that suits your data. Reuse one aiohttp.ClientSession for the batch, limit concurrency, and handle status codes and timeouts deliberately. Async requests can overlap network waits; they do not guarantee a fixed speedup or override a site’s access controls.
What asyncio does—and what it does not do
asyncio is Python’s library for writing concurrent code, and Python describes it as often a good fit for I/O-bound network code. It coordinates coroutines while they wait for operations such as HTTP responses; it is not itself an HTTP client or an HTML parser. Python’s asyncio documentation
- asyncio: schedules and coordinates asynchronous work.
- aiohttp: sends asynchronous HTTP requests and reads their responses.
- An HTML parser: extracts fields from the response text after it has been fetched. Choose one based on your project’s needs; the cited asyncio and aiohttp documentation does not compare parser libraries.
This pattern helps when you have multiple independent URLs and much of the work is waiting on the network. It does not make CPU-heavy parsing asynchronous, guarantee a particular speedup, or bypass a site’s restrictions. Results depend on the workload, network, server behavior, and implementation.
Windows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallOutdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchInstall aiohttp and prepare your URL list
Install aiohttp in the same Python environment that will run your script:
#1 Best Overall
python -m pip install aiohttp
Save the following as scrape_async.py. Replace the example URLs with pages you are permitted to access. The script fetches HTML and prints each response’s status and body length; the parsing step comes later.
Fetch multiple pages with one reusable session
import asyncio
import aiohttp
URLS = [
"https://example.com/",
"https://www.python.org/",
]
async def fetch(session, url):
try:
async with session.get(url) as response:
body = await response.text()
return {
"url": url,
"status": response.status,
"body": body,
}
except (aiohttp.ClientError, asyncio.TimeoutError) as exc:
return {
"url": url,
"error": f"{type(exc).__name__}: {exc}",
}
async def main():
timeout = aiohttp.ClientTimeout(total=30)
connector = aiohttp.TCPConnector(limit=5)
async with aiohttp.ClientSession(
timeout=timeout,
connector=connector,
) as session:
results = await asyncio.gather(
*(fetch(session, url) for url in URLS)
)
for result in results:
if "error" in result:
print(f"FAILED {result['url']}: {result['error']}")
else:
print(
f"{result['status']} {result['url']} "
f"({len(result['body'])} characters)"
)
if __name__ == "__main__":
asyncio.run(main())
Run it from a normal terminal with python scrape_async.py. A successful request produces a status and character count for each page. A failed request produces an error entry for that URL rather than stopping the whole batch.
Why use one ClientSession?
The session owns a connection pool and supports connection reuse across requests. Reusing it avoids repeatedly creating and closing session-level resources for every URL. The aiohttp quickstart explicitly advises: “Don’t create a session per request.” aiohttp Client Quickstart
What gather and the connection limit mean
asyncio.gather schedules the independent fetch coroutines and waits for their results. The connector’s limit=5 bounds simultaneous connections for this session; it is an example setting, not a universally appropriate rate. Pick a conservative bound for the target site and your workload, and respect any site-specific instructions. A connection limit is not a substitute for an intentional request pace when the site needs one.
Rank #2
Python 3.11 and newer: TaskGroup
For a structured-concurrency approach, Python 3.11+ provides asyncio.TaskGroup. Tasks created inside its context are awaited when the context exits. The example above uses gather because it collects a result per URL, including handled request errors. With a task group, decide explicitly how you want exceptions from individual requests to affect the batch; an unhandled task error can cancel sibling work.
Read responses, check status, and extract data
The example awaits response.text() before leaving the response context, so it has the complete HTML string available. A returned HTTP status is not necessarily a successful page: for example, the request can complete and still return an error status. Decide which statuses your use case accepts and record or handle the others rather than treating every body as valid content.
Once a response is retrieved, pass its text to your chosen HTML parser and extract only the fields you need. Keep fetching and parsing conceptually separate: aiohttp retrieves the document; your parser interprets it. If a target returns non-HTML content, a login page, or a challenge page, parsing it as the expected page can yield empty or misleading fields.
Recommended Free Tools
Convenient whole-body reads
Aiohttp’s text(), json(), and read() methods load the response body into memory. They are convenient when pages or API responses are modest in size and you need the full content. Use the method appropriate to the payload: HTML is commonly read as text; JSON responses can be decoded with json(); raw bytes can be read with read(). aiohttp Client Quickstart
Stream large responses
For very large bodies, consume response.content incrementally rather than materializing the entire response with text() or read(). For example, this writes chunks to a file while the response is open:
async def download_to_file(session, url, path):
async with session.get(url) as response:
response.raise_for_status()
with open(path, "wb") as output:
async for chunk in response.content.iter_chunked(64 * 1024):
output.write(chunk)
Streaming changes how you handle the body; it does not remove the need to check the response, choose a sensible concurrency bound, and handle I/O errors. If you need to parse a large HTML document, decide whether to save it, process it incrementally with a suitable parser, or accept whole-body memory use.
Check site rules and choose a responsible request pace
Before collecting pages, inspect the target’s robots.txt and applicable site terms. Python’s urllib.robotparser can read a robots file and answer whether a user agent may fetch a URL with can_fetch. It also exposes crawl-delay and request-rate values when present. Python urllib.robotparser documentation
A robots file is a useful signal for crawler behavior, not a complete legal determination. The Python documentation describes the parser API; it does not establish whether scraping a particular site or dataset is lawful. Legal requirements depend on the jurisdiction, target, data, and intended use. Do not treat concurrency as permission: use a conservative pace, reduce load if the site signals trouble, and stop if access is denied.
The documentation does not prescribe a universal connection limit, timeout, or retry policy. The example’s five-connection limit and 30-second total timeout are starting values to adapt, not official requirements. Avoid automatic aggressive retries: repeated requests can increase load, and retrying an access-denied response does not make access appropriate.
Run it in scripts, notebooks, and larger jobs
Normal Python scripts
Use asyncio.run(main()) once at the top level of a regular script. It creates and manages the event loop for that entry point, then closes it when the coroutine finishes. Keep the batch’s session inside an async context manager so it is closed after the work is done.
Notebooks and already-running event loops
Some notebooks and interactive environments already run an event loop. In that case, do not call asyncio.run() from inside the running loop. Instead, run or await the coroutine using the environment’s existing loop—for example, in a notebook cell that supports top-level await, use await main().
Persist results deliberately
The sample prints a summary and keeps successful bodies in memory only until the program exits. For a real collection job, decide what to save, how to identify a page, and how to represent failures. Write output incrementally when a batch could be large, and avoid retaining every full HTML body if the next stage only needs a few extracted fields.
Best Value
Common errors and practical fixes
- “Session is closed” or a closed-connector error: a request is running after its session context has exited. Keep all request tasks within the
ClientSessioncontext and await them before leaving it. - Timeout exceptions: the server may be slow, unreachable, or not responding within the configured total timeout. Check the URL and connectivity, then adjust the timeout to fit the workload rather than retrying without limit.
- Client connection or DNS errors: verify the hostname, network access, and URL scheme. Capture the exception per URL so one failed host does not erase the status of the rest of the batch.
- Non-success status or unexpected HTML: inspect the status and a small portion of the response before parsing. The target may have moved the page, denied access, or returned a different document than expected.
- Too many simultaneous requests: lower the connector limit and use a slower request pace. Asyncio makes it easy to initiate many waits, but the target server and your own network still have finite capacity.
- Memory rises during a large batch: avoid collecting every full body in a list. Extract and persist only necessary fields, process results in bounded batches, or stream large responses through
response.content. asyncio.run()reports an existing event loop: remove that call from the nested environment and await the coroutine through the loop already in use.
Performance, reliability, and cost considerations
Async I/O can improve throughput when many independent requests spend time waiting on network responses, because those waits can overlap. It does not guarantee a fixed multiplier or make every scraper faster: a single URL, CPU-heavy parsing, a slow target, rate limits, connection overhead, and local resource limits can change the result. There is no general speedup figure established here; measure your actual workload if performance matters.
Reliability comes from controlling the batch as much as from concurrency: bound simultaneous connections, set a timeout appropriate to the work, record per-URL errors, and avoid unbounded retries or in-memory accumulation. Treat server responses as data to inspect, not proof that the intended page was fetched. Asyncio and aiohttp are software libraries; their use does not remove the ordinary network, operational, or legal considerations of collecting web data.
Or skip the browser setup
If your task is to capture a rendered page as an image or PDF rather than extract arbitrary fields from HTML, ScreenshotNeo offers a one-request screenshot API. It is not a replacement for an aiohttp scraper when you need to parse page data. Here is the Python request pattern for a screenshot:
The Tool Desk
Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
See the ScreenshotNeo API documentation for request options. It removes cookie or consent banners, newsletter popups, and chat widgets before capture; each of those steps can be turned off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and the response identifies the page verdict and billing status in headers. An MCP server provides take_screenshot, get_page_info, and capture_pdf tools for AI agents. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 shots. Sign up for ScreenshotNeo’s free plan.
Frequently Asked Questions
Does asyncio parse HTML?
No. It coordinates asynchronous work; aiohttp retrieves pages, and a separate parser extracts information from the HTML.
Does robots.txt decide whether scraping is legal?
No. The robot parser reports rules and directives in robots.txt; it is not a legal assessment.
Can I use ScreenshotNeo to extract fields from a page?
ScreenshotNeo is a screenshot and PDF capture service. Use an HTTP client and parser when you need to extract arbitrary page data.
Quick wins for a faster PC:
Scan for outdated or missing drivers - takes under a minuteDriver Scan →Clear out junk files and repair common Windows errorsFree Scan →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

