Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

To fetch multiple web pages without waiting for each response in sequence, use Python’s asyncio to coordinate concurrent tasks and aiohttp to make asynchronous HTTP requests. Parse each returned HTML document separately with the parser that suits your data. Reuse one aiohttp.ClientSession for the batch, limit concurrency, and handle status codes and timeouts deliberately. Async requests can overlap network waits; they do not guarantee a fixed speedup or override a site’s access controls.

What asyncio does—and what it does not do

asyncio is Python’s library for writing concurrent code, and Python describes it as often a good fit for I/O-bound network code. It coordinates coroutines while they wait for operations such as HTTP responses; it is not itself an HTTP client or an HTML parser. Python’s asyncio documentation

  • asyncio: schedules and coordinates asynchronous work.
  • aiohttp: sends asynchronous HTTP requests and reads their responses.
  • An HTML parser: extracts fields from the response text after it has been fetched. Choose one based on your project’s needs; the cited asyncio and aiohttp documentation does not compare parser libraries.

This pattern helps when you have multiple independent URLs and much of the work is waiting on the network. It does not make CPU-heavy parsing asynchronous, guarantee a particular speedup, or bypass a site’s restrictions. Results depend on the workload, network, server behavior, and implementation.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Install aiohttp and prepare your URL list

Install aiohttp in the same Python environment that will run your script:

python -m pip install aiohttp

Save the following as scrape_async.py. Replace the example URLs with pages you are permitted to access. The script fetches HTML and prints each response’s status and body length; the parsing step comes later.

Fetch multiple pages with one reusable session

import asyncio
import aiohttp

URLS = [
    "https://example.com/",
    "https://www.python.org/",
]

async def fetch(session, url):
    try:
        async with session.get(url) as response:
            body = await response.text()
            return {
                "url": url,
                "status": response.status,
                "body": body,
            }
    except (aiohttp.ClientError, asyncio.TimeoutError) as exc:
        return {
            "url": url,
            "error": f"{type(exc).__name__}: {exc}",
        }

async def main():
    timeout = aiohttp.ClientTimeout(total=30)
    connector = aiohttp.TCPConnector(limit=5)

    async with aiohttp.ClientSession(
        timeout=timeout,
        connector=connector,
    ) as session:
        results = await asyncio.gather(
            *(fetch(session, url) for url in URLS)
        )

    for result in results:
        if "error" in result:
            print(f"FAILED {result['url']}: {result['error']}")
        else:
            print(
                f"{result['status']} {result['url']} "
                f"({len(result['body'])} characters)"
            )

if __name__ == "__main__":
    asyncio.run(main())

Run it from a normal terminal with python scrape_async.py. A successful request produces a status and character count for each page. A failed request produces an error entry for that URL rather than stopping the whole batch.

Why use one ClientSession?

The session owns a connection pool and supports connection reuse across requests. Reusing it avoids repeatedly creating and closing session-level resources for every URL. The aiohttp quickstart explicitly advises: “Don’t create a session per request.” aiohttp Client Quickstart

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What gather and the connection limit mean

asyncio.gather schedules the independent fetch coroutines and waits for their results. The connector’s limit=5 bounds simultaneous connections for this session; it is an example setting, not a universally appropriate rate. Pick a conservative bound for the target site and your workload, and respect any site-specific instructions. A connection limit is not a substitute for an intentional request pace when the site needs one.

Python 3.11 and newer: TaskGroup

For a structured-concurrency approach, Python 3.11+ provides asyncio.TaskGroup. Tasks created inside its context are awaited when the context exits. The example above uses gather because it collects a result per URL, including handled request errors. With a task group, decide explicitly how you want exceptions from individual requests to affect the batch; an unhandled task error can cancel sibling work.

Read responses, check status, and extract data

The example awaits response.text() before leaving the response context, so it has the complete HTML string available. A returned HTTP status is not necessarily a successful page: for example, the request can complete and still return an error status. Decide which statuses your use case accepts and record or handle the others rather than treating every body as valid content.

Once a response is retrieved, pass its text to your chosen HTML parser and extract only the fields you need. Keep fetching and parsing conceptually separate: aiohttp retrieves the document; your parser interprets it. If a target returns non-HTML content, a login page, or a challenge page, parsing it as the expected page can yield empty or misleading fields.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Convenient whole-body reads

Aiohttp’s text(), json(), and read() methods load the response body into memory. They are convenient when pages or API responses are modest in size and you need the full content. Use the method appropriate to the payload: HTML is commonly read as text; JSON responses can be decoded with json(); raw bytes can be read with read(). aiohttp Client Quickstart

Stream large responses

For very large bodies, consume response.content incrementally rather than materializing the entire response with text() or read(). For example, this writes chunks to a file while the response is open:

async def download_to_file(session, url, path):
    async with session.get(url) as response:
        response.raise_for_status()
        with open(path, "wb") as output:
            async for chunk in response.content.iter_chunked(64 * 1024):
                output.write(chunk)

Streaming changes how you handle the body; it does not remove the need to check the response, choose a sensible concurrency bound, and handle I/O errors. If you need to parse a large HTML document, decide whether to save it, process it incrementally with a suitable parser, or accept whole-body memory use.

Check site rules and choose a responsible request pace

Before collecting pages, inspect the target’s robots.txt and applicable site terms. Python’s urllib.robotparser can read a robots file and answer whether a user agent may fetch a URL with can_fetch. It also exposes crawl-delay and request-rate values when present. Python urllib.robotparser documentation

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A robots file is a useful signal for crawler behavior, not a complete legal determination. The Python documentation describes the parser API; it does not establish whether scraping a particular site or dataset is lawful. Legal requirements depend on the jurisdiction, target, data, and intended use. Do not treat concurrency as permission: use a conservative pace, reduce load if the site signals trouble, and stop if access is denied.

The documentation does not prescribe a universal connection limit, timeout, or retry policy. The example’s five-connection limit and 30-second total timeout are starting values to adapt, not official requirements. Avoid automatic aggressive retries: repeated requests can increase load, and retrying an access-denied response does not make access appropriate.

Run it in scripts, notebooks, and larger jobs

Normal Python scripts

Use asyncio.run(main()) once at the top level of a regular script. It creates and manages the event loop for that entry point, then closes it when the coroutine finishes. Keep the batch’s session inside an async context manager so it is closed after the work is done.

Notebooks and already-running event loops

Some notebooks and interactive environments already run an event loop. In that case, do not call asyncio.run() from inside the running loop. Instead, run or await the coroutine using the environment’s existing loop—for example, in a notebook cell that supports top-level await, use await main().

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Persist results deliberately

The sample prints a summary and keeps successful bodies in memory only until the program exits. For a real collection job, decide what to save, how to identify a page, and how to represent failures. Write output incrementally when a batch could be large, and avoid retaining every full HTML body if the next stage only needs a few extracted fields.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Common errors and practical fixes

  • “Session is closed” or a closed-connector error: a request is running after its session context has exited. Keep all request tasks within the ClientSession context and await them before leaving it.
  • Timeout exceptions: the server may be slow, unreachable, or not responding within the configured total timeout. Check the URL and connectivity, then adjust the timeout to fit the workload rather than retrying without limit.
  • Client connection or DNS errors: verify the hostname, network access, and URL scheme. Capture the exception per URL so one failed host does not erase the status of the rest of the batch.
  • Non-success status or unexpected HTML: inspect the status and a small portion of the response before parsing. The target may have moved the page, denied access, or returned a different document than expected.
  • Too many simultaneous requests: lower the connector limit and use a slower request pace. Asyncio makes it easy to initiate many waits, but the target server and your own network still have finite capacity.
  • Memory rises during a large batch: avoid collecting every full body in a list. Extract and persist only necessary fields, process results in bounded batches, or stream large responses through response.content.
  • asyncio.run() reports an existing event loop: remove that call from the nested environment and await the coroutine through the loop already in use.

Performance, reliability, and cost considerations

Async I/O can improve throughput when many independent requests spend time waiting on network responses, because those waits can overlap. It does not guarantee a fixed multiplier or make every scraper faster: a single URL, CPU-heavy parsing, a slow target, rate limits, connection overhead, and local resource limits can change the result. There is no general speedup figure established here; measure your actual workload if performance matters.

Reliability comes from controlling the batch as much as from concurrency: bound simultaneous connections, set a timeout appropriate to the work, record per-URL errors, and avoid unbounded retries or in-memory accumulation. Treat server responses as data to inspect, not proof that the intended page was fetched. Asyncio and aiohttp are software libraries; their use does not remove the ordinary network, operational, or legal considerations of collecting web data.

Or skip the browser setup

If your task is to capture a rendered page as an image or PDF rather than extract arbitrary fields from HTML, ScreenshotNeo offers a one-request screenshot API. It is not a replacement for an aiohttp scraper when you need to parse page data. Here is the Python request pattern for a screenshot:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

See the ScreenshotNeo API documentation for request options. It removes cookie or consent banners, newsletter popups, and chat widgets before capture; each of those steps can be turned off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and the response identifies the page verdict and billing status in headers. An MCP server provides take_screenshot, get_page_info, and capture_pdf tools for AI agents. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 shots. Sign up for ScreenshotNeo’s free plan.

Frequently Asked Questions

Does asyncio parse HTML?

No. It coordinates asynchronous work; aiohttp retrieves pages, and a separate parser extracts information from the HTML.

Does robots.txt decide whether scraping is legal?

No. The robot parser reports rules and directives in robots.txt; it is not a legal assessment.

Can I use ScreenshotNeo to extract fields from a page?

ScreenshotNeo is a screenshot and PDF capture service. Use an HTTP client and parser when you need to extract arbitrary page data.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.