Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Asynchronous web scraping overlaps network waits instead of handling every request one at a time. An event loop runs coroutines, and while one HTTP request is waiting for DNS, a connection, or response bytes, another task can run. This can improve the use of time for I/O-bound crawls, but it is not parallel CPU execution, does not accelerate parsing by itself, and has no universal speedup percentage. Real throughput depends on latency, server limits, connection settings, response size, and your error-handling design.

How asynchronous scraping works

A conventional scraper sends a request, waits for the response, parses it, and then starts the next request. An asynchronous scraper represents that work as coroutines. At each await, the event loop can suspend the current coroutine and run another ready task.

  • Concurrency: multiple operations make progress during overlapping waits.
  • Parallelism: work executes simultaneously on multiple CPU cores. Asyncio does not provide this for ordinary Python code.
  • I/O-bound work: HTTP, DNS, file, and database waits are where async usually helps.
  • CPU-bound work: parsing, compression, or machine-learning transforms may need processes, threads, or an external worker; adding await alone will not make them faster.

Async also does not bypass robots.txt, authentication, rate limits, bot defenses, or a site’s terms. It merely changes how your program waits.

When async is a good fit

Use it for many independent HTTP requests

Price checks, link checks, sitemap fetches, and API-backed crawls often spend most of their time waiting on remote systems. A shared client session can reuse connections, while bounded tasks keep memory and server load predictable.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Prefer another design for CPU-heavy pipelines

If each response requires expensive HTML parsing or image analysis, keep network fetching asynchronous and move CPU work to a bounded process or worker pool. Measure the complete pipeline rather than assuming that more concurrent requests improve end-to-end time.

A bounded Python scraper with aiohttp

Install the client with python -m pip install aiohttp. This example uses one session, a semaphore, a connector limit, a timeout, status checks, and explicit cleanup. Replace the example URLs with targets you are allowed to fetch.

import asyncio
from typing import Iterable

import aiohttp

URLS = [
    "https://example.com/",
    "https://www.python.org/",
    "https://www.wikipedia.org/",
]

async def fetch(session: aiohttp.ClientSession, url: str,
               gate: asyncio.Semaphore) -> tuple[str, int | None, str | None]:
    async with gate:
        try:
            async with session.get(url, allow_redirects=True) as response:
                text = await response.text(errors="replace")
                if response.status >= 400:
                    return url, response.status, None
                return url, response.status, text
        except (aiohttp.ClientError, asyncio.TimeoutError) as exc:
            return url, None, f"{type(exc).__name__}: {exc}"

async def main(urls: Iterable[str]) -> None:
    timeout = aiohttp.ClientTimeout(total=30)
    connector = aiohttp.TCPConnector(limit=20, limit_per_host=4)
    gate = asyncio.Semaphore(10)
    async with aiohttp.ClientSession(
        timeout=timeout,
        connector=connector,
        headers={"User-Agent": "ExampleResearchBot/1.0"},
    ) as session:
        tasks = [asyncio.create_task(fetch(session, url, gate)) for url in urls]
        for url, status, result in await asyncio.gather(*tasks):
            if status is None:
                print(f"{url}: failed: {result}")
            elif result is None:
                print(f"{url}: HTTP {status}")
            else:
                print(f"{url}: HTTP {status}, {len(result)} characters")

if __name__ == "__main__":
    asyncio.run(main(URLS))

The semaphore limits entries into the request section to 10. The connector adds separate total and per-host safeguards. Both are intentional: a task limit alone does not necessarily constrain every connection created by a more complex program.

Concurrency controls and backpressure

Semaphores

asyncio.Semaphore(n) protects a section so no more than n coroutines enter it at once. Choose a value from the target’s documented limits, your bandwidth, response sizes, and observed error rates—not from a universal “safe” number.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

aiohttp connector limits

The current aiohttp client reference documents a total connection limit of 100 by default and a per-host limit of 0, meaning no per-host cap by default. Those are library defaults, not recommendations. Set both explicitly for a crawl.

Batches and queues

Creating one task for every URL in a million-item crawl consumes memory and makes cancellation difficult. Feed URLs through a bounded asyncio.Queue, or process batches and keep only a controlled number of tasks alive. Backpressure is part of correctness: it prevents your producer from outrunning the network and storage stages.

Delays and politeness

Concurrency and request rate are different. Add per-host delays or a rate limiter where required, honor published crawl instructions, and stop or slow down when a service returns throttling responses.

Failures, cancellation, and cleanup

gather() behavior

By default, asyncio.gather() propagates the first exception while other submitted awaitables may continue running. Catch expected request errors inside the worker, as the example does, or call gather(..., return_exceptions=True) when partial results are the desired output.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

TaskGroup for structured failure handling

Python’s asyncio.TaskGroup provides stronger structured-concurrency behavior: when one grouped task fails, remaining tasks are cancelled and the group raises after cleanup. Use it when a partial crawl is invalid if any required task fails. Use a queue and durable result store when individual failures should be recorded and the crawl should continue.

Always close sessions

An async with ClientSession(...) block closes the session and connector even when a task raises. Reusing one session for a logical unit of work avoids needless handshakes and prevents leaked sockets.

Retries without making an outage worse

Retry transient network errors, connection resets, and selected 5xx responses with exponential backoff and random jitter. Do not blindly retry 401, 403, 404, malformed requests, or a persistent parsing error. Respect Retry-After when supplied. Cap attempts, record the final reason, and make writes idempotent so a retry cannot duplicate data.

Set separate connect and total timeouts when the client and workload warrant it. A long total timeout multiplied by thousands of tasks can hold memory for a long time; a very short timeout turns slow but healthy pages into false failures.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Choosing aiohttp, Scrapy, or another design

Question Direct asyncio plus aiohttp Scrapy
Best scope A focused fetch, parse, and export workflow A crawler with scheduling, downloader components, middleware, pipelines, and broader orchestration
Concurrency controls Connector limits, semaphores, queues, and application logic Framework concurrency and delay settings plus project configuration
Runtime integration Your application’s asyncio event loop Its documented runner and reactor/event-loop configuration
Operations You build persistence, monitoring, retries, and scheduling Many crawler concerns are framework components; deployment still remains your responsibility

Scrapy supports async def callbacks and other coroutine entry points. Its documentation warns that asyncio-dependent libraries may require asyncio support to be enabled, and that runner choice depends on the Twisted reactor or an existing asyncio loop. Do not start a second event loop inside an already running application; use the integration path documented for the Scrapy version you installed. Scrapy’s project site also describes pushing spiders to Scrapy Cloud and scheduling runs as a managed deployment option.

Access rules and responsible crawling

Python’s urllib.robotparser.RobotFileParser can answer whether a user agent may fetch a URL under the site’s robots.txt rules. That is a technical signal, not a complete legal assessment. Also check terms, authentication requirements, privacy obligations, copyright constraints, and applicable law. Identify your bot, cache responses where appropriate, avoid collecting unnecessary personal data, and provide a contact address.

Common problems and fixes

“Event loop is already running”

This occurs when code calls asyncio.run() from an environment that already owns a loop, such as a notebook, async web server, or Scrapy integration. Make the caller async and use await main(), or follow the host framework’s documented runner.

Too many connections or 429 responses

Lower the semaphore and connector limits, add per-host pacing, honor Retry-After, and verify that redirects are not multiplying requests. A library’s default limit is not evidence that a target permits that rate.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Tasks never finish

Set explicit timeouts, inspect DNS and TLS errors, and ensure every queue item calls task_done() in a finally block. Check that response bodies are consumed or the response context is exited so connections return to the pool.

Memory grows during a large crawl

Do not build an unbounded task list or retain every response body. Use a bounded queue, stream or discard content after extraction, and persist results incrementally.

Works in a script but not in Scrapy

Check the installed Scrapy documentation for reactor and coroutine configuration. Mixing a Twisted runner, an asyncio-only client, and an already running loop without the supported adapter commonly causes runtime errors.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup

If your job is to capture rendered pages rather than build a crawler, ScreenshotNeo provides a website screenshot API and MCP server. One GET request returns PNG, JPEG, WebP, or PDF. Before capture it accepts cookie or consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each cleanup step can be turned off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers report the page verdict and whether it was billed.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Its MCP server exposes take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, or another MCP client. Options include full-page lazy-image loading, CSS-selector element capture, dark mode, device presets or custom viewports, retina scale, PDF paper settings and page ranges, custom CSS or JavaScript, clicks, waits, blocked resources, headers, cookies, user agents, authorization, timezone, geolocation, transparent backgrounds, resizing, chosen cache TTLs, signed image links, asynchronous jobs with signed webhooks, bulk capture of up to 100 URLs per call, a usage API, and an OpenAPI specification.

Use the parameter names documented at ScreenshotNeo’s documentation. cURL:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

The Free plan includes 1,000 shots per month with no card; paid plans start at $5 for 3,000 shots. Every feature is on every plan. Create a free ScreenshotNeo account.

FAQ

Does asynchronous scraping hide my IP address?

No. Async changes scheduling inside your client; it is not a proxy, anonymity, or identity service.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Can async scraping bypass JavaScript challenges?

Not by itself. An HTTP client may receive a challenge page instead of the content. Browser automation or an authorized rendering service may be required, subject to the site’s rules.

Should every scraper use the maximum concurrency available?

No. The useful setting is the highest bounded level that remains reliable and permitted for the target and your own network and storage systems.

Frequently Asked Questions

Does asynchronous scraping hide my IP address?

No. Async changes scheduling inside your client; it is not a proxy, anonymity, or identity service.

Can async scraping bypass JavaScript challenges?

Not by itself. An HTTP client may receive a challenge page instead of the content. Browser automation or an authorized rendering service may be required, subject to the site’s rules.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Should every scraper use the maximum concurrency available?

No. The useful setting is the highest bounded level that remains reliable and permitted for the target and your own network and storage systems.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.