October DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsWindows FixRecommendedWindows errors stealing your time? Find the fix fastScan stability, cleanup and performance issues.Fix NowOctober DealsAmazon USDeal season is back - check today's better picksAmazon US: current deals, useful picks and tech finds.See Picks×
Skip to content
aiohttp

Python Asyncio for Web Scraping and Browser Automation

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use asyncio with an asynchronous HTTP client such as aiohttp when the data you need is available from ordinary web requests. Use an async browser automation library such as Playwright when the result depends on browser rendering, interaction, or browser-visible output. If your project needs a crawling framework, consider Scrapy and check its event-loop requirements before combining it with Playwright.

How do I use asyncio for web scraping?

asyncio is Python’s library for writing concurrent code with async and await. It is often useful for IO-bound work such as waiting for network responses: while one request waits, the event loop can run other eligible tasks. It also provides APIs for network I/O, subprocesses, queues, and synchronization. See the Python asyncio documentation.

For a scraper, the common pattern is to create asynchronous request tasks, await their responses, and parse the returned content. Concurrency does not make CPU-heavy parsing or blocking synchronous calls non-blocking. Keep blocking work out of the event loop or move it to an appropriate worker mechanism if it becomes a bottleneck.

Runnable example: fetch several pages with aiohttp

aiohttp is an asyncio-based HTTP client/server library. Its client flow uses a ClientSession, awaits a request, then reads the response body. Install it with python -m pip install aiohttp, save the following as scrape.py, then run python scrape.py:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
import asyncio
import aiohttp

URLS = [
    "https://example.com/",
    "https://www.iana.org/domains/reserved",
]

async def fetch(session, url):
    async with session.get(url) as response:
        response.raise_for_status()
        html = await response.text()
        return url, response.status, html

async def main():
    timeout = aiohttp.ClientTimeout(total=30)
    connector = aiohttp.TCPConnector(limit=10)

    async with aiohttp.ClientSession(
        timeout=timeout,
        connector=connector,
    ) as session:
        results = await asyncio.gather(
            *(fetch(session, url) for url in URLS)
        )

    for url, status, html in results:
        print(f"{status} {url}: {len(html)} characters")

if __name__ == "__main__":
    asyncio.run(main())

The connector limit and timeout in this example are explicit example settings, not universal tuning recommendations. Choose limits and timeouts based on the target, your workload, and the site’s rules. raise_for_status() makes unsuccessful HTTP responses visible instead of treating their bodies as successful page data.

Why reuse a session?

A ClientSession owns client-side state and connection-pooling behavior. Reusing one session across a batch is the ordinary aiohttp pattern; creating a new session for every URL adds unnecessary setup and forfeits reuse. Close the session with an asynchronous context manager as shown so its resources are released.

Bound work and handle failures

Launching an unbounded task for every URL can overwhelm your machine or place an unnecessary load on a site. Set a deliberate concurrency limit, and consider a semaphore or bounded queue when the number of inputs is large. Decide how your crawler should handle timeouts, network errors, non-success status codes, retries, and malformed content. Retry only errors that are plausibly transient, and use a limit and delay policy rather than retrying indefinitely. There is no single concurrency or retry setting that is appropriate for every target.

Should I use aiohttp or Playwright?

Choose based on where the information comes from and what output you need. A page that happens to use JavaScript does not automatically require a browser: if its data is available through reproducible HTTP requests, direct fetching may still be the simpler route. A browser is warranted when browser execution or interaction is part of the requirement.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Need Likely approach Why
Data is present in ordinary HTTP responses asyncio with aiohttp Fetch responses directly and parse the data you need.
Browser rendering, interaction, or a browser-visible artifact Playwright’s async API Drive a real browser engine to perform actions or observe the rendered result.
A crawler needs crawling-framework components Scrapy, with its asyncio support as appropriate Use the framework when its crawling components suit the project, while checking reactor and event-loop compatibility.

Scrapy recommends reproducing a page’s underlying data requests when practical: doing so can reduce parsing time and network transfer while yielding structured, complete data. A headless browser is useful when browser behavior is required, including when you need a screenshot as seen in a browser. Read Scrapy’s guide to selecting dynamically loaded content before deciding which route fits a site.

How do I automate a browser with Python asyncio?

Playwright’s async Python API drives Chromium, Firefox, and WebKit. Its driver runs in a subprocess, so browser automation has more operational setup and resource overhead than fetching response bodies directly. Install Playwright and its browser binaries using the current instructions in the Playwright Python documentation. The following standalone pattern opens a page and prints its title:

import asyncio
from playwright.async_api import async_playwright

async def main():
    async with async_playwright() as playwright:
        browser = await playwright.chromium.launch()
        page = await browser.new_page()
        await page.goto("https://example.com/", wait_until="domcontentloaded")
        print(await page.title())
        await browser.close()

if __name__ == "__main__":
    asyncio.run(main())

asyncio.run(main()) is the standard entry point for a standalone coroutine program. If a framework, notebook, or other host already manages an event loop, do not start a second one blindly; integrate with the host’s async model instead. Use await at asynchronous boundaries, and select page navigation and waiting conditions that match the behavior you need. A navigation event alone does not guarantee that every application-specific request or widget has finished.

Prefer the request behind a page when possible

When a browser page obtains structured data from an ordinary endpoint, inspect the page’s network behavior and determine whether the same request can be made directly. Reproducing that request can avoid launching a browser and transferring or parsing a full rendered page. Use Playwright instead if the data or artifact genuinely depends on browser execution, client-side interactions, or rendered state that you cannot reliably reproduce with requests.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How do Scrapy and Playwright work together?

Scrapy is a crawling framework; its asyncio support does not mean every Playwright and Scrapy configuration can share any event loop. For integration, Scrapy recommends scrapy-playwright when you want browser handling while retaining more Scrapy components. Check the integration’s current installation and usage instructions alongside the Scrapy version and reactor your project uses.

Windows event-loop compatibility

On Windows, Playwright’s documentation requires ProactorEventLoop because its driver runs in a subprocess. Scrapy’s Windows asyncio reactor uses SelectorEventLoop. Those requirements conflict when combined in that configuration. Before adopting the combination, verify the configured reactor and loop for the versions and environment you will actually run.

Scrapy documents running without its Twisted reactor as a way to avoid this particular conflict, but that choice comes with feature limitations. It is not a universal fix: check which reactor-dependent components your project needs and confirm the current Scrapy guidance at Scrapy’s asyncio documentation.

Or skip the browser setup

If your goal is to capture a website screenshot rather than build and maintain browser automation, ScreenshotNeo offers a screenshot API and MCP server. Its one-call HTTP request can return an image or PDF; its documented API details and parameters are at ScreenshotNeo’s API documentation.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
curl -G "https://api.screenshotneo.com/v1/shot" 
  -d access_key=YOUR_API_KEY 
  --data-urlencode url=https://stripe.com 
  -o shot.webp

ScreenshotNeo accepts cookie or consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each of those steps can be turned off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and responses identify the page verdict and billing status in headers. Its MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 screenshots.

Create a free ScreenshotNeo account to try it with 1,000 screenshots a month and no card.

Performance, reliability, and responsible crawling

Async concurrency is not a speed guarantee

Async code can overlap time spent waiting on network I/O, but it does not guarantee a specific speedup or throughput. Actual results depend on the target, connection behavior, response sizes, parsing work, and concurrency choices. Measure your own workload; do not raise concurrency simply because the program uses asyncio.

Make failures observable

  • Set a total request timeout and handle timeout exceptions.
  • Check response status before parsing a body as valid page data.
  • Record the URL and failure type so one bad response does not disappear silently.
  • Limit retries and distinguish transient network failures from persistent errors.
  • Validate extracted fields; a successful HTTP response can still contain an error page or unexpected markup.
  • For browser automation, close pages and browsers reliably, including on exceptions, and wait for the specific state your task needs rather than adding arbitrary sleeps by default.

Respect access rules

Concurrency does not bypass access controls or grant permission to collect data. Review the target site’s applicable terms, access policies, and legal requirements, and keep request rates appropriate. The technical documentation cited here does not determine what is permitted for any particular site.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Troubleshooting common problems

“RuntimeError: asyncio.run() cannot be called from a running event loop”

Your environment already owns an event loop. In a notebook or async framework, call and await the coroutine through that environment’s supported mechanism instead of invoking asyncio.run() again.

Playwright works alone but fails with Scrapy on Windows

Check whether the configured Scrapy reactor uses SelectorEventLoop while Playwright needs ProactorEventLoop. Review Scrapy’s documented no-Twisted-reactor option and its limitations, or use the recommended Scrapy integration after confirming its requirements for your environment.

A page is blank or missing data after a successful request

Inspect the response status, body, and network requests. The requested URL may return a shell whose data is loaded separately, a redirect, or a different response than expected. Reproduce the underlying data request if practical; otherwise use a browser and wait for the required rendered state.

Requests hang or the crawler consumes too many resources

Set finite timeouts, bound the number of in-flight requests, and ensure sessions and browser resources are closed. Large batches should not become one unbounded gather of tasks; feed work through a bounded queue or process it in controlled batches.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Retries keep failing

Confirm whether the failure is transient before retrying. Persistent status errors, changed page structure, access denial, and invalid URLs will not be repaired by an endless retry loop. Surface the failure and decide whether to skip, stop, or investigate it.

Frequently Asked Questions

Does a JavaScript website always require Playwright?

No. Use direct HTTP requests if they expose the required data; use a browser when the required result depends on browser execution or interaction.

Can I use aiohttp and Playwright in the same Python project?

Yes, but when Scrapy is also involved, event-loop and reactor compatibility matters, particularly on Windows. Verify the configuration and integration guidance for the versions you use.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Read next

Recommended PC Tool
Recommended PC Tool
PC Slower Than It Used to Be?Free scan - under a minute
Outdated Drivers Are Slowing You DownFree scan - exact matches

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.