Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

There is no single best Python web-scraping library. The right choice depends on whether you need to fetch static HTML, parse markup, run JavaScript, make concurrent requests, or coordinate a large crawl. For a simple static page, combine Requests or HTTPX with Beautiful Soup. Choose Playwright or Selenium when browser rendering is required, and Scrapy when you need a complete crawl framework.

This guide maps each library to its proper role, shows runnable starting code, explains trade-offs without unsupported speed claims, and gives a practical decision process.

Choose by the job, not by the package name

Need Good starting point What it provides Important limitation
Fetch static pages Requests or HTTPX HTTP requests and response objects Neither client parses HTML or executes browser JavaScript.
Extract data from HTML Beautiful Soup or Scrapy selectors (Parsel) Text, attributes and nodes using parser APIs Beautiful Soup is forgiving and approachable; Scrapy documentation notes a speed drawback compared with its selectors, but no general benchmark winner is established.
Fetch concurrently HTTPX Async requests and concurrency patterns Concurrency must respect the target site’s capacity and access rules; it still does not render JavaScript.
Render JavaScript or interact with a page Playwright or Selenium Real browser execution, clicks, waits and scripted interaction Browser binaries, startup time and runtime overhead add operational complexity.
Coordinate a multi-page crawl Scrapy Requests, extraction, scheduling and crawl workflow It is a framework, not a drop-in replacement for an HTTP client or standalone parser.

These are different layers rather than interchangeable products. The role-based comparison is summarized at this tool overview; Scrapy’s official selector documentation explains its CSS and XPath support and Parsel/lxml foundation at docs.scrapy.org.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Start with static HTML: Requests plus Beautiful Soup

First inspect the response you receive. If the required text or attributes are already in the returned HTML, a browser is unnecessary. Requests is a synchronous HTTP client; Beautiful Soup builds a navigable object from the markup and handles malformed HTML reasonably well.

import requests
from bs4 import BeautifulSoup

url = "https://example.com/news"
response = requests.get(url, timeout=30, headers={"User-Agent": "Mozilla/5.0"})
response.raise_for_status()
soup = BeautifulSoup(response.text, "html.parser")
for item in soup.select("article h2 a"):
    print(item.get_text(" ", strip=True), item.get("href"))

Use a session when requesting several pages so connection handling and shared headers are centralized. Set a timeout, check status codes, and select stable attributes rather than brittle positional selectors.

When Beautiful Soup is the better fit

  • You are learning or writing a small one-off extractor.
  • The page is mostly static and forgiving parsing is useful.
  • You want a high-level API such as select(), find() and get_text().

When to use Parsel selectors instead

Scrapy selectors expose CSS and XPath and are backed by Parsel, which uses lxml underneath. They are a natural choice inside Scrapy and can also suit code that needs precise XPath expressions. Scrapy’s documentation describes Beautiful Soup as popular and tolerant of bad markup, while stating: “BeautifulSoup is a very popular web scraping library among Python programmers which constructs a Python object based on the structure of the HTML code and also deals with bad markup reasonably well, but it has one drawback: it’s slow.” Treat that as documentation guidance, not a universal benchmark; workload, parser settings and I/O often dominate.

Use HTTPX when asynchronous fetching matters

HTTPX offers synchronous and asynchronous interfaces. Async is valuable when your program spends most of its time waiting for many independent network responses. It does not turn a client-rendered application into static HTML.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
import asyncio
import httpx

async def fetch(client, url):
    response = await client.get(url, timeout=30)
    response.raise_for_status()
    return url, response.text

async def main():
    urls = ["https://example.com/a", "https://example.com/b"]
    limits = httpx.Limits(max_connections=10, max_keepalive_connections=5)
    async with httpx.AsyncClient(limits=limits, headers={"User-Agent": "my-research-bot/1.0"}) as client:
        for url, html in await asyncio.gather(*(fetch(client, u) for u in urls)):
            print(url, len(html))

asyncio.run(main())

Bound concurrency, retry only transient failures, and add backoff. A large task queue is not permission to ignore robots directives, terms, authentication boundaries or rate limits for the site you access.

Choose Playwright or Selenium for JavaScript-rendered pages

If the data is absent from the initial response because scripts populate it in a browser, use browser automation. Playwright and Selenium can wait for elements, click controls, submit forms and capture the rendered DOM. Browser automation is heavier than HTTP plus parsing, so confirm the need by inspecting the response or browser network activity first.

Playwright example

from playwright.sync_api import sync_playwright

with sync_playwright() as p:
    browser = p.chromium.launch(headless=True)
    page = browser.new_page()
    page.goto("https://example.com/dashboard", wait_until="networkidle", timeout=60_000)
    page.locator("table tbody tr").first.wait_for()
    for row in page.locator("table tbody tr").all():
        print(row.inner_text())
    browser.close()

Install the package and its browser binaries according to the current Playwright documentation. Selenium is a sound alternative when your organization already uses WebDriver, Grid or Selenium-specific integrations. Neither should be selected merely because it is perceived as faster: the available evidence does not establish a controlled, cross-workload winner.

Browser edge cases

  • Wait for a meaningful selector, not an arbitrary sleep, whenever possible.
  • Handle consent dialogs and authentication explicitly and lawfully.
  • Expect more CPU, memory and failure modes than an HTTP client.
  • Persist only the data you need; rendered pages can contain secrets and personal information.

Use Scrapy for an actual crawl

Scrapy coordinates requests, scheduling, callbacks, extraction, item pipelines and crawl policy. It is the strongest fit when you are following links across many pages rather than writing a single fetch-and-parse script.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
import scrapy

class ArticleSpider(scrapy.Spider):
    name = "articles"
    start_urls = ["https://example.com/blog"]

    def parse(self, response):
        for article in response.css("article"):
            yield {
                "title": article.css("h2 a::text").get(default="").strip(),
                "url": response.urljoin(article.css("h2 a::attr(href)").get()),
            }
        next_page = response.css("a.next::attr(href)").get()
        if next_page:
            yield response.follow(next_page, callback=self.parse)

Scrapy selectors support both CSS and XPath; see the official selector guide. Scrapy can be combined with browser tooling for pages that require JavaScript, but adding a browser to every request sacrifices the efficiency that made a crawl framework useful.

Project scale questions

  • How many URLs are there, and are they discovered through links?
  • Do you need deduplication, throttling, retries, pipelines or resumable jobs?
  • Will extraction, storage and monitoring be maintained as a service?

The Scrapy project site currently reports version 2.19.0 in September 2026 and describes an experimental aiohttp-based download handler. Verify the release page and compatibility before pinning versions because project details change.

A practical selection workflow

  1. Inspect the response. Request one page and search its HTML for the value you need.
  2. Parse with the smallest layer. Use Beautiful Soup or selectors if the value is present.
  3. Add async only for a demonstrated I/O workload. HTTPX can provide concurrency, subject to site limits.
  4. Render only when required. Choose Playwright or Selenium for scripts, clicks or browser-only state.
  5. Adopt Scrapy when crawl coordination is the problem. A parser alone does not schedule and govern a crawl.
  6. Measure your own workload. Compare success rate, extraction correctness, resource use and maintenance cost; do not claim a universal fastest library without a reproducible benchmark.

Reliability, ethics and operating costs

  • Use explicit timeouts and status checks; log URL, status, latency and parser failures.
  • Respect robots.txt where applicable, contractual terms, authentication boundaries and applicable law.
  • Throttle requests and identify your client honestly. Retries should be limited and backoff-based.
  • Cache responses where permitted, normalize encodings, and write tests against representative HTML fixtures.
  • Expect browser automation to consume substantially more compute and require browser-version maintenance than direct HTTP.

No source here provides an independent benchmark, package-wide cost model or legal advice for a particular target. Make those decisions with target-specific documentation and your own measurements.

Or skip the browser setup

If your goal is a clean screenshot rather than extracting fields, ScreenshotNeo provides a single website-screenshot API call. It accepts cookie and consent banners like a visitor, then removes more than 60 known consent platforms, newsletter popups and chat widgets before capture. Bot checks, blank pages, timeouts, failed loads and cache hits are not billed, and response headers identify the page verdict and billing result. Its MCP server exposes take_screenshot, get_page_info and capture_pdf to Claude, Cursor and other MCP clients.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

See the ScreenshotNeo API documentation for all options, including PNG/JPEG/WebP or PDF output, full-page and element capture, device presets, retina scale, dark mode, custom CSS and JavaScript, clicks, waits, request blocking, cookies, headers, geolocation, timezone, transparent backgrounds, resizing, caching, signed links, asynchronous webhooks, bulk capture and usage reporting.

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

The Free plan includes 1,000 screenshots monthly without a card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Common failures and fixes

HTML contains no target data

The page likely renders it with JavaScript or loads an API call after startup. Inspect network requests, then use the underlying permitted endpoint or browser automation.

Selectors return nothing

Print a response excerpt, verify the selector against the actual document, account for namespaces or shadow DOM, and avoid assuming a browser’s post-render DOM equals the original response.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Requests time out or returns 403

Check the URL, timeout and redirect behavior; slow down, use a truthful User-Agent, and follow the site’s access policy. Do not attempt to defeat bot protection.

Scrapy crawl loops or grows unexpectedly

Restrict allowed domains, normalize URLs, filter duplicates, cap pagination and inspect which callback schedules each request.

Browser runs locally but fails in production

Install compatible browser binaries, run headless with required system dependencies, set realistic timeouts, and capture console/page errors for diagnosis.

Frequently Asked Questions

Should I use Beautiful Soup or Scrapy?

Use Beautiful Soup for a small extractor and Scrapy when scheduling, following links, throttling and pipelines are central. They solve different layers, so the page count and workflow matter more than a blanket winner.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Which Python library scrapes JavaScript-rendered pages?

Playwright and Selenium run a browser and can wait for scripts and interactions. HTTPX, Requests and Beautiful Soup alone do not execute client-side JavaScript.

Is HTTPX faster than Requests?

HTTPX supports asynchronous patterns, which can improve an I/O-bound design. The available sources do not establish a universal speed ranking; measure your workload and respect target limits.

What should I install first?

For static HTML, start with Requests and Beautiful Soup. Add HTTPX for an async architecture, Playwright or Selenium for browser-only content, and Scrapy for a coordinated crawl.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.