Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

There is no single best scraper. Choose the tool from the page and the size of the job: use Requests with Beautiful Soup when the fields are already in the initial HTML, Scrapy for repeatable crawls across many URLs, and Selenium when a real browser must run JavaScript or perform actions. A hybrid—direct HTTP first, browser automation only where necessary—usually gives the best balance for mixed sites.

Choose by page type and scale

Start by answering two questions: does the response HTML already contain the data, and how many pages must you process? The answers determine your Python stack more reliably than a library popularity list.

Situation Recommended approach Why it fits Main trade-off
One or a few mostly static pages Requests + Beautiful Soup A small, visible request-and-parse pipeline is easy to inspect and debug. You must add retries, throttling, pagination and storage yourself.
Many pages or domains Scrapy Schedulers, spiders, selectors, exports, caching, cookies, sessions and pipelines are built for crawls. There is more project structure to learn and maintain.
JavaScript-rendered or interactive pages Selenium WebDriver A supported browser executes JavaScript and can click, scroll, submit forms and preserve browser state. Browser processes use more CPU and RAM, and timing can be flaky.
Mixed or partially protected workflow Requests or an API, plus targeted Selenium Fast direct HTTP handles ordinary pages while a browser is reserved for rendered or interactive steps. Session, cookies and hand-off logic add moving parts.

Beautiful Soup is an HTML/XML parser; it does not fetch a URL. Requests performs the HTTP request, and Beautiful Soup turns the returned markup into a navigable tree. Scrapy also uses Request and Response objects, but adds the crawler machinery around them. Selenium WebDriver drives a browser natively.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Inspect the page before writing a scraper

  1. Open the target URL in a browser and view the page source, not only the rendered Elements panel. Search the source for a distinctive value you need.
  2. If the value is present in the initial source, try Requests first. If it is absent, open Developer Tools, select the Network tab, reload, and look for an XHR or Fetch response containing the data. A documented endpoint is preferable to rendering every page.
  3. Check whether pagination, “load more” controls, login state, or scrolling changes the request pattern. Record the URL, method, query parameters, required headers and cookies.
  4. Decide the output before scaling: JSON Lines, CSV, a database table, or an item pipeline. A clear schema prevents a crawl from producing inconsistent records.

Method 1: Requests and Beautiful Soup for static HTML

This is the right starting point when the server sends the content you need in the first response. Install the two packages:

python -m pip install requests beautifulsoup4 lxml

The following script fetches a page, checks the HTTP status, extracts article cards and writes JSON. Replace the selectors with those from your target site.

import json
import requests
from bs4 import BeautifulSoup

url = 'https://example.com/news'
headers = {'User-Agent': 'my-research-bot/1.0 (+https://example.com/contact)'}
response = requests.get(url, headers=headers, timeout=30)
response.raise_for_status()

soup = BeautifulSoup(response.text, 'lxml')
records = []
for card in soup.select('article.card'):
    link = card.select_one('a.card__link')
    title = card.select_one('h2, h3')
    if not link or not title:
        continue
    records.append({
        'title': title.get_text(' ', strip=True),
        'url': link.get('href'),
    })

with open('items.json', 'w', encoding='utf-8') as output:
    json.dump(records, output, ensure_ascii=False, indent=2)
print(f'Wrote {len(records)} records')

Make a small script dependable

  • Use a finite timeout on every request and call raise_for_status() so a 403 or 500 response is not parsed as if it were a page.
  • Reuse a requests.Session() when fetching several pages. It preserves cookies and can reuse connections.
  • Add bounded retries with increasing delays for transient 429 and 5xx responses. Do not retry a permanent 404 indefinitely.
  • Throttle requests and identify the client with a meaningful user agent. Enable and honor the site’s robots policy for your use case.
  • Normalize relative links with urllib.parse.urljoin, and validate that required fields are present before writing a record.

Requests plus Beautiful Soup is deliberately unopinionated. You supply the pagination loop, duplicate detection, checkpointing, data validation and storage. That simplicity is an advantage for a one-off extraction and a maintenance burden for a long-running crawl.

Method 2: Scrapy for repeatable multi-page crawls

Choose Scrapy when link following, pagination, concurrency, retries, caching, exports and pipelines are core requirements. Create a project and spider:

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
python -m pip install scrapy
scrapy startproject catalog
cd catalog
scrapy genspider products example.com

Replace catalog/spiders/products.py with a spider such as this:

import scrapy

class ProductsSpider(scrapy.Spider):
    name = 'products'
    allowed_domains = ['example.com']
    start_urls = ['https://example.com/products']

    def parse(self, response):
        for card in response.css('article.card'):
            yield {
                'name': card.css('h2::text, h3::text').get(default='').strip(),
                'url': response.urljoin(card.css('a::attr(href)').get()),
                'price': card.css('.price::text').get(default='').strip(),
            }
        next_page = response.css('a.next::attr(href)').get()
        if next_page:
            yield response.follow(next_page, callback=self.parse)

Run it as JSON Lines:

scrapy crawl products -O products.jsonl

Settings that matter in production

  • Set ROBOTSTXT_OBEY = True when you want Scrapy’s RobotsTxtMiddleware to filter disallowed requests.
  • Configure download delays or AutoThrottle, retry rules, concurrency and a clear user agent instead of sending an uncontrolled burst.
  • Turn on HTTP caching while developing selectors so you can iterate without repeatedly downloading the same pages.
  • Use item pipelines for validation, deduplication and database writes. Feed exports are convenient for JSON, JSON Lines, CSV and XML.
  • Persist progress and make parsing idempotent. A restart should not silently duplicate every item already stored.

Scrapy is not automatically “better” than Selenium. It is better matched to a crawler workload: asynchronous scheduling and structured Request/Response processing let it coordinate many ordinary pages without opening a browser for each one.

Method 3: Selenium for JavaScript and interaction

Use Selenium when the required data appears only after JavaScript executes, when a click or scroll reveals it, or when the workflow needs forms and browser state. Install Selenium:

python -m pip install selenium

Selenium Manager can obtain a compatible driver for supported browsers. This example waits for a meaningful element rather than assuming that the browser’s load event means a single-page application is finished.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

options = webdriver.ChromeOptions()
options.add_argument('--headless=new')
options.add_argument('--window-size=1440,1200')
driver = webdriver.Chrome(options=options)

try:
    driver.get('https://example.com/dashboard')
    wait = WebDriverWait(driver, 20)
    wait.until(EC.visibility_of_element_located((By.CSS_SELECTOR, 'main .result')))
    rows = []
    for row in driver.find_elements(By.CSS_SELECTOR, 'main .result'):
        rows.append({
            'title': row.find_element(By.CSS_SELECTOR, 'h2, h3').text.strip(),
            'text': row.text.strip(),
        })
    print(rows)
finally:
    driver.quit()

Reliable browser waits

Prefer explicit waits tied to a condition—an element becoming visible, a button becoming clickable, or a loading marker disappearing. A fixed sleep can be too short on a busy run and unnecessarily slow on a fast one. document.readyState only describes the document load; an application may still be fetching and rendering data afterward.

For scrolling, perform a bounded number of scrolls and wait for the item count to increase after each one. For a click, wait for clickability, click, then wait for the expected DOM change. Capture diagnostics on failure: the current URL, page title, a screenshot and the HTML source. Always call quit() in a finally block so orphaned browser processes do not accumulate.

Use a hybrid when only part of the site needs a browser

A practical pattern is to discover the data request in the browser, reproduce it with Requests, and reserve Selenium for the step that genuinely requires a browser. For example, use Selenium to sign in or trigger a token, then pass the resulting cookies to a Requests session for hundreds of detail pages. If an endpoint is undocumented or changes frequently, keep the browser path for that portion and make the hand-off explicit. This reduces resource use without assuming that every page is static.

Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server when your deliverable is a reliable image or PDF rather than extracted fields. One GET request returns PNG, JPEG, WebP or PDF. Before capture it accepts the cookie or consent banner like a visitor and removes more than 60 known consent platforms, newsletter popups and chat widgets; each cleanup step can be switched off. Only clean shots are billed: bot checks or CAPTCHAs, blank pages, timeouts, failed loads and cache hits cost nothing, and the response identifies the outcome with X-Page-Verdict and X-Billed headers.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Call the API with cURL (the complete option reference is in the ScreenshotNeo documentation):

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

The same request in Python:

import requests

r = requests.get(
    'https://api.screenshotneo.com/v1/shot',
    params={'access_key': 'YOUR_API_KEY', 'url': 'https://stripe.com'},
    timeout=90,
)
r.raise_for_status()
open('shot.webp', 'wb').write(r.content)

And in Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`HTTP ${res.status}`);
const fs = await import('node:fs/promises');
await fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer()));

The service also provides an MCP server for AI agents such as Claude, Cursor and other MCP clients, with take_screenshot, get_page_info and capture_pdf tools. Options include full-page capture with lazy images loaded, CSS-selector element capture, dark mode, 12 device presets or a custom viewport, retina scale, PDF paper size/margins/orientation/page ranges, custom CSS and JavaScript, pre-capture clicks, hidden selectors, waits for a selector/delay/network idle, blocking ads/trackers/requests/resource types, custom headers/cookies/user agent/Authorization, timezone and geolocation, transparent backgrounds, resizing, selectable cache TTL, signed image links, asynchronous jobs with signed webhooks, bulk capture of up to 100 URLs per call, a usage API and an OpenAPI specification. Parameter names used by other screenshot APIs are accepted to ease migration.

Plan Allowance and price
Free 1,000 screenshots/month, no card
Starter $5 for 3,000
Growth $15 for 15,000
Pro $39 for 60,000
Scale $99 for 250,000
Business $249 for 1,000,000

Yearly billing gives two months free, and every feature is available on every plan. If screenshots replace a browser-rendering pipeline, the free allowance lets you validate the workflow without a card. Create a free ScreenshotNeo account to start with 1,000 screenshots a month.

Performance, reliability and cost decisions

HTTP parsers

Requests has the lowest per-page overhead because it does not start a browser. Keep concurrency conservative, reuse sessions, cache during development and measure server response time before increasing workers. Parsing errors are generally easier to reproduce because the raw response can be saved.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Scrapy crawlers

Scrapy’s scheduler, asynchronous requests, caching and feed exports reduce the amount of infrastructure you must write. Its speed still depends on the target site’s latency, allowed concurrency and your parsing and storage work; there is no universal cross-tool speed number. Monitor response codes, queue depth, retry counts and item-validation failures.

Browsers

Selenium consumes substantially more resources than direct HTTP and introduces nondeterministic timing. Limit parallel browser instances, reuse a session only when state can safely be shared, and shut down every driver. Store screenshots and HTML on failures so a selector change can be diagnosed instead of retried blindly.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Troubleshooting common failures

Beautiful Soup finds no elements

Inspect the saved response.text. If the value is missing, it is probably rendered by JavaScript or supplied by a separate request; find that request or switch only that step to Selenium. If the value is present, check selectors, namespaces and whether the markup differs between pages.

Requests returns 403 or 429

Do not respond by rapidly rotating requests. Verify the site’s rules, identify your user agent, slow the crawl, honor retry-after guidance and use an authenticated, documented API when available. A browser is not a guaranteed or appropriate way around access controls.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Scrapy follows links forever

Restrict allowed_domains, select only the intended next-page link, canonicalize URLs and maintain a visited or item key. Inspect redirect chains and query parameters that create calendar or filter loops.

Selenium times out

Confirm the browser and driver start in the same environment, then wait for a specific element or state rather than page-load completion. Check iframe boundaries, shadow DOM, consent dialogs and authentication redirects. Increase the timeout only after identifying which condition is genuinely slow.

The browser sees a CAPTCHA or blank page

Treat it as an access or availability failure, not a selector bug. Stop retry storms, follow the site’s permitted access path and record the failure for review. For a screenshot workflow, ScreenshotNeo reports bot checks, blank pages and failed loads as unbilled outcomes through its response headers.

Results are duplicated or incomplete

Assign a stable key such as a canonical URL, validate required fields before writing, checkpoint progress and make reruns idempotent. Compare expected pagination counts with stored item counts and retain the source URL for each record.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Questions developers still ask

Can a scraper use a login?

Only when you are authorized to access the account and the site’s terms permit automation. Keep credentials out of source code, use a session or browser profile deliberately, and protect any cookies or tokens written to disk.

When is an API preferable to scraping?

Use a documented API when it supplies the fields, pagination and usage rights you need. It is usually more stable than selectors tied to presentation markup; scraping remains useful when no suitable interface exists and the access is permitted.

How should I test selector changes?

Keep representative HTML fixtures or cached responses, run parser tests against them, and add a small live smoke test separately. This lets you detect a markup change without making every unit test dependent on network availability.

Frequently Asked Questions

What should I learn first: Beautiful Soup, Scrapy or Selenium?

Learn Requests with Beautiful Soup first if you are new and your target is static. Move to Scrapy when you need a maintained crawl, and add Selenium only for browser execution or interaction.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Is Selenium faster than Scrapy?

They solve different problems, and no authoritative cross-tool benchmark applies to every site. Scrapy avoids browser overhead for ordinary HTTP pages; Selenium is necessary when JavaScript or interaction is the requirement.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.