Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

For a small, static page, the most direct Python workflow is Requests plus Beautiful Soup: request the HTML with a timeout, verify the HTTP status, select the fields you need, validate the results, and save structured data. Use Python’s standard-library urllib when you need minimal dependencies, and move to Scrapy when you need pagination, link following, scheduling, concurrency controls, or feed exports. If the data is inserted by JavaScript, look for an official API or feed before using browser rendering.

Choose the right Python scraping approach

Start by identifying how the target delivers data and how much you need to collect. An API, downloadable file, RSS feed, or other documented endpoint is usually more stable than extracting presentation HTML. If no suitable interface exists and you are authorized to collect the information, choose a tool that matches the job.

Situation Good starting point Reason
One or a few static pages Requests + Beautiful Soup Requests handles HTTP retrieval and response details; Beautiful Soup parses and searches HTML or XML.
Minimal dependencies or standard-library-only code urllib.request Python can open URLs and read responses without installing a third-party HTTP client. urllib.robotparser can read robots.txt rules.
Pagination, many pages, recurring runs, or structured exports Scrapy It provides spiders, callbacks, selectors, link following, scheduling, crawl-delay and concurrency settings, pipelines, and feed exports.
Content created in the browser Documented API or feed first; browser rendering if necessary A plain HTTP response may not contain content inserted by JavaScript.

Do not begin with browser automation for an ordinary static page. It adds setup and resource use when the server already returns the required HTML.

Before you write a scraper

Define fields and scope

Write down the exact fields you need, the permitted pages, the expected output format, and a low request rate. Restrict collection to what your project requires. Keep a small sample for manual checking so a selector failure cannot silently produce an empty or misleading dataset.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Check for a better data interface

Search the site’s developer documentation, page source, feeds, and downloadable data. An API may provide stable field names, pagination, authentication, and usage rules that HTML extraction does not.

Review access signals

Read the site’s terms and robots.txt, identify your crawler with a clear user agent, and stop if the site signals overload or denies access. The Robots Exclusion Protocol standardized by RFC 9309 is a crawler-preference protocol, not authentication or a legal permission slip. A permitted path does not by itself make collection or reuse lawful, and a disallowed path is a clear signal to avoid crawling it.

Scrape a static page with Requests and Beautiful Soup

Install the libraries

Create an isolated environment, then install the two packages:

python -m venv .venv
# macOS/Linux
source .venv/bin/activate
# Windows PowerShell
# .venvScriptsActivate.ps1
python -m pip install requests beautifulsoup4

Fetch, check, parse, and extract

This complete example expects product cards with the illustrative selectors article.product, h2, and .price. Inspect the authorized target and change the selectors to match its actual markup.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
import csv
import requests
from bs4 import BeautifulSoup

url = "https://example.com/catalog"
headers = {
    "User-Agent": "ExampleResearchBot/1.0 (contact: [email protected])"
}

response = requests.get(url, headers=headers, timeout=10)
response.raise_for_status()

# Requests derives response.text from the response encoding. Inspect or set
# response.encoding if the server's declaration is incorrect.
soup = BeautifulSoup(response.text, "html.parser")
records = []

for card in soup.select("article.product"):
    title_node = card.select_one("h2")
    price_node = card.select_one(".price")
    if not title_node or not price_node:
        continue
    records.append({
        "title": title_node.get_text(" ", strip=True),
        "price": price_node.get_text(" ", strip=True),
    })

if not records:
    raise RuntimeError("No records found; verify the URL and CSS selectors")

with open("products.csv", "w", newline="", encoding="utf-8") as output:
    writer = csv.DictWriter(output, fieldnames=["title", "price"])
    writer.writeheader()
    writer.writerows(records)

print(f"Saved {len(records)} records")

raise_for_status() turns unsuccessful HTTP responses into exceptions instead of allowing an error page to be parsed as if it were data. The timeout is essential: Requests notes that nearly all production code should set one. Its timeout measures inactivity between bytes, not a total deadline for receiving the entire response.

Normalize and validate fields

Use get_text(" ", strip=True) to collapse formatting whitespace, then convert values deliberately. For example, parse a price into a decimal only after removing the correct currency symbol and thousands separator; parse dates with an explicit format or a known date parser. Check required fields, record counts, duplicate identifiers, and representative samples. Save the source URL and retrieval time when provenance matters.

Use Python’s standard library with urllib

When third-party packages are not allowed, urllib.request can retrieve a page and urllib.robotparser can inspect robots.txt. You still need status, timeout, encoding, parsing, and validation logic.

from html.parser import HTMLParser
from urllib.parse import urljoin
from urllib.request import Request, urlopen

class LinkParser(HTMLParser):
    def __init__(self):
        super().__init__()
        self.links = []

    def handle_starttag(self, tag, attrs):
        if tag == "a":
            href = dict(attrs).get("href")
            if href:
                self.links.append(href)

url = "https://example.com/"
request = Request(url, headers={"User-Agent": "ExampleResearchBot/1.0"})
with urlopen(request, timeout=10) as response:
    if response.status != 200:
        raise RuntimeError(f"HTTP status {response.status}")
    charset = response.headers.get_content_charset() or "utf-8"
    html = response.read().decode(charset, errors="replace")

parser = LinkParser()
parser.feed(html)
for href in parser.links:
    print(urljoin(url, href))

The standard library’s html.parser is useful for simple structures, but Beautiful Soup is generally more convenient for searching a document tree and handling varied markup.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Scale to multiple pages with Scrapy

Scrapy is designed for crawls rather than a one-off request. It uses Request and Response objects, callbacks, selectors, link following, asynchronous scheduling, and feed exports.

Create a project and spider

python -m pip install scrapy
scrapy startproject catalog_crawler
cd catalog_crawler
scrapy genspider products example.com

Replace the generated spider with a version that follows a “next” link and yields structured items:

import scrapy

class ProductsSpider(scrapy.Spider):
    name = "products"
    allowed_domains = ["example.com"]
    start_urls = ["https://example.com/catalog"]

    custom_settings = {
        "DOWNLOAD_DELAY": 1.0,
        "CONCURRENT_REQUESTS_PER_DOMAIN": 2,
        "ROBOTSTXT_OBEY": True,
        "FEEDS": {"products.jsonl": {"format": "jsonlines", "encoding": "utf8"}},
    }

    def parse(self, response):
        for card in response.css("article.product"):
            title = card.css("h2::text").get()
            price = card.css(".price::text").get()
            if title and price:
                yield {
                    "url": response.url,
                    "title": " ".join(title.split()),
                    "price": " ".join(price.split()),
                }

        next_url = response.css("a.next::attr(href)").get()
        if next_url:
            yield response.follow(next_url, callback=self.parse)

Run it with scrapy crawl products. Adjust the selectors and domain to the permitted target. Delay, per-domain concurrency, and AutoThrottle settings help prevent an aggressive crawl; validation and item pipelines can enforce consistent output as the project grows. Scrapy’s project site currently labels version 2.19.0 as its latest release (September 2026); pin and review the version appropriate for your deployment rather than assuming this label will remain current.

Handle pagination, JavaScript, and changing markup

Pagination and link discovery

Prefer a documented page parameter or API cursor. In HTML crawls, follow only links that match your allowed domain and URL patterns. Put a maximum-page or maximum-item guard in recurring jobs, and deduplicate canonical URLs so tracking parameters do not create an accidental loop.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

JavaScript-rendered content

If the HTML lacks the records you see in a browser, inspect network requests for a documented data endpoint or feed. Use that endpoint when authorized. If no suitable endpoint exists and browser execution is appropriate, use a rendering tool designed for JavaScript-heavy pages; Scrapy’s ecosystem lists browser rendering as an extension path. Rendering consumes more resources and introduces browser-specific failure modes, so reserve it for pages that actually require execution.

Markup changes

Selectors are coupled to page structure. Assert that required fields exist, alert when record counts fall outside an expected range, retain failed URLs, and review a sample after site redesigns. Never treat an empty result as a successful crawl.

Reliability and failure troubleshooting

Symptom Likely cause Fix
requests.exceptions.Timeout No bytes arrived within the inactivity timeout, or the server is slow. Set a suitable connect/read timeout, retry narrowly with backoff where permitted, and keep a maximum job deadline outside Requests.
HTTP 403, 429, or 5xx Access policy, rate limiting, temporary failure, or server error. Stop or slow down, respect the site’s instructions, verify authorization, and do not attempt to bypass access controls.
Parser returns zero items Wrong selector, an error page, or content loaded by JavaScript. Log status and final URL, save a diagnostic response, inspect the HTML, then use the documented endpoint or an appropriate renderer.
Text is garbled Incorrect or missing encoding declaration. Inspect response.encoding and the document’s encoding metadata; decode using the verified charset.
Duplicate or runaway requests Pagination links, fragments, or tracking parameters create repeated URLs. Normalize and deduplicate URLs, restrict allowed domains and paths, and set crawl limits.
Data shape changes The site changed its markup or returned a different template. Require key fields, monitor counts, preserve failed samples, and update selectors deliberately.

Treat every response as untrusted external input. Do not execute returned scripts, interpolate scraped text into shell commands, or use untrusted values as unrestricted filesystem paths.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup

When your real requirement is a screenshot or PDF of a JavaScript-heavy page rather than structured text, ScreenshotNeo provides a website screenshot API and MCP server. It accepts consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be disabled. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and the response reports the page verdict and billing status in headers. Its MCP tools—take_screenshot, get_page_info, and capture_pdf—work with Claude, Cursor, and other MCP clients.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

One GET request is enough. See the ScreenshotNeo API documentation for all options.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

ScreenshotNeo also supports full-page captures with lazy images, CSS-selector element captures, dark mode, 12 device presets and custom viewports, retina scale, PDF paper and page settings, custom CSS or JavaScript, pre-capture clicks, selector hiding, waits, request and resource blocking, headers, cookies, user agents, Authorization, timezone and geolocation, transparent backgrounds, resizing, chosen cache TTLs, signed image links, asynchronous jobs with signed webhooks, bulk capture of up to 100 URLs per call, usage reporting, and an OpenAPI specification. Every plan includes every feature: 1,000 screenshots per month are free without a card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.

Responsible, lawful scraping

  • Prefer an API, feed, or downloadable dataset when one is available.
  • Read robots.txt and terms, identify your crawler, and use conservative rates.
  • Do not bypass authentication, CAPTCHAs, paywalls, technical blocks, or access controls.
  • Collect only necessary data; protect personal information and define retention and deletion rules.
  • Keep provenance, retrieval times, and error logs so results can be audited.

There is no universal legal answer. Copyright, contract terms, privacy and data-protection law, access controls, intended use, and jurisdiction all matter. The U.S. Copyright Office’s Fair Use Index is a resource for U.S. fair-use decisions and cases, not a blanket ruling that scraping is allowed. For a consequential project, obtain advice based on the specific facts and jurisdiction.

Operational checklist

  • Data source or API checked first.
  • Fields, URL scope, output schema, and retention defined.
  • Clear user agent, robots.txt review, and low request rate configured.
  • Timeouts, status checks, retries, and crawl limits implemented.
  • Selectors tested against representative pages with missing-field handling.
  • Encoding, dates, numbers, duplicates, and record counts validated.
  • Logs, failed URLs, samples, and source timestamps retained.
  • JavaScript rendering used only when an endpoint cannot provide the authorized data.

Frequently Asked Questions

How do I scrape a website with BeautifulSoup?

Fetch the authorized page with an HTTP client such as Requests, call raise_for_status(), create BeautifulSoup(response.text, "html.parser"), select elements with CSS selectors, normalize their text, validate required fields, and save the records.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How do I scrape a page that uses JavaScript?

First inspect the browser’s network activity for a documented API or feed. If none is suitable and browser execution is authorized, use a rendering tool; ordinary Requests parsing cannot see content that never arrives in the initial HTML.

Is web scraping legal?

It depends on the site, data, purpose, and jurisdiction. Robots.txt is guidance for crawlers, not legal permission. Review terms, copyright, privacy, access-control rules, and obtain jurisdiction-specific advice for high-stakes work.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.