October DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsSlow PC?RecommendedPC slow today? Run a repair scan before it gets worseResolve common Windows issues and optimize system performance.Scan NowOctober DealsAmazon USDeal season is back - check today's better picksAmazon US: current deals, useful picks and tech finds.See Picks×
Skip to content
World desk9 min

How to Build a Web Scraper in Python: A Practical Beginner’s Guide

A practical Python scraping workflow: fetch and parse static HTML, extract validated records, follow pagination within clear limits, handle errors, and choose between Requests, Beautiful Soup and Scrapy.
Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A small Python web scraper has four jobs: fetch a page, parse its HTML, extract and validate the fields you need, then save them in a structured format. For a permitted static page or a modest set of pages, requests and Beautiful Soup are a clear starting point. This guide builds that workflow, adds bounded pagination and error handling, and explains when Scrapy is a better fit.

Before you scrape: define the task and check access

Choose a site and data you are permitted to access. Prefer an official API or downloadable dataset if one meets the need. Before writing code, decide which fields to collect, which pages are in scope, how many pages to visit, and what output format you want. A defined boundary helps prevent an accidental site-wide crawl.

  • Inspect the site’s published terms and its robots.txt file. Python’s standard library includes urllib.robotparser for parsing robots.txt; the Python 3.14.8 documentation describes it alongside urllib.request and urllib.parse: Python urllib documentation.
  • Robots.txt communicates crawler access preferences, but it is not a complete statement of legal permission. Google explains that the file tells search engine crawlers which URLs they may access and is used mainly to manage crawler traffic; it also says robots.txt is not a mechanism for keeping a URL out of search results: Google Search Central: Robots.txt Introduction and Guide.
  • Use conservative request volume, honor applicable terms and access controls, and stop if access is denied. There is no universal request rate or legal rule that applies to every site and jurisdiction.

Keep credentials out of source code. Do not attempt to evade a site’s blocks or access controls.

Install the Python packages

Use a virtual environment for a small project, then install Requests and Beautiful Soup:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
#1 Best Overall
Elebase USB to USB C Adapter for iPhone 18 Pro Max,USBC Car Charger Adapter
  • Read Before You Buy — No Video Output: These adapters support charging and USB 2.0 data transfer, but cannot transmit video signals. Except for standard USB webcams (which use USB data only), they are not compatible with HDMI/DisplayPort cables, video-capable USB-C hubs, or docking stations with video output.
  • Convert USB-A Ports to USB-C: Designed to connect USB-C earphones, cables, flash drives, card readers, and other USB-C accessories to standard USB-A ports. Plug-and-play with no drivers or software required.
  • Aluminum Alloy Housing: Built with a sturdy aluminum alloy shell that aids in heat dissipation and protects against daily wear and scratches. Designed to maintain a stable and secure connection.
  • Compact & Travel-Friendly: The ultra-compact design allows the adapter to stay plugged into your device without blocking adjacent ports or adding bulk, reducing wear and tear on your original USB ports.
  • 12-Month Warranty: Backed by a 12-month manufacturer warranty for peace of mind. Designed to meet strict quality control standards for reliable everyday performance.
python -m venv .venv
# macOS or Linux:
source .venv/bin/activate
# Windows PowerShell:
.venvScriptsActivate.ps1
python -m pip install requests beautifulsoup4

Requests provides HTTP request/response handling, including timeouts, status codes, headers, sessions and exceptions. Beautiful Soup parses HTML into a searchable tree. Its documentation describes multiple parser backends; specifying one makes your parsing choice explicit, and different parsers can produce different trees for malformed markup: Requests documentation and Beautiful Soup documentation.

Build a bounded scraper for a page and its next links

The example below fetches a starting page and follows a page’s <a rel="next"> link, if present, without leaving the starting host or exceeding a page limit. It extracts each page title and its links into JSON Lines. The selector for the next link is a convention, not a guarantee: inspect the target site’s HTML and adjust it to the actual markup. The code is an illustrative starting pattern, not a claim that it has been run against a live target.

import json
import time
from urllib.parse import urljoin, urlparse

import requests
from bs4 import BeautifulSoup

START_URL = "https://example.com/"
MAX_PAGES = 5
DELAY_SECONDS = 1
TIMEOUT_SECONDS = 10
OUTPUT_FILE = "pages.jsonl"


def same_host(url, host):
    return urlparse(url).netloc == host


def scrape(start_url):
    host = urlparse(start_url).netloc
    session = requests.Session()
    visited = set()
    current_url = start_url
    records = []

    for _ in range(MAX_PAGES):
        if not current_url or current_url in visited:
            break
        if not same_host(current_url, host):
            print(f"Stopping: URL is outside the starting host: {current_url}")
            break

        visited.add(current_url)
        try:
            response = session.get(current_url, timeout=TIMEOUT_SECONDS)
            response.raise_for_status()
        except requests.exceptions.Timeout:
            print(f"Timed out: {current_url}")
            break
        except requests.exceptions.HTTPError as exc:
            status = exc.response.status_code if exc.response is not None else "unknown"
            print(f"HTTP error {status}: {current_url}")
            break
        except requests.exceptions.RequestException as exc:
            print(f"Request failed for {current_url}: {exc}")
            break

        soup = BeautifulSoup(response.text, "html.parser")
        title = soup.title.get_text(" ", strip=True) if soup.title else None
        links = []
        for anchor in soup.select("a[href]"):
            absolute_url = urljoin(response.url, anchor["href"])
            if urlparse(absolute_url).scheme in {"http", "https"}:
                links.append({
                    "text": anchor.get_text(" ", strip=True),
                    "url": absolute_url,
                })

        if not title:
            print(f"Expected title is missing: {response.url}")
        records.append({"url": response.url, "title": title, "links": links})

        next_anchor = soup.select_one('a[rel~="next"][href]')
        next_url = urljoin(response.url, next_anchor["href"]) if next_anchor else None
        if next_url and (not same_host(next_url, host) or next_url in visited):
            next_url = None
        current_url = next_url
        if current_url:
            time.sleep(DELAY_SECONDS)

    return records


if __name__ == "__main__":
    pages = scrape(START_URL)
    with open(OUTPUT_FILE, "w", encoding="utf-8") as output:
        for page in pages:
            output.write(json.dumps(page, ensure_ascii=False) + "n")
    print(f"Saved {len(pages)} page record(s) to {OUTPUT_FILE}")

Change START_URL, the extraction selectors, and the next-page selector to match the site you are allowed to access. The example uses a finite timeout, checks HTTP status with raise_for_status(), keeps a visited set, restricts navigation to the starting host, caps pages, and pauses between requests. The page limit is a safety boundary for this sample, not a recommended universal crawl size or rate.

Rank #2
Anker USB-C Hub, 5-in-1 USB Hub for Laptops, 4K HDMI Multiport Adapter
  • 5-in-1 USB-C Hub: Experience comprehensive connectivity featuring a Power Delivery input, two USB-A 2.0 ports, a USB-A 3.0 port, and an HDMI port. (Note: The USB-C power delivery input port is only for connecting an external wall charger to power your laptop and cannot power peripheral devices.)
  • 90W Pass-Through Charging: Achieve optimal charging with 90W pass-through power to your laptop, supported by a total input of 100W, with the hub reserving 10W for operational efficiency. (Note: Wall charger not included.)
  • Quick Data Transfers: Accelerate your productivity with rapid data transfers using a high-speed 5Gbps USB 3.0 port and two 480Mbps USB 2.0 ports.
  • 4K HDMI Display: Enhance your visual experience with a hub capable of delivering 4K resolution at 30Hz in both mirror and extend modes. Please note that this hub is compatible with MacBook (macOS 12 and newer), Windows 10 and 11, ChromeOS, and laptops equipped with DP Alt Mode and Power Delivery. Note: This device is not compatible with Linux.
  • What You Get: Anker USB-C Hub (5-in-1, 4K HDMI), welcome guide, 18-month warranty, and our friendly customer service.

How the scraper works

Fetch: make the HTTP request

requests.Session() lets the script reuse a session across requests. Each call uses a finite timeout, since a network request can otherwise wait indefinitely, and raise_for_status() makes unsuccessful HTTP status codes visible rather than treating the response as usable page content. Requests documents timeouts, status codes, response headers and its exception types in its API documentation.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Parse: turn HTML into a tree

BeautifulSoup(response.text, "html.parser") selects Python’s built-in HTML parser explicitly. For HTML that is malformed, parser choice can affect the tree Beautiful Soup constructs, so use a deliberate backend rather than relying on an unspecified default. Beautiful Soup supports other backends too; consult its documentation if you need a different one.

Extract and normalize: create consistent records

The code uses soup.title for the document title and soup.select("a[href]") for anchors with an href. It converts relative links to absolute URLs with urljoin, and trims whitespace from visible link text. For a real task, select elements that correspond to your actual fields, such as a product name or article date, and give every record the same keys.

Rank #3
Sale
Anker USB C Hub, 7in1 Multi-Port USB Adapter, 4K@60Hz USBC to HDMI Splitter
  • Sleek 7-in-1 USB-C Hub: Features an HDMI port, two USB-A 3.0 ports, and a USB-C data port, each providing 5Gbps transfer speeds. It also includes a USB-C PD input port for charging up to 100W and dual SD and TF card slots, all in a compact design.
  • Flawless 4K@60Hz Video with HDMI: Delivers exceptional clarity and smoothness with its 4K@60Hz HDMI port, making it ideal for high-definition presentations and entertainment. (Note: Only the HDMI port supports video projection; the USB-C port is for data transfer only.)
  • Double Up on Efficiency: The two USB-A 3.0 ports and a USB-C port support a fast 5Gbps data rate, significantly boosting your transfer speeds and improving productivity.
  • Fast and Reliable 85W Charging: Offers high-capacity, speedy charging for laptops up to 85W, so you spend less time tethered to an outlet and more time being productive.
  • What You Get: Anker USB-C Hub (7-in-1), welcome guide, 18-month warranty, and our friendly customer service.

Validate important fields before treating a record as complete. This example reports a missing title; you could also skip the record or mark it invalid, depending on the task. Do not silently store an empty or missing value as if it were valid data.

Follow pages without turning a script into an open-ended crawl

The sample follows only the declared next link, only on the starting host, and only until the page cap or a stopping condition. A visited set prevents loops. Sites use many pagination patterns, so inspect the page markup and choose a selector that reflects the target’s real structure; there is no universal pagination algorithm.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Save records in a usable format

JSON Lines writes one JSON object per line, which is convenient for processing records incrementally. The output file is UTF-8, and ensure_ascii=False preserves readable non-ASCII text. For a small task, CSV may be easier to open in a spreadsheet; pick a format that matches your downstream use and define how missing fields should be represented.

Rank #4
Sale
UGREEN USB to USB C Adapter Combo 4-Pack, 10Gbps USB C Converter Space Gray
  • Dual Converters, Infinite Potential:Includes 2× USB C male to USB A female adapters and 2× USB A male to USB C female adapters. Perfect for a wide range of uses—tablets with Bluetooth keyboards, expand USB ports on macbook, and more. Two different converters for all your daily needs
  • Next-Level 10Gbps & 3A Charging: No more slow 480Mbps, this usb to usb c adapter has a transfer speed of up to 10Gbps, allowing you to do more transferring in less time. This usb adapter fits both USB A and USB C charger, supporting up to 3A fast charging
  • Upgraded Exquisite Craftsmanship: With an aluminum alloy housing and metal connector, the usbc to usb adapter is extremely durable and sturdy. Rigorously tested to withstand more than 10,000 times of plugging and unplugging, ensuring long-lasting performance
  • Broad Compatible: The usb c to usb adapter widely supports all USB C/ USB A devices like laptops, tablets, cellphones, car chargers, and phone chargers. Such as compatible with MacBook Pro/Air 2023/2022, Thunderbolt 4/3 Devices,Apple MagSafe Watch 9/8/7/SE/Ultra, iPad Pro 2022/2021, Samsung Galaxy S23/S20/S10, and iPhone 17/16/15 Pro. Plug and play
  • Please Note: To reach 10Gbps speed, keep the cable under 3.3 ft. For USB A Male to USB C adapters, try flipping the USB C connector. USB C Male to USB A adapters support bidirectional 10Gbps transfer within 3.3 ft

Use the standard library or Requests and Beautiful Soup?

Approach Useful for What you get Trade-off
Python standard library A small task where minimizing dependencies matters urllib.request for opening URLs, urllib.parse for URL operations, and urllib.robotparser for robots.txt parsing You assemble the request and parsing workflow from lower-level pieces.
Requests + Beautiful Soup One page or a modest set of static HTML pages A higher-level HTTP client and a convenient HTML tree/search API You still write your own pagination, crawl boundaries, validation, output and operational handling.
Scrapy A larger, repeatable crawl that needs a broader crawling framework Request/response abstractions, project setup and deployment workflows More framework and project structure than a one-page script needs.

This is a scale-based recommendation, not a claim that one tool is best for every site. Compare expected URL count, repeat frequency, concurrency needs, control over requests, error/retry handling, output integration and maintenance effort. Scrapy’s official site describes its framework and deployment options; its request/response reference documents URL, status, headers, body and decoded text: Scrapy and Scrapy request and response. The site listed Scrapy 2.19.0 in September 2026; check the official documentation for current release details before pinning a version.

If the information you need is absent from the HTML fetched by Requests, changing selectors will not make it appear. First look for a documented API, structured data, or another permitted data source. The available evidence does not establish a browser-rendering tool as a universal fix for any particular target.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

When to move from a script to Scrapy

A hand-written loop is often enough for a handful of static pages. Consider Scrapy when you need a repeatable multi-page crawl and want its request/response model and project workflow rather than building all crawl management yourself. Scrapy is a framework, so assess the setup and maintenance cost as well as the features; a larger tool is not automatically a better fit for a small extraction task.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Best Value
Anker USB C Hub, 5-in-1 USBC to HDMI Splitter with 4K Display
  • 5-in-1 Connectivity: Equipped with a 4K HDMI port, a 5 Gbps USB-C data port, two 5 Gbps USB-A ports, and a USB C 100W PD-IN port. Note: The USB C 100W PD-IN port supports only charging and does not support data transfer devices such as headphones or speakers.
  • Powerful Pass-Through Charging: Supports up to 85W pass-through charging so you can power up your laptop while you use the hub. Note: Pass-through charging requires a charger (not included). Note: To achieve full power for iPad, we recommend using a 45W wall charger.
  • Transfer Files in Seconds: Move files to and from your laptop at speeds of up to 5 Gbps via the USB-C and USB-A data ports. Note: The USB C 5Gbps Data port does not support video output.
  • HD Display: Connect to the HDMI port to stream or mirror content to an external monitor in resolutions of up to 4K@30Hz. Note: The USB-C ports do not support video output.
  • What You Get: Anker 332 USB-C Hub (5-in-1), welcome guide, our worry-free 18-month warranty, and friendly customer service.

Before scaling up, make sure the data source and access are appropriate, define domain and page boundaries, validate records, and decide how the crawl will handle failures and persist output. Do not increase concurrency or volume simply because a framework makes it easy to do so.

Common problems and fixes

  • The request times out. The host may be slow or unreachable, or the chosen timeout may be too short for the task. Keep the timeout finite; check connectivity and the target’s availability, then adjust cautiously. Do not retry indefinitely.
  • The script reports an HTTP error. Inspect the status code and response context. A denial or other error is not a signal to bypass access controls; stop or use an authorized data source.
  • The title or fields are missing. The HTML may have changed, the selector may not match, or the field may not be present in the fetched response. Inspect the returned HTML and update selectors only to match actual markup; check for an official API or structured data if the content is absent.
  • Pagination repeats or leaves the intended scope. Retain the visited set and page cap, resolve relative URLs against the response URL, and check the hostname before following a link. Confirm that the next-link selector identifies the intended navigation element.
  • Output contains blank or inconsistent records. Validate required fields before writing, use consistent record keys, and decide explicitly whether incomplete records should be skipped, flagged or retained with missing values.
  • Parsed structure differs across runs or machines. Specify the parser backend, as in the example. Beautiful Soup documents that parser choice can change the tree for malformed HTML.
  • The useful content is not in the response HTML. A static HTML parser cannot extract content it did not receive. Check for an authorized API, structured data or other permitted interface rather than assuming a different CSS selector will solve it.

Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server, not a web-scraping or structured-data extraction replacement. It can be useful when the task is to capture a page visually instead of extracting fields. One GET request returns an image or PDF; the example below saves a WebP screenshot. The ScreenshotNeo API documentation describes the request options.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp

ScreenshotNeo accepts cookie or consent banners as a visitor and removes more than 60 known consent platforms, newsletter popups and chat widgets before capture; each of those steps can be turned off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads and cache hits are not billed, and responses report the page verdict and billing status in headers. Its MCP server offers take_screenshot, get_page_info and capture_pdf for AI agents and MCP clients. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 screenshots.

Sign up for ScreenshotNeo’s free plan: 1,000 screenshots a month, no card required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Frequently Asked Questions

Can I scrape a page if its robots.txt allows crawling?

Not on that fact alone. Robots.txt expresses crawler access preferences; it does not settle legal permission or replace checking the site’s terms and applicable restrictions.

What should I try if the page content is not in the HTML I fetched?

Look for a documented API, structured data or another permitted source. A selector only finds markup that is present in the response.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

More from the Wire

  1. Shenzhen desk3 min
    HONOR Expands Beyond Smartphones With Humanoid Robot RevealHONOR said it unveiled its first humanoid robot at MWC 2026 and named shopping assistance, workplace inspections, and supportive companionship as intended uses. Later Robotics D1 claims and a reported…
  2. Cupertino desk5 min
    Apple Unveils AirPods Max 2: The Upgrade That Should Have Happened Years AgoAirPods Max 2 adds H2-powered audio features and Apple claims up to 1.5× more effective ANC, but its design, Smart Case, and 20-hour battery rating are unchanged. Wired lossless audio…
  3. Cupertino desk4 min
    Apple’s OLED Touch MacBooks Are Coming—but the Dynamic Island Is the Real GambleApple has not announced an OLED touchscreen MacBook, but reports point to high-end models arriving in late 2026 or early 2027. The reported Mac Dynamic Island could be useful, but…
Recommended PC Tool
Recommended PC Tool
Outdated Drivers Are Slowing You DownFree scan - exact matches
PC Slower Than It Used to Be?Free scan - under a minute

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.