A small Python web scraper has four jobs: fetch a page, parse its HTML, extract and validate the fields you need, then save them in a structured format. For a permitted static page or a modest set of pages, requests and Beautiful Soup are a clear starting point. This guide builds that workflow, adds bounded pagination and error handling, and explains when Scrapy is a better fit.
Before you scrape: define the task and check access
Choose a site and data you are permitted to access. Prefer an official API or downloadable dataset if one meets the need. Before writing code, decide which fields to collect, which pages are in scope, how many pages to visit, and what output format you want. A defined boundary helps prevent an accidental site-wide crawl.
- Inspect the site’s published terms and its
robots.txtfile. Python’s standard library includesurllib.robotparserfor parsing robots.txt; the Python 3.14.8 documentation describes it alongsideurllib.requestandurllib.parse: Python urllib documentation. - Robots.txt communicates crawler access preferences, but it is not a complete statement of legal permission. Google explains that the file tells search engine crawlers which URLs they may access and is used mainly to manage crawler traffic; it also says robots.txt is not a mechanism for keeping a URL out of search results: Google Search Central: Robots.txt Introduction and Guide.
- Use conservative request volume, honor applicable terms and access controls, and stop if access is denied. There is no universal request rate or legal rule that applies to every site and jurisdiction.
Keep credentials out of source code. Do not attempt to evade a site’s blocks or access controls.
Install the Python packages
Use a virtual environment for a small project, then install Requests and Beautiful Soup:
The Tool Desk
Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →#1 Best Overall
- Read Before You Buy — No Video Output: These adapters support charging and USB 2.0 data transfer, but cannot transmit video signals. Except for standard USB webcams (which use USB data only), they are not compatible with HDMI/DisplayPort cables, video-capable USB-C hubs, or docking stations with video output.
- Convert USB-A Ports to USB-C: Designed to connect USB-C earphones, cables, flash drives, card readers, and other USB-C accessories to standard USB-A ports. Plug-and-play with no drivers or software required.
- Aluminum Alloy Housing: Built with a sturdy aluminum alloy shell that aids in heat dissipation and protects against daily wear and scratches. Designed to maintain a stable and secure connection.
- Compact & Travel-Friendly: The ultra-compact design allows the adapter to stay plugged into your device without blocking adjacent ports or adding bulk, reducing wear and tear on your original USB ports.
- 12-Month Warranty: Backed by a 12-month manufacturer warranty for peace of mind. Designed to meet strict quality control standards for reliable everyday performance.
python -m venv .venv
# macOS or Linux:
source .venv/bin/activate
# Windows PowerShell:
.venvScriptsActivate.ps1
python -m pip install requests beautifulsoup4
Requests provides HTTP request/response handling, including timeouts, status codes, headers, sessions and exceptions. Beautiful Soup parses HTML into a searchable tree. Its documentation describes multiple parser backends; specifying one makes your parsing choice explicit, and different parsers can produce different trees for malformed markup: Requests documentation and Beautiful Soup documentation.
Build a bounded scraper for a page and its next links
The example below fetches a starting page and follows a page’s <a rel="next"> link, if present, without leaving the starting host or exceeding a page limit. It extracts each page title and its links into JSON Lines. The selector for the next link is a convention, not a guarantee: inspect the target site’s HTML and adjust it to the actual markup. The code is an illustrative starting pattern, not a claim that it has been run against a live target.
import json
import time
from urllib.parse import urljoin, urlparse
import requests
from bs4 import BeautifulSoup
START_URL = "https://example.com/"
MAX_PAGES = 5
DELAY_SECONDS = 1
TIMEOUT_SECONDS = 10
OUTPUT_FILE = "pages.jsonl"
def same_host(url, host):
return urlparse(url).netloc == host
def scrape(start_url):
host = urlparse(start_url).netloc
session = requests.Session()
visited = set()
current_url = start_url
records = []
for _ in range(MAX_PAGES):
if not current_url or current_url in visited:
break
if not same_host(current_url, host):
print(f"Stopping: URL is outside the starting host: {current_url}")
break
visited.add(current_url)
try:
response = session.get(current_url, timeout=TIMEOUT_SECONDS)
response.raise_for_status()
except requests.exceptions.Timeout:
print(f"Timed out: {current_url}")
break
except requests.exceptions.HTTPError as exc:
status = exc.response.status_code if exc.response is not None else "unknown"
print(f"HTTP error {status}: {current_url}")
break
except requests.exceptions.RequestException as exc:
print(f"Request failed for {current_url}: {exc}")
break
soup = BeautifulSoup(response.text, "html.parser")
title = soup.title.get_text(" ", strip=True) if soup.title else None
links = []
for anchor in soup.select("a[href]"):
absolute_url = urljoin(response.url, anchor["href"])
if urlparse(absolute_url).scheme in {"http", "https"}:
links.append({
"text": anchor.get_text(" ", strip=True),
"url": absolute_url,
})
if not title:
print(f"Expected title is missing: {response.url}")
records.append({"url": response.url, "title": title, "links": links})
next_anchor = soup.select_one('a[rel~="next"][href]')
next_url = urljoin(response.url, next_anchor["href"]) if next_anchor else None
if next_url and (not same_host(next_url, host) or next_url in visited):
next_url = None
current_url = next_url
if current_url:
time.sleep(DELAY_SECONDS)
return records
if __name__ == "__main__":
pages = scrape(START_URL)
with open(OUTPUT_FILE, "w", encoding="utf-8") as output:
for page in pages:
output.write(json.dumps(page, ensure_ascii=False) + "n")
print(f"Saved {len(pages)} page record(s) to {OUTPUT_FILE}")
Change START_URL, the extraction selectors, and the next-page selector to match the site you are allowed to access. The example uses a finite timeout, checks HTTP status with raise_for_status(), keeps a visited set, restricts navigation to the starting host, caps pages, and pauses between requests. The page limit is a safety boundary for this sample, not a recommended universal crawl size or rate.
Rank #2
- 5-in-1 USB-C Hub: Experience comprehensive connectivity featuring a Power Delivery input, two USB-A 2.0 ports, a USB-A 3.0 port, and an HDMI port. (Note: The USB-C power delivery input port is only for connecting an external wall charger to power your laptop and cannot power peripheral devices.)
- 90W Pass-Through Charging: Achieve optimal charging with 90W pass-through power to your laptop, supported by a total input of 100W, with the hub reserving 10W for operational efficiency. (Note: Wall charger not included.)
- Quick Data Transfers: Accelerate your productivity with rapid data transfers using a high-speed 5Gbps USB 3.0 port and two 480Mbps USB 2.0 ports.
- 4K HDMI Display: Enhance your visual experience with a hub capable of delivering 4K resolution at 30Hz in both mirror and extend modes. Please note that this hub is compatible with MacBook (macOS 12 and newer), Windows 10 and 11, ChromeOS, and laptops equipped with DP Alt Mode and Power Delivery. Note: This device is not compatible with Linux.
- What You Get: Anker USB-C Hub (5-in-1, 4K HDMI), welcome guide, 18-month warranty, and our friendly customer service.
How the scraper works
Fetch: make the HTTP request
requests.Session() lets the script reuse a session across requests. Each call uses a finite timeout, since a network request can otherwise wait indefinitely, and raise_for_status() makes unsuccessful HTTP status codes visible rather than treating the response as usable page content. Requests documents timeouts, status codes, response headers and its exception types in its API documentation.
Parse: turn HTML into a tree
BeautifulSoup(response.text, "html.parser") selects Python’s built-in HTML parser explicitly. For HTML that is malformed, parser choice can affect the tree Beautiful Soup constructs, so use a deliberate backend rather than relying on an unspecified default. Beautiful Soup supports other backends too; consult its documentation if you need a different one.
Extract and normalize: create consistent records
The code uses soup.title for the document title and soup.select("a[href]") for anchors with an href. It converts relative links to absolute URLs with urljoin, and trims whitespace from visible link text. For a real task, select elements that correspond to your actual fields, such as a product name or article date, and give every record the same keys.
Rank #3
- Sleek 7-in-1 USB-C Hub: Features an HDMI port, two USB-A 3.0 ports, and a USB-C data port, each providing 5Gbps transfer speeds. It also includes a USB-C PD input port for charging up to 100W and dual SD and TF card slots, all in a compact design.
- Flawless 4K@60Hz Video with HDMI: Delivers exceptional clarity and smoothness with its 4K@60Hz HDMI port, making it ideal for high-definition presentations and entertainment. (Note: Only the HDMI port supports video projection; the USB-C port is for data transfer only.)
- Double Up on Efficiency: The two USB-A 3.0 ports and a USB-C port support a fast 5Gbps data rate, significantly boosting your transfer speeds and improving productivity.
- Fast and Reliable 85W Charging: Offers high-capacity, speedy charging for laptops up to 85W, so you spend less time tethered to an outlet and more time being productive.
- What You Get: Anker USB-C Hub (7-in-1), welcome guide, 18-month warranty, and our friendly customer service.
Validate important fields before treating a record as complete. This example reports a missing title; you could also skip the record or mark it invalid, depending on the task. Do not silently store an empty or missing value as if it were valid data.
Follow pages without turning a script into an open-ended crawl
The sample follows only the declared next link, only on the starting host, and only until the page cap or a stopping condition. A visited set prevents loops. Sites use many pagination patterns, so inspect the page markup and choose a selector that reflects the target’s real structure; there is no universal pagination algorithm.
Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minutePC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Save records in a usable format
JSON Lines writes one JSON object per line, which is convenient for processing records incrementally. The output file is UTF-8, and ensure_ascii=False preserves readable non-ASCII text. For a small task, CSV may be easier to open in a spreadsheet; pick a format that matches your downstream use and define how missing fields should be represented.
Rank #4
- Dual Converters, Infinite Potential:Includes 2× USB C male to USB A female adapters and 2× USB A male to USB C female adapters. Perfect for a wide range of uses—tablets with Bluetooth keyboards, expand USB ports on macbook, and more. Two different converters for all your daily needs
- Next-Level 10Gbps & 3A Charging: No more slow 480Mbps, this usb to usb c adapter has a transfer speed of up to 10Gbps, allowing you to do more transferring in less time. This usb adapter fits both USB A and USB C charger, supporting up to 3A fast charging
- Upgraded Exquisite Craftsmanship: With an aluminum alloy housing and metal connector, the usbc to usb adapter is extremely durable and sturdy. Rigorously tested to withstand more than 10,000 times of plugging and unplugging, ensuring long-lasting performance
- Broad Compatible: The usb c to usb adapter widely supports all USB C/ USB A devices like laptops, tablets, cellphones, car chargers, and phone chargers. Such as compatible with MacBook Pro/Air 2023/2022, Thunderbolt 4/3 Devices,Apple MagSafe Watch 9/8/7/SE/Ultra, iPad Pro 2022/2021, Samsung Galaxy S23/S20/S10, and iPhone 17/16/15 Pro. Plug and play
- Please Note: To reach 10Gbps speed, keep the cable under 3.3 ft. For USB A Male to USB C adapters, try flipping the USB C connector. USB C Male to USB A adapters support bidirectional 10Gbps transfer within 3.3 ft
Use the standard library or Requests and Beautiful Soup?
| Approach | Useful for | What you get | Trade-off |
|---|---|---|---|
| Python standard library | A small task where minimizing dependencies matters | urllib.request for opening URLs, urllib.parse for URL operations, and urllib.robotparser for robots.txt parsing |
You assemble the request and parsing workflow from lower-level pieces. |
| Requests + Beautiful Soup | One page or a modest set of static HTML pages | A higher-level HTTP client and a convenient HTML tree/search API | You still write your own pagination, crawl boundaries, validation, output and operational handling. |
| Scrapy | A larger, repeatable crawl that needs a broader crawling framework | Request/response abstractions, project setup and deployment workflows | More framework and project structure than a one-page script needs. |
This is a scale-based recommendation, not a claim that one tool is best for every site. Compare expected URL count, repeat frequency, concurrency needs, control over requests, error/retry handling, output integration and maintenance effort. Scrapy’s official site describes its framework and deployment options; its request/response reference documents URL, status, headers, body and decoded text: Scrapy and Scrapy request and response. The site listed Scrapy 2.19.0 in September 2026; check the official documentation for current release details before pinning a version.
If the information you need is absent from the HTML fetched by Requests, changing selectors will not make it appear. First look for a documented API, structured data, or another permitted data source. The available evidence does not establish a browser-rendering tool as a universal fix for any particular target.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.When to move from a script to Scrapy
A hand-written loop is often enough for a handful of static pages. Consider Scrapy when you need a repeatable multi-page crawl and want its request/response model and project workflow rather than building all crawl management yourself. Scrapy is a framework, so assess the setup and maintenance cost as well as the features; a larger tool is not automatically a better fit for a small extraction task.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Best Value
- 5-in-1 Connectivity: Equipped with a 4K HDMI port, a 5 Gbps USB-C data port, two 5 Gbps USB-A ports, and a USB C 100W PD-IN port. Note: The USB C 100W PD-IN port supports only charging and does not support data transfer devices such as headphones or speakers.
- Powerful Pass-Through Charging: Supports up to 85W pass-through charging so you can power up your laptop while you use the hub. Note: Pass-through charging requires a charger (not included). Note: To achieve full power for iPad, we recommend using a 45W wall charger.
- Transfer Files in Seconds: Move files to and from your laptop at speeds of up to 5 Gbps via the USB-C and USB-A data ports. Note: The USB C 5Gbps Data port does not support video output.
- HD Display: Connect to the HDMI port to stream or mirror content to an external monitor in resolutions of up to 4K@30Hz. Note: The USB-C ports do not support video output.
- What You Get: Anker 332 USB-C Hub (5-in-1), welcome guide, our worry-free 18-month warranty, and friendly customer service.
Before scaling up, make sure the data source and access are appropriate, define domain and page boundaries, validate records, and decide how the crawl will handle failures and persist output. Do not increase concurrency or volume simply because a framework makes it easy to do so.
Common problems and fixes
- The request times out. The host may be slow or unreachable, or the chosen timeout may be too short for the task. Keep the timeout finite; check connectivity and the target’s availability, then adjust cautiously. Do not retry indefinitely.
- The script reports an HTTP error. Inspect the status code and response context. A denial or other error is not a signal to bypass access controls; stop or use an authorized data source.
- The title or fields are missing. The HTML may have changed, the selector may not match, or the field may not be present in the fetched response. Inspect the returned HTML and update selectors only to match actual markup; check for an official API or structured data if the content is absent.
- Pagination repeats or leaves the intended scope. Retain the visited set and page cap, resolve relative URLs against the response URL, and check the hostname before following a link. Confirm that the next-link selector identifies the intended navigation element.
- Output contains blank or inconsistent records. Validate required fields before writing, use consistent record keys, and decide explicitly whether incomplete records should be skipped, flagged or retained with missing values.
- Parsed structure differs across runs or machines. Specify the parser backend, as in the example. Beautiful Soup documents that parser choice can change the tree for malformed HTML.
- The useful content is not in the response HTML. A static HTML parser cannot extract content it did not receive. Check for an authorized API, structured data or other permitted interface rather than assuming a different CSS selector will solve it.
Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server, not a web-scraping or structured-data extraction replacement. It can be useful when the task is to capture a page visually instead of extracting fields. One GET request returns an image or PDF; the example below saves a WebP screenshot. The ScreenshotNeo API documentation describes the request options.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp
ScreenshotNeo accepts cookie or consent banners as a visitor and removes more than 60 known consent platforms, newsletter popups and chat widgets before capture; each of those steps can be turned off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads and cache hits are not billed, and responses report the page verdict and billing status in headers. Its MCP server offers take_screenshot, get_page_info and capture_pdf for AI agents and MCP clients. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 screenshots.
Sign up for ScreenshotNeo’s free plan: 1,000 screenshots a month, no card required.
Quick wins for a faster PC:
Repair Windows errors before they cause bigger problemsFix Now →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Clear out junk files and repair common Windows errorsFree Scan →Frequently Asked Questions
Can I scrape a page if its robots.txt allows crawling?
Not on that fact alone. Robots.txt expresses crawler access preferences; it does not settle legal permission or replace checking the site’s terms and applicable restrictions.
What should I try if the page content is not in the HTML I fetched?
Look for a documented API, structured data or another permitted source. A selector only finds markup that is present in the response.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




