Do these 3 things before closing this tab:
1Clear out junk files and repair common Windows errors2Fix the driver behind crashes, sound loss and screen glitches3Repair Windows errors before they cause bigger problemsUse functions to divide a scraper into clear steps: retrieve a page, parse its HTML, clean the extracted values, and save the results. Keeping those jobs separate makes the code easier to read, test, and adapt when a page changes. This guide assumes you know basic programming; the official Python 3.14.7 tutorial is aimed at people new to Python, rather than people new to programming.
What functions do in a scraper
A function gives a named responsibility to a piece of code. Instead of writing one long script that requests a page, searches its markup, fixes values, and writes a file, define a function for each stage and pass data from one stage to the next.
A useful starting pipeline is:
fetch_page(url)retrieves a page and returns its text.parse_items(html)extracts fields from the HTML and returns structured records.clean_item(item)normalizes or validates a record.save_items(items, path)writes the results to a destination.
This is a design choice, not a mandatory architecture. A tiny one-off script may need fewer functions; a larger scraper may split responsibilities further, such as retry policy, logging, or validation.
Choose retrieval and parsing tools
HTTP retrieval and HTML parsing solve different problems. Python’s urllib.request opens URLs and returns response content; the standard-library urllib package also includes URL parsing and error-handling modules. Python’s urllib.request documentation describes the request interface.
Free tools Windows power users keep installed
One-click scans. No signup required.
#1 Best Overall
Requests is a third-party HTTP client with a higher-level API. Its documentation describes sessions, automatic response decoding, connection pooling, and timeout support; the project documentation surfaced as release 2.34.2 and states support for Python 3.10 and newer. Check the current release and your installed version before relying on version-specific behavior.
For parsing, Python includes basic HTML parsing facilities, while Beautiful Soup provides a dedicated way to parse HTML and XML and navigate or search the resulting document tree. Its documentation surfaced as version 4.15.0, but version references there are not fully consistent; check the version you install rather than assuming a compatibility range.
| Task | Standard-library option | Third-party option |
|---|---|---|
| Retrieve a page | urllib.request; no extra HTTP package to install. |
Requests; a higher-level client with documented sessions, decoding, connection pooling, and timeouts. |
| Parse markup | Built-in HTML parsing facilities for basic parsing tasks. | Beautiful Soup; an HTML/XML tree-navigation and search interface. |
These are interface and dependency trade-offs, not a speed ranking. You can pair either retrieval approach with a parser; selecting one does not require selecting the other.
A small function-based scraper
The example below fetches a page with Requests, parses article-like cards with Beautiful Soup, normalizes their title and link, and saves JSON. It is illustrative: real websites use different markup, and you must replace the example selectors with ones that match a page you are permitted to access. Install the dependencies in your environment with python -m pip install requests beautifulsoup4.
Rank #2
import json
from pathlib import Path
from urllib.parse import urljoin
import requests
from bs4 import BeautifulSoup
def fetch_page(url):
"""Retrieve a page and return its decoded HTML text."""
response = requests.get(url, timeout=20)
response.raise_for_status()
return response.text
def parse_items(html, base_url):
"""Extract records from article cards in the supplied HTML."""
soup = BeautifulSoup(html, "html.parser")
items = []
for card in soup.select("article"):
heading = card.select_one("h2")
link = card.select_one("a[href]")
if heading is None or link is None:
continue
items.append({
"title": heading.get_text(" ", strip=True),
"url": urljoin(base_url, link["href"]),
})
return items
def clean_item(item):
"""Trim fields and reject incomplete records."""
title = " ".join(item.get("title", "").split())
url = item.get("url", "").strip()
if not title or not url:
return None
return {"title": title, "url": url}
def save_items(items, path):
"""Write records as UTF-8 JSON."""
Path(path).write_text(
json.dumps(items, ensure_ascii=False, indent=2),
encoding="utf-8",
)
def scrape(url, output_path):
html = fetch_page(url)
raw_items = parse_items(html, url)
items = [cleaned for item in raw_items if (cleaned := clean_item(item))]
save_items(items, output_path)
return items
if __name__ == "__main__":
page_url = "https://example.com/"
records = scrape(page_url, "items.json")
print(f"Saved {len(records)} records")
The sample uses Python’s built-in html.parser backend through Beautiful Soup. The selector article and the h2 heading are assumptions for demonstration, not selectors that work on every site. If a record has no heading or link, the parser skips it; the cleaner rejects empty values. The URL is resolved against the page address so relative links become absolute.
Why the boundaries help
fetch_pageowns network behavior and HTTP status handling. A timeout or non-success status becomes visible at the retrieval stage.parse_itemsturns markup into ordinary Python dictionaries; it does not make requests or write files.clean_itemapplies data-quality rules in one place, so those rules can be changed without rewriting the parser.save_itemscontrols the output format and destination independently of where records came from.
Use the standard library instead
If avoiding an additional HTTP dependency matters, the retrieval function can use urllib.request. For example, replace the Requests import and fetch_page with:
from urllib.request import Request, urlopen
def fetch_page(url):
request = Request(url, headers={"User-Agent": "ExampleScraper/1.0"})
with urlopen(request, timeout=20) as response:
return response.read().decode("utf-8")
This minimal variant decodes as UTF-8; pages with a different declared encoding may need encoding-aware handling. For either client, use a timeout so a stalled network request does not wait indefinitely, and handle errors at the boundary where they arise.
Make the functions safer and easier to change
Keep inputs and outputs explicit
Pass a URL into the retrieval function and return the page text, rather than hiding a particular URL in a global variable. Pass HTML into the parser and return records. Explicit inputs and outputs let you reason about each step without running the whole scraper.
Validate the structure you depend on
Selectors are assumptions about a page, not guarantees. Skip incomplete cards when that is appropriate, or raise a clear exception if a required field is missing. If a site redesign changes its markup, the parsing function is the natural place to update the selector.
Separate failures from empty results
An empty list can mean the page genuinely had no matching records, or that your selector no longer matches the markup. For important jobs, log the URL and record count, and consider treating an unexpectedly empty result as a warning or error. Do not silently save a successful-looking empty file after a failed request.
Handle exceptions at a useful level
Requests raises exceptions for network problems and raise_for_status() turns HTTP error responses into exceptions. Catch exceptions around a scrape or batch when you can decide whether to report, retry, or skip that page. Avoid catching every exception inside each small helper and returning an empty value: that can hide the difference between a missing field and a failed network call.
Check crawler guidance before automating requests
Before sending automated traffic, inspect the target site’s terms and crawler guidance, use conservative request volume, and plan for errors. Python’s urllib.robotparser can answer whether a user agent may fetch a URL under the rules it parses from robots.txt; its documented helpers include crawl_delay and request_rate. The linked Python documentation is for prerelease Python 3.16.0a0, so check details against the stable Python installation you use: urllib.robotparser documentation.
The Robots Exclusion Protocol specifies how crawlers interpret rules such as allow and disallow. RFC 9309 states: “These rules are not a form of access authorization.” See the IETF RFC 9309. A robots.txt check is crawler guidance, not a legal determination, a security barrier, or permission to access data. Whether a particular scraping activity is permitted depends on the site, jurisdiction, data, terms, and access method.
Common errors and fixes
| Symptom | Likely cause | What to check |
|---|---|---|
| Connection or timeout exception | The host is unavailable, the network is slow, or the request is being blocked. | Confirm the URL and connectivity, use a finite timeout, reduce request frequency, and decide whether a limited retry is appropriate. |
HTTP error after raise_for_status() |
The server returned an unsuccessful status, such as a missing page or denied request. | Check the URL and response status; do not treat the response as valid page data without understanding the status. |
| Parser returns no records | The example selectors do not match the page, the content is different, or the page is rendered client-side. | Inspect the response HTML and update selectors. A parser cannot extract markup that was never present in the retrieved response. |
| Broken or relative links | The page uses relative URLs. | Resolve links against the page URL with urljoin, as in the example. |
| Unreadable characters | The response bytes were decoded with the wrong character encoding. | Use the HTTP client’s encoding handling or inspect the page’s declared encoding; do not assume every page is UTF-8 when manually decoding bytes. |
| Package import error | A dependency is missing from the Python environment running the script. | Install packages into that environment and verify the interpreter used to run the script matches the one used for installation. |
Performance, reliability, and cost considerations
Functions themselves do not make HTTP requests faster. They make it easier to change and inspect the places where network work, parsing, and output happen. Requests documents connection pooling and sessions; when making multiple requests to the same host, consult its session guidance rather than building assumptions about performance into the scraper. The available documentation does not establish a universal speed advantage for one parser or retrieval library.
Reliability depends on the target site’s availability and markup, your request policy, and how the script responds to timeouts, status errors, and incomplete records. A finite timeout, explicit status handling, conservative volume, and visible logging make failures easier to diagnose. Scraping cost depends on your infrastructure and workload; the sources here do not establish a general cost figure.
Or skip the browser setup
If your goal is a screenshot rather than structured data extracted from HTML, ScreenshotNeo is a website screenshot API and MCP server for developers. One GET request can return a PNG, JPEG, WebP, or PDF. Its capture options include full-page screenshots with lazy images loaded, CSS-selector element capture, viewport and device presets, dark mode, custom CSS and JavaScript, waits, and PDF settings. See the ScreenshotNeo API documentation for parameters and setup.
Best Value
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
open("shot.webp", "wb").write(r.content)
ScreenshotNeo accepts cookie and consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each of those steps can be turned off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers report the page verdict and billing status. It also offers an MCP server for AI agents, including Claude, Cursor, and other MCP clients. The free plan includes 1,000 screenshots a month without a card; paid plans start at $5 for 3,000 screenshots.
Sign up for ScreenshotNeo’s free plan to try 1,000 screenshots a month with no card.
Frequently Asked Questions
Can functions scrape a page that renders its content with JavaScript?
Not if the HTML returned by the HTTP request does not contain the content you need; the parser only sees the response it receives. The screenshot workflow above is for capturing a rendered visual page, not for extracting structured records.
Do I need Requests and Beautiful Soup together?
No. They handle separate jobs: Requests retrieves pages, while Beautiful Soup parses markup. You can choose standard-library retrieval and still use a parser, or use Requests with a different parsing approach.
Quick wins for a faster PC:
Clear out junk files and repair common Windows errorsFree Scan →Scan for outdated or missing drivers - takes under a minuteDriver Scan →Repair Windows errors before they cause bigger problemsFix Now →Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

