Recommended Free Tools
You can fetch a static public page or file and pull out the fields you need using only Python’s standard library: urllib.request to retrieve the response, urllib.robotparser to check crawl rules, and html.parser, json, or csv to read the content. No pip install is required. This guide walks through that workflow in the order you should run it, with the checks that keep it from failing silently.
What “static” and “zero dependencies” cover
“Static” here means the server returns the data in the initial HTTP response, without a browser running JavaScript to assemble it. “Zero dependencies” means every module in the example ships with Python. It does not mean the approach works on every website, that you need no setup at all (you still need a Python 3 interpreter), or that it can read pages that render their content in the browser. If a value appears only after scripts run, the standard-library method will not see it, and you will need a different tool.
Step 1: Check robots.txt and identify the response format
Before you send a request, answer two questions: what does the site’s robots.txt allow, and what kind of content will the URL return? A single URL can return HTML, plain text, JSON, CSV, or binary data, so do not assume markup. Open the URL in a browser’s developer tools or with a single manual request and note the Content-Type header. Then check the crawl rules programmatically:
from urllib.robotparser import RobotFileParser
USER_AGENT = "StaticDataLearner/1.0"
TARGET = "https://example.org/data/page.html"
rp = RobotFileParser()
rp.set_url("https://example.org/robots.txt")
rp.read()
if not rp.can_fetch(USER_AGENT, TARGET):
raise SystemExit("robots.txt disallows this URL for this user agent")
Use the same user-agent string in the robots check and in the request header, so the check describes the client you actually run. The parser answers only one question: whether the robots.txt rules permit that user agent to fetch that URL. It says nothing about the site’s terms of service, access controls, or privacy obligations. How the parser handles a robots.txt file that is missing or returns an error depends on your Python release, so read that release’s urllib.robotparser documentation before relying on the result.
#1 Best Overall
Step 2: Fetch the response and handle failures
Use urllib.request.Request so you can set headers, and open it with a context manager so the connection closes even when parsing fails. Catch HTTPError before URLError, because HTTPError is a subclass of it and would otherwise be swallowed by the broader handler.
import urllib.request
import urllib.error
def fetch(url, user_agent=USER_AGENT, timeout=10):
req = urllib.request.Request(url, headers={"User-Agent": user_agent})
try:
with urllib.request.urlopen(req, timeout=timeout) as resp:
body = resp.read()
return {
"status": resp.status,
"final_url": resp.geturl(),
"content_type": resp.headers.get_content_type(),
"charset": resp.headers.get_content_charset(),
"body": body,
}
except urllib.error.HTTPError as e:
raise RuntimeError(f"HTTP {e.code} for {url}") from e
except urllib.error.URLError as e:
raise RuntimeError(f"Could not reach {url}: {e.reason}") from e
Three details matter here. First, timeout limits individual blocking socket operations; it is not a ceiling on the total time a download takes, and the standard library does not guarantee that a connection will be quick. Second, a plain GET request follows redirects by default, so compare final_url with what you asked for. Third, read() returns bytes, not text. Store the raw body first and decode it in a separate step, so that a decoding problem never hides a network problem.
Step 3: Decode the bytes deliberately
urlopen cannot know how the byte stream is encoded on its own. The server may declare a charset in the Content-Type header, the document may declare one itself, or neither may be present. The example above reads the header value with get_content_charset(), which returns None when the header has no charset parameter. Use this order:
Rank #2
- If the header declares a charset, decode with that charset and
errors="strict". - If the header has no charset and the format is HTML, look for a charset declaration in the document itself, such as a meta tag near the top of the page, and decode with that value.
- If neither is present, choose a fallback deliberately, such as UTF-8 for JSON, and verify the result against the page in a browser.
Do not treat a fixed UTF-8 decode as guaranteed for every site. Strict decoding is useful because a wrong charset then raises an error you can see, instead of producing garbled text that passes your checks.
Windows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallOutdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchStep 4: Parse according to the resource
Choose the parser from the response format, not from habit. The table below compares the three formats this approach handles.
| Format | Typical Content-Type | How the data is embedded | Standard-library tool | Main fragility |
|---|---|---|---|---|
| HTML | text/html | Values mixed with markup and layout, so you locate them by tag and attribute | html.parser.HTMLParser |
Markup changes break selectors; the parser does not build a full document tree |
| JSON | application/json | Structured keys and nested objects or arrays | json.loads |
Field names and nesting can change without notice |
| CSV | text/csv | Rows of delimited values with an optional header row | csv.DictReader |
Column order, quoting, and delimiters vary between exporters |
The Content-Type values in the table are the usual ones, not guarantees. Confirm them against the actual response header for the URL you are fetching.
HTML with html.parser
HTMLParser is a callback parser. You subclass it and override handlers such as handle_starttag, handle_endtag, and handle_data, and the parser calls them as it reads the document. It accepts invalid markup without raising an exception, but it does not check that end tags match start tags, and it does not call end-tag handlers for every implicitly closed element. Write your handlers to tolerate that. The example below collects the text of every second-level heading. Note that handle_data can be called several times for one text node, so the code appends to the current value rather than replacing it.
from html.parser import HTMLParser
class H2Collector(HTMLParser):
def __init__(self):
super().__init__()
self.in_h2 = False
self.headings = []
def handle_starttag(self, tag, attrs):
if tag == "h2":
self.in_h2 = True
self.headings.append("")
def handle_endtag(self, tag):
if tag == "h2":
self.in_h2 = False
def handle_data(self, data):
if self.in_h2:
self.headings[-1] += data
parser = H2Collector()
parser.feed(html_text)
parser.close()
headings = [h.strip() for h in parser.headings if h.strip()]
This parser is not a browser DOM and does not run JavaScript. The more deeply nested and irregular the markup, the more state your handlers must track. For a page with a stable, simple structure, callbacks are compact. For a page where you need to reason about ancestors and siblings, the same code grows quickly, and you should consider whether the page offers a cleaner data source such as a linked file or an API.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
JSON with json.loads
json.loads accepts either a string or bytes, so you can pass the body directly. Use the charset from Step 3 when you decode it yourself, and check that the structure is what you expect before using it.
import json
data = json.loads(body)
items = data.get("items")
if not isinstance(items, list):
raise ValueError("Expected a list under 'items'; the response structure may have changed")
names = [item["name"] for item in items if "name" in item]
CSV with csv.DictReader
For delimited data, decode the body to text and wrap it in io.StringIO created with newline="", which is the form the csv documentation recommends for reading. DictReader uses the first row as field names.
import csv
import io
text = body.decode(charset or "utf-8")
reader = csv.DictReader(io.StringIO(text, newline=""))
rows = list(reader)
if not rows or "date" not in reader.fieldnames:
raise ValueError("CSV is empty or missing the 'date' column")
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Step 5: Validate the result, then save it
Treat an empty result as a failure, not a success. A changed class name in HTML or a renamed key in JSON will often produce an empty list rather than an exception, and an unchecked empty list is the most common silent error in extraction scripts. Compare the count of extracted items with what you expect, and fail loudly when it falls outside a reasonable range.
When saving, write text with an explicit encoding and keep the raw response if you may need to re-parse it later:
Do these 3 things before closing this tab:
1Scan for outdated or missing drivers - takes under a minute2Repair Windows errors before they cause bigger problems3Fix the driver behind crashes, sound loss and screen glitchesBest Value
import json
with open("headings.json", "w", encoding="utf-8") as f:
json.dump(headings, f, ensure_ascii=False, indent=2)
with open("raw_response.bin", "wb") as f:
f.write(body)
When the static approach stops working
Most failures fall into a small set of patterns. Check these in order before changing your parser.
- The status is 200 but the data is missing. The values are probably injected by JavaScript after the page loads. Open the page with JavaScript disabled, or search the raw response for a known value. If it is absent, the static method cannot see it.
- HTTP 403, 429, or repeated 5xx responses. The server may be refusing automated clients or limiting request rates. Stop, read the site’s published access rules, and reduce your request frequency. Do not switch headers to get around a refusal.
- Garbled accented characters or symbols. The charset is wrong. Recheck the header and the document’s own declaration, as described in Step 3.
- The final URL differs from the requested one. The request was redirected. Check whether the new location is still within what robots.txt permits, and run the robots check against the final URL.
- Timeouts or long pauses. Connection setup can take arbitrarily long. Lower the expectations of a single request, add a retry limit with a pause between attempts, and log the failure rather than retrying without bound.
- Extraction returns zero items. The markup or keys have changed. Save the raw response, inspect it, and update the selectors or key names, then rerun the validation in Step 5.
Compliance and responsibility
A working script is not the same as permission. Public visibility does not settle whether a collection complies with a site’s terms, its access controls, the privacy expectations of the people described in the data, or the law in your jurisdiction. Check those questions separately from the technical steps above, and limit collection to the fields and frequency you actually need.
Quick Recap
“
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




