Hardware FixRecommendedDevice not working? Your driver may be the problemCheck updates for common hardware issues.Fix DriversOctober DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsSlow PC?RecommendedPC slow today? Run a repair scan before it gets worseResolve common Windows issues and optimize system performance.Scan Now×
Skip to content
World desk7 min

Extracting Static Public Data with Python (Zero Dependencies)

Fetch a static public page or file and extract its fields using only Python's standard library, with robots.txt checks, safe decoding, and parsers chosen by response format.
Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

You can fetch a static public page or file and pull out the fields you need using only Python’s standard library: urllib.request to retrieve the response, urllib.robotparser to check crawl rules, and html.parser, json, or csv to read the content. No pip install is required. This guide walks through that workflow in the order you should run it, with the checks that keep it from failing silently.

What “static” and “zero dependencies” cover

“Static” here means the server returns the data in the initial HTTP response, without a browser running JavaScript to assemble it. “Zero dependencies” means every module in the example ships with Python. It does not mean the approach works on every website, that you need no setup at all (you still need a Python 3 interpreter), or that it can read pages that render their content in the browser. If a value appears only after scripts run, the standard-library method will not see it, and you will need a different tool.

Step 1: Check robots.txt and identify the response format

Before you send a request, answer two questions: what does the site’s robots.txt allow, and what kind of content will the URL return? A single URL can return HTML, plain text, JSON, CSV, or binary data, so do not assume markup. Open the URL in a browser’s developer tools or with a single manual request and note the Content-Type header. Then check the crawl rules programmatically:

from urllib.robotparser import RobotFileParser

USER_AGENT = "StaticDataLearner/1.0"
TARGET = "https://example.org/data/page.html"

rp = RobotFileParser()
rp.set_url("https://example.org/robots.txt")
rp.read()
if not rp.can_fetch(USER_AGENT, TARGET):
    raise SystemExit("robots.txt disallows this URL for this user agent")

Use the same user-agent string in the robots check and in the request header, so the check describes the client you actually run. The parser answers only one question: whether the robots.txt rules permit that user agent to fetch that URL. It says nothing about the site’s terms of service, access controls, or privacy obligations. How the parser handles a robots.txt file that is missing or returns an error depends on your Python release, so read that release’s urllib.robotparser documentation before relying on the result.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Step 2: Fetch the response and handle failures

Use urllib.request.Request so you can set headers, and open it with a context manager so the connection closes even when parsing fails. Catch HTTPError before URLError, because HTTPError is a subclass of it and would otherwise be swallowed by the broader handler.

import urllib.request
import urllib.error

def fetch(url, user_agent=USER_AGENT, timeout=10):
    req = urllib.request.Request(url, headers={"User-Agent": user_agent})
    try:
        with urllib.request.urlopen(req, timeout=timeout) as resp:
            body = resp.read()
            return {
                "status": resp.status,
                "final_url": resp.geturl(),
                "content_type": resp.headers.get_content_type(),
                "charset": resp.headers.get_content_charset(),
                "body": body,
            }
    except urllib.error.HTTPError as e:
        raise RuntimeError(f"HTTP {e.code} for {url}") from e
    except urllib.error.URLError as e:
        raise RuntimeError(f"Could not reach {url}: {e.reason}") from e

Three details matter here. First, timeout limits individual blocking socket operations; it is not a ceiling on the total time a download takes, and the standard library does not guarantee that a connection will be quick. Second, a plain GET request follows redirects by default, so compare final_url with what you asked for. Third, read() returns bytes, not text. Store the raw body first and decode it in a separate step, so that a decoding problem never hides a network problem.

Step 3: Decode the bytes deliberately

urlopen cannot know how the byte stream is encoded on its own. The server may declare a charset in the Content-Type header, the document may declare one itself, or neither may be present. The example above reads the header value with get_content_charset(), which returns None when the header has no charset parameter. Use this order:

  1. If the header declares a charset, decode with that charset and errors="strict".
  2. If the header has no charset and the format is HTML, look for a charset declaration in the document itself, such as a meta tag near the top of the page, and decode with that value.
  3. If neither is present, choose a fallback deliberately, such as UTF-8 for JSON, and verify the result against the page in a browser.

Do not treat a fixed UTF-8 decode as guaranteed for every site. Strict decoding is useful because a wrong charset then raises an error you can see, instead of producing garbled text that passes your checks.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Step 4: Parse according to the resource

Choose the parser from the response format, not from habit. The table below compares the three formats this approach handles.

Format Typical Content-Type How the data is embedded Standard-library tool Main fragility
HTML text/html Values mixed with markup and layout, so you locate them by tag and attribute html.parser.HTMLParser Markup changes break selectors; the parser does not build a full document tree
JSON application/json Structured keys and nested objects or arrays json.loads Field names and nesting can change without notice
CSV text/csv Rows of delimited values with an optional header row csv.DictReader Column order, quoting, and delimiters vary between exporters

The Content-Type values in the table are the usual ones, not guarantees. Confirm them against the actual response header for the URL you are fetching.

HTML with html.parser

HTMLParser is a callback parser. You subclass it and override handlers such as handle_starttag, handle_endtag, and handle_data, and the parser calls them as it reads the document. It accepts invalid markup without raising an exception, but it does not check that end tags match start tags, and it does not call end-tag handlers for every implicitly closed element. Write your handlers to tolerate that. The example below collects the text of every second-level heading. Note that handle_data can be called several times for one text node, so the code appends to the current value rather than replacing it.

from html.parser import HTMLParser

class H2Collector(HTMLParser):
    def __init__(self):
        super().__init__()
        self.in_h2 = False
        self.headings = []

    def handle_starttag(self, tag, attrs):
        if tag == "h2":
            self.in_h2 = True
            self.headings.append("")

    def handle_endtag(self, tag):
        if tag == "h2":
            self.in_h2 = False

    def handle_data(self, data):
        if self.in_h2:
            self.headings[-1] += data

parser = H2Collector()
parser.feed(html_text)
parser.close()
headings = [h.strip() for h in parser.headings if h.strip()]

This parser is not a browser DOM and does not run JavaScript. The more deeply nested and irregular the markup, the more state your handlers must track. For a page with a stable, simple structure, callbacks are compact. For a page where you need to reason about ancestors and siblings, the same code grows quickly, and you should consider whether the page offers a cleaner data source such as a linked file or an API.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

JSON with json.loads

json.loads accepts either a string or bytes, so you can pass the body directly. Use the charset from Step 3 when you decode it yourself, and check that the structure is what you expect before using it.

import json

data = json.loads(body)
items = data.get("items")
if not isinstance(items, list):
    raise ValueError("Expected a list under 'items'; the response structure may have changed")
names = [item["name"] for item in items if "name" in item]

CSV with csv.DictReader

For delimited data, decode the body to text and wrap it in io.StringIO created with newline="", which is the form the csv documentation recommends for reading. DictReader uses the first row as field names.

import csv
import io

text = body.decode(charset or "utf-8")
reader = csv.DictReader(io.StringIO(text, newline=""))
rows = list(reader)
if not rows or "date" not in reader.fieldnames:
    raise ValueError("CSV is empty or missing the 'date' column")
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Step 5: Validate the result, then save it

Treat an empty result as a failure, not a success. A changed class name in HTML or a renamed key in JSON will often produce an empty list rather than an exception, and an unchecked empty list is the most common silent error in extraction scripts. Compare the count of extracted items with what you expect, and fail loudly when it falls outside a reasonable range.

When saving, write text with an explicit encoding and keep the raw response if you may need to re-parse it later:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
import json

with open("headings.json", "w", encoding="utf-8") as f:
    json.dump(headings, f, ensure_ascii=False, indent=2)

with open("raw_response.bin", "wb") as f:
    f.write(body)

When the static approach stops working

Most failures fall into a small set of patterns. Check these in order before changing your parser.

  • The status is 200 but the data is missing. The values are probably injected by JavaScript after the page loads. Open the page with JavaScript disabled, or search the raw response for a known value. If it is absent, the static method cannot see it.
  • HTTP 403, 429, or repeated 5xx responses. The server may be refusing automated clients or limiting request rates. Stop, read the site’s published access rules, and reduce your request frequency. Do not switch headers to get around a refusal.
  • Garbled accented characters or symbols. The charset is wrong. Recheck the header and the document’s own declaration, as described in Step 3.
  • The final URL differs from the requested one. The request was redirected. Check whether the new location is still within what robots.txt permits, and run the robots check against the final URL.
  • Timeouts or long pauses. Connection setup can take arbitrarily long. Lower the expectations of a single request, add a retry limit with a pause between attempts, and log the failure rather than retrying without bound.
  • Extraction returns zero items. The markup or keys have changed. Save the raw response, inspect it, and update the selectors or key names, then rerun the validation in Step 5.

Compliance and responsibility

A working script is not the same as permission. Public visibility does not settle whether a collection complies with a site’s terms, its access controls, the privacy expectations of the people described in the data, or the law in your jurisdiction. Check those questions separately from the technical steps above, and limit collection to the fields and frequency you actually need.

“

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

More from the Wire

  1. World desk4 min
    How to Spot an AI Voice Scam Before Sending MoneyDon’t rely on how a caller sounds. Pause, call back through a known number, and verify the emergency with another trusted person before sending money.
  2. Mountain View desk4 min
    Google’s SynthID Detector: How to Check AI-Generated Images, Video and AudioGoogle’s SynthID Detector looks for an embedded watermark in supported images, video and audio. Here is what its results do—and do not—show.
  3. Redmond desk20 min
    How to create a link to File or Folder in Windows 11Windows 11 gives you several ways to point to a file or folder without moving or duplicating it. You can create a desktop shortcut,…
Recommended PC Tool
Recommended PC Tool
PC Slower Than It Used to Be?Free scan - under a minute
Outdated Drivers Are Slowing You DownFree scan - exact matches

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.