Hardware FixRecommendedDevice not working? Your driver may be the problemCheck updates for common hardware issues.Fix DriversOctober DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsSlow PC?RecommendedPC slow today? Run a repair scan before it gets worseResolve common Windows issues and optimize system performance.Scan Now×
Skip to content
World desk8 min

Data Parsing: How to Turn Web Data into Structured Data

Turn HTML and XML into structured data by matching the parser to the source, defining an output schema, and validating extracted values against real pages.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

To turn web data into structured data, first identify its shape, then choose a parser that matches it: use an HTML tree parser for page elements, a table reader for HTML tables, and an XML reader for XML. Extract into fields you define, normalize values, and validate the result against real source pages before using it downstream. Parsing creates a usable representation; it does not guarantee that the extracted data is complete, correct, or stable.

What parsing does—and what it does not do

Parsing converts source text or markup into a representation a program can inspect and transform. An HTML parser builds a tree of elements; a table reader can turn an HTML table into rows and columns; an XML reader can map nodes and attributes into records.

Parsing is only one part of extraction. You still need to identify the right content, decide how it maps to your output fields, handle missing or inconsistent values, and check that the result matches the source. A successful parse can still produce the wrong table, omit a field, or reflect a page structure that has since changed.

Choose a parser based on the source and target

Source shape Practical starting point Result and important caveat
HTML page with information in headings, links, or containers Beautiful Soup with a selected parser A navigable parse tree from which you select text or attributes. Different parsers may build different trees from malformed HTML.
HTML table pandas read_html() A list of DataFrames, even if the page contains just one table. Inspect the list and select the intended table.
XML with repeating, relatively shallow records pandas read_xml() A DataFrame built from nodes and attributes. Deeply nested XML may need to be flattened first.
Recurring extraction from pages that can change A maintained workflow with validation and error reporting Selectors and assumptions can break when the source changes; detect missing or unexpected output and revise the extraction.

These are starting points, not universal solutions. Consider input shape, desired output (such as a parse tree, DataFrame, CSV, or JSON), markup quality, dependencies, and how you will detect changes. Beautiful Soup describes itself as “a Python library for pulling data out of HTML and XML files.” Its documentation discusses lxml, html5lib, and Python’s built-in html.parser; the same malformed input can produce different trees with different parsers. Choose by checking the tree you actually get, not by assuming one parser is always best.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
#1 Best Overall
Sale
Storytelling with Data: A Data Visualization Guide for Business Professionals
  • Wiley
  • Language: english
  • Book - storytelling with data: a data visualization guide for business professionals

A practical workflow for turning a page into records

  1. Inspect a representative source. Determine whether the information is a table, repeated set of records, linked attribute, or nested structure. Check whether the useful content is present in the markup you have; a page may also rely on scripts, and no single parsing method applies to every dynamically rendered page.
  2. Define the output schema. Write down field names and expected types before extraction. Decide how to represent missing values, duplicates, and inconsistent formats. Include source context such as the page URL or a record identifier when it matters.
  3. Choose a parser for that shape. Use a tree parser for selected page elements, a table reader for HTML tables, or an XML reader for XML. Check the documented input and output behavior of the tool.
  4. Extract and normalize deliberately. Select only the target fields, trim whitespace, normalize values consistently, and convert types with explicit handling for invalid or missing values.
  5. Validate against the source. Check required fields, record counts or expected records, usable types, and a few representative values by comparing them with the original page or document.
  6. Monitor recurring runs. Alert on empty output, missing required fields, or unexpected changes. Review the source and update selectors or transformations when the page structure changes.

Extract page elements with Beautiful Soup

Beautiful Soup is useful when target values are spread across page elements rather than organized in one table. The following example parses saved HTML, selects repeated article cards by CSS class, and emits JSON records. It uses Python’s built-in html.parser; if the markup is malformed, inspect the resulting tree and compare another supported parser if necessary.

from bs4 import BeautifulSoup
import json

html = """<main>
  <article class="card">
    <h2>Example item</h2>
    <a href="/items/1">Details</a>
  </article>
</main>"""

soup = BeautifulSoup(html, "html.parser")
records = []

for card in soup.select("article.card"):
    title = card.select_one("h2")
    link = card.select_one("a[href]")
    if title is None or link is None:
        continue
    records.append({
        "title": title.get_text(" ", strip=True),
        "href": link["href"].strip(),
    })

print(json.dumps(records, ensure_ascii=False, indent=2))

This example intentionally skips cards without both required elements. For production extraction, do not let skipped records go unnoticed: count them or report them as validation failures if those fields are required. If the target data is an attribute, select the relevant element and read that attribute rather than its visible text.

Read an HTML table into pandas

Use read_html() when the data is actually represented in an HTML table. In pandas 3.0.6 documentation, the function accepts HTML strings, files, or URLs and returns a list of DataFrames. The list behavior applies even when only one table is found.

from io import StringIO
import pandas as pd

html = """<table>
  <thead><tr><th>Product</th><th>Price</th></tr></thead>
  <tbody>
    <tr><td>A</td><td>12.50</td></tr>
    <tr><td>B</td><td>15.00</td></tr>
  </tbody>
</table>"""

tables = pd.read_html(StringIO(html))
if not tables:
    raise ValueError("No HTML tables were found")

df = tables[0]
required = {"Product", "Price"}
missing = required.difference(df.columns)
if missing:
    raise ValueError(f"Missing expected columns: {sorted(missing)}")

df["Price"] = pd.to_numeric(df["Price"], errors="raise")
print(df.to_json(orient="records", indent=2))

When a page has multiple tables, inspect each candidate’s headers and rows before choosing an index; table order alone may not identify the intended data. The example uses an in-memory HTML string. For a file or URL, use the input form supported by your pandas version and account for access, network, and page-rendering requirements separately from parsing.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Parse XML into a DataFrame

For repeating, shallow XML records, pandas read_xml() can read nodes and attributes into a DataFrame. XML does not have one universal record structure, so select the repeating element that represents one row.

from io import StringIO
import pandas as pd

xml = """<catalog>
  <item id="1"><name>A</name><price>12.50</price></item>
  <item id="2"><name>B</name><price>15.00</price></item>
</catalog>"""

df = pd.read_xml(StringIO(xml), xpath=".//item")
print(df.to_dict(orient="records"))

The selected item nodes become records; the id attribute and child elements become columns. pandas documentation says read_xml() works best with flatter, shallow XML. If the values you need are deeply nested, transform the structure first—for example, with an appropriate stylesheet—or use a workflow that explicitly maps nested nodes into your target schema.

Validate the output before relying on it

Validation is an application-level responsibility: the parsing functions do not automatically prove that the result conforms to your business schema. Add checks appropriate to the data and fail visibly when essential assumptions stop holding.

  • Presence: Are all required fields or columns present and non-empty where expected?
  • Shape: Did you find the intended table, nodes, or repeated records rather than unrelated page content?
  • Types and formats: Can numeric, date, and identifier fields be interpreted consistently without silently turning invalid values into plausible output?
  • Coverage: Is the number of records plausible, and do records with missing fields get reported rather than silently discarded?
  • Accuracy: Do sample values match the source page or document?
  • Context and privacy: Is source information retained where useful, and is any personal data handled with appropriate safeguards?

Why web extraction breaks, and how to make it maintainable

Real pages include navigation, ads, tracking scripts, and nested elements that can obscure the content you want. A selector that once matched the right element may later match nothing—or a different element—after a page redesign. Malformed markup can also yield different trees depending on the parser. For recurring jobs, keep a representative fixture, test required fields and expected output shape, and log failures and source context so a change can be diagnosed.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Choose a workflow by its input shape, output needs, tolerance for malformed markup, dependencies, and ease of maintenance. There is no basis here for a universal speed or accuracy ranking among parsers. If processing volume, accuracy, or personal data is important, account for those requirements explicitly rather than assuming that parsing alone resolves them.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup

Parsing extracts structured fields from HTML or XML; a screenshot is an image or PDF, not structured data. Screenshot capture can still help when you need a visual record of a page to inspect or archive alongside an extraction. ScreenshotNeo is a website screenshot API and MCP server; it is not a replacement for an HTML or XML parser.

For example, this cURL request saves a screenshot of a page as WebP. See the ScreenshotNeo API documentation for request options.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
  • Cookie banners are accepted and more than 60 known consent platforms, newsletter popups, and chat widgets are removed before capture; each of those steps can be turned off.
  • Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing; responses identify the page verdict and billing status in headers.
  • An MCP server provides take_screenshot, get_page_info, and capture_pdf tools for AI agents and MCP clients.
  • The Free plan includes 1,000 shots per month with no card; paid plans start at $5 for 3,000 shots. Yearly billing gives two months free, and every feature is on every plan.

Sign up for 1,000 free screenshots a month with no card.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Frequently asked questions

Can parsing guarantee that extracted data is correct?

No. Parsing turns markup into a representation; selection, normalization, source quality, and validation determine whether the fields are useful and accurate.

Should I use Beautiful Soup or pandas read_html()?

Use Beautiful Soup for targeted elements distributed through a page, and read_html() when the information is in an HTML table. The source structure and desired output should decide.

Why does my web scraper stop working after a page changes?

Page changes can invalidate selectors or change the markup tree. Monitor required fields and output shape, then inspect the updated page and revise the extraction rules.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

More from the Wire

  1. Shenzhen desk3 min
    HONOR Expands Beyond Smartphones With Humanoid Robot RevealHONOR said it unveiled its first humanoid robot at MWC 2026 and named shopping assistance, workplace inspections, and supportive companionship as intended uses. Later Robotics D1 claims and a reported…
  2. Cupertino desk5 min
    Apple Unveils AirPods Max 2: The Upgrade That Should Have Happened Years AgoAirPods Max 2 adds H2-powered audio features and Apple claims up to 1.5× more effective ANC, but its design, Smart Case, and 20-hour battery rating are unchanged. Wired lossless audio…
  3. Cupertino desk4 min
    Apple’s OLED Touch MacBooks Are Coming—but the Dynamic Island Is the Real GambleApple has not announced an OLED touchscreen MacBook, but reports point to high-end models arriving in late 2026 or early 2027. The reported Mac Dynamic Island could be useful, but…
Recommended PC Tool
Recommended PC Tool
Crashes, No Sound, or Screen Glitches?Free driver scan
Windows Errors? Fix Them Before They SpreadFree repair scan

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.