Use pandas.read_html() for real HTML tables and Beautiful Soup CSS selectors for repeated cards, products, or list items. A table becomes a list of pandas DataFrame objects; serialize the selected frame with orient='records' for the usual array of JSON objects, orient='values' for unlabeled nested arrays, or orient='table' when consumers need JSON Table Schema metadata. For non-table markup, select each repeated container, extract its child fields, normalize them, and emit one dictionary per item.
Choose the parser from the HTML structure
Inspect the source before writing selectors. A semantic table has <table>, <thead>, <tbody>, <tr>, and usually <th>/<td>. Product grids, search results, news feeds, and navigation menus are often repeated <li>, <article>, or <div> elements instead. The two cases need different extraction strategies.
| Markup | Recommended method | Typical output | Main risk |
|---|---|---|---|
Semantic <table> |
pandas.read_html() |
DataFrame converted to records, values, or table-schema JSON | Multiple tables, merged cells, or malformed markup |
Repeated cards, rows, or <li> elements |
Beautiful Soup select() |
One dictionary per matched container | Selectors break after a redesign |
| Hosted extraction workflow | A selector-based API such as Microlink’s documented table/list primitives | Named, typed fields returned as JSON | Service availability, limits, and cost must be checked for your use case |
Keep the source URL and retrieval timestamp with every extraction. They make a later validation failure diagnosable and distinguish a changed page from a parser bug.
Scrape a semantic HTML table with pandas
Install the parser stack
python -m pip install pandas requests beautifulsoup4 lxml html5lib
read_html() can accept a URL, an HTML string, or a file and always returns a list of DataFrames, even when the page contains only one table. Installing Beautiful Soup and html5lib gives pandas fallback parsers when lxml cannot handle malformed markup.
Do these 3 things before closing this tab:
1Clear out junk files and repair common Windows errors2Scan for outdated or missing drivers - takes under a minute3Repair Windows errors before they cause bigger problems#1 Best Overall
Fetch once, then inspect every table
from io import StringIO
from datetime import datetime, timezone
import requests
import pandas as pd
url = "https://example.com/prices"
response = requests.get(url, timeout=30, headers={"User-Agent": "table-export/1.0"})
response.raise_for_status()
html = response.text
retrieved_at = datetime.now(timezone.utc).isoformat()
tables = pd.read_html(StringIO(html))
print(f"found {len(tables)} tables at {retrieved_at}")
for index, frame in enumerate(tables):
print(index, frame.shape)
print(frame.head(2).to_string(index=False))
Do not assume index zero is the intended table. Select by index only after inspecting the list, or use a distinctive match string or table attributes when the page contains several candidates. A quick row/column count check catches many accidental selections.
Normalize columns and missing values
frame = tables[0].copy()
# Flatten labels if pandas created a MultiIndex from multi-row headers.
if isinstance(frame.columns, pd.MultiIndex):
frame.columns = [" ".join(str(part) for part in column if str(part) != "nan").strip()
for column in frame.columns]
else:
frame.columns = [str(column).strip() for column in frame.columns]
# Normalize text without converting genuine missing values to the string "nan".
for column in frame.select_dtypes(include="object").columns:
frame[column] = frame[column].map(
lambda value: " ".join(value.split()) if isinstance(value, str) else value
)
# Use explicit names that downstream code can rely on.
frame = frame.rename(columns={"Product name": "product_name", "Price": "price"})
frame = frame.where(pd.notna(frame), None)
Expect complications from rowspan, colspan, nested tables, footnote rows, currency symbols, localized decimal separators, and header rows repeated inside the body. Inspect representative records before converting types. A value such as €1.234,50 needs a locale-aware conversion policy; stripping punctuation blindly can change its meaning.
Emit the JSON shape your consumer expects
# Array of objects: the normal API and database-import shape.
records_json = frame.to_json(orient="records", force_ascii=False, indent=2)
# Nested arrays: labels and index are intentionally discarded.
values_json = frame.to_json(orient="values", force_ascii=False, indent=2)
# JSON Table Schema-compatible envelope with fields and data.
table_json = frame.to_json(orient="table", force_ascii=False, indent=2)
print(records_json)
With records, a row such as Product/Price becomes {"product_name":"...","price":...}. values produces arrays such as [["...", 12.5]]; use it only when the receiver already knows column order. table preserves schema metadata and is appropriate when typed field definitions matter. The records orientation is usually the safest contract for a downstream REST API.
Validate before storing
required = {"product_name", "price"}
missing = required - set(frame.columns)
if missing:
raise ValueError(f"missing columns: {sorted(missing)}")
if not 1 <= len(frame) <= 10_000:
raise ValueError(f"unexpected row count: {len(frame)}")
sample = frame.iloc[0].to_dict()
print("validated first row:", sample)
Compare the exported row count with the visible page, check a first, middle, and last row, and verify that links, dates, and numeric fields have the expected types. Keep an alert for zero rows and implausibly large result sets; both often indicate a selector or page-layout change.
Recommended Free Tools
Rank #2
Turn repeated cards or list items into objects with Beautiful Soup
Identify a stable item selector
Select the repeated container, not a page-wide wrapper. Prefer stable attributes such as data-testid, semantic classes, or an item-specific role over generated CSS-module names. Beautiful Soup supports descendant selectors such as body a and direct-child selectors such as head > title.
from bs4 import BeautifulSoup
from urllib.parse import urljoin
import requests
url = "https://example.com/catalog"
response = requests.get(url, timeout=30, headers={"User-Agent": "card-export/1.0"})
response.raise_for_status()
soup = BeautifulSoup(response.text, "html.parser")
items = soup.select("article.product-card")
if not items:
raise RuntimeError("No product cards matched; inspect the current HTML")
rows = []
for item in items:
title_node = item.select_one(".product-card__title")
price_node = item.select_one(".product-card__price")
link_node = item.select_one("a.product-card__link[href]")
rows.append({
"title": title_node.get_text(" ", strip=True) if title_node else None,
"price": price_node.get_text(" ", strip=True) if price_node else None,
"url": urljoin(url, link_node["href"]) if link_node else None,
})
print(rows[:2])
select_one() returns the first matching child and None when a field is absent. That explicit missing value is preferable to crashing or silently shifting fields between records. Use get_text(" ", strip=True) to collapse nested markup and preserve word boundaries.
Normalize and serialize the card records
import json
import re
from decimal import Decimal, InvalidOperation
def clean_space(value):
return re.sub(r"\s+", " ", value).strip() if isinstance(value, str) else value
def parse_price(value):
if value is None:
return None
cleaned = re.sub(r"[^0-9.,-]", "", value)
# Apply a locale-specific rule in production; this example expects a dot decimal.
try:
return float(Decimal(cleaned.replace(",", "")))
except InvalidOperation:
return None
for row in rows:
row["title"] = clean_space(row["title"])
row["price"] = parse_price(row["price"])
with open("products.json", "w", encoding="utf-8") as output:
json.dump(rows, output, ensure_ascii=False, indent=2)
For dates, parse with an explicitly chosen format or ISO-8601 policy. For images, extract src, data-src, or srcset according to the site’s lazy-loading convention. For links, urljoin() converts relative paths into absolute URLs using the response URL.
Handle pagination and lazy content
One request sees only server-rendered HTML. Follow a next link until it disappears, or call the site’s documented JSON endpoint if one exists. Infinite-scroll pages may require browser automation or an endpoint discovered in network tools; do not claim that a static parser captured content that was never in the response. Deduplicate by a stable ID or canonical URL when pages overlap.
Make the extraction resilient
- Stable selectors: anchor to semantic elements and durable attributes; avoid positional selectors such as
:nth-child()unless the layout is fixed. - Cardinality checks: fail loudly on zero matches and review unusually high counts.
- Field checks: require keys that every record needs, while allowing optional fields to be
null. - Schema versioning: record the extractor version, source URL, retrieval time, and schema version beside the JSON.
- Encoding and whitespace: preserve UTF-8, collapse layout whitespace, and keep meaningful punctuation.
- Ethics and operations: respect the site’s terms, robots guidance, authentication requirements, and rate limits. Cache responses where appropriate.
When a hosted selector workflow is useful
A hosted extraction service can fetch pages, match every row or card with selectorAll, and declare named child fields with CSS selectors. This moves scheduling, retries, and deployment out of your code, while local Python gives you maximum control over cleaning and validation. Compare parser tolerance, schema features, operational control, and cost before committing. Availability and limits vary by provider, so verify them for your region and workload.
Or skip the browser setup
If the real prerequisite is obtaining clean page HTML or an image before processing it, ScreenshotNeo provides a single screenshot/PDF endpoint and an MCP server for AI agents. It accepts consent banners as a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each cleanup step can be disabled. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing result.
Use the returned image or PDF as an audit artifact alongside your extracted JSON. ScreenshotNeo is not a replacement for parsing table cells; it is a way to capture the rendered page without maintaining browser setup. See the parameter reference in the ScreenshotNeo documentation.
cURL
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
Node.js
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`HTTP ${res.status}`);
const data = Buffer.from(await res.arrayBuffer());
await Bun.write('shot.webp', data);
ScreenshotNeo’s MCP tools—take_screenshot, get_page_info, and capture_pdf—work with Claude, Cursor, and other MCP clients. It includes full-page and element capture, custom CSS/JavaScript, waits, request blocking, headers, cookies, user agents, geolocation, resizing, caching, signed links, asynchronous webhooks, bulk capture of up to 100 URLs per call, and a usage API. Plans are Free (1,000 shots/month, no card), Starter $5 for 3,000, Growth $15 for 15,000, Pro $39 for 60,000, Scale $99 for 250,000, and Business $249 for 1,000,000; yearly billing provides two months free, and every feature is on every plan. Sign up free to get 1,000 screenshots a month without a card.
Free tools Windows power users keep installed
One-click scans. No signup required.
Troubleshooting common failures
“No tables found” from pandas
The page may render its table with JavaScript, return a bot-check page, or use divs styled as a table. Save and inspect response.text. If no literal <table> exists, use a browser-rendered source or the card-selector approach.
The wrong table was exported
Print every DataFrame’s shape and first rows, then select by a verified index or distinctive content. Pages often include navigation, pricing, and hidden responsive tables.
Cards return zero records
Inspect the downloaded HTML, confirm the selector in a browser’s Elements panel, and check whether content is injected after load. Replace brittle generated classes with stable attributes and add a zero-match alert.
Columns or fields shift
Merged headers, optional badges, and nested links can alter positions. Select by names or child selectors, not by child index. Normalize missing fields to null and validate required keys.
Numbers and dates are wrong
Keep raw text until you have chosen locale rules. Currency symbols, thousands separators, percentages, and localized dates require explicit parsers; log values that fail conversion instead of coercing them silently.
Best Value
Requests time out or receive a challenge
Use a reasonable timeout, a descriptive user agent, retries with backoff, and caching where permitted. A challenge page is not valid source data; detect it by status, title, or known markers and stop rather than saving it as a successful extraction.
FAQ
Does read_html() return JSON directly?
No. It returns a list of DataFrames. Choose the intended frame, clean it, and call to_json().
Which orientation should an API receive?
Use records unless the consumer explicitly requires unlabeled arrays or JSON Table Schema.
Quick wins for a faster PC:
Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Clear out junk files and repair common Windows errorsFree Scan →Can Beautiful Soup parse JavaScript-rendered cards?
Only if the rendered cards are present in the HTML you give it. It does not execute page JavaScript.
How do I detect a site redesign?
Run row-count, required-key, and sample-value assertions on every extraction and alert when they fail.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

