Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
Use a layered extractor. Save the raw response, classify its content, read semantic formats such as JSON-LD, Microdata, and RDFa, then use CSS or XPath selectors for fields that are only present in page structure. If the server response does not contain the data, render the page in a browser and inspect the resulting DOM or network responses. Normalize every value, validate it against visible content, and keep field-level provenance so template changes can be diagnosed.
What “structured data” means
Structured data has two separate layers:
- Vocabulary: the names and relationships for entities, often Schema.org types such as
Article,Product,Event, orPerson. - Encoding: the syntax carrying those terms. Common encodings are JSON-LD, Microdata, and RDFa.
A page can expose all three, none, or several conflicting copies. A successful HTTP request only proves that a response arrived; it does not prove that the desired record is in that response or that the values are current.
A resilient extraction workflow
- Record the response. Store the URL, retrieval timestamp, status, headers, content type, and raw bytes. Do not discard the original before parsing.
- Classify the payload. Handle HTML, XML, JSON, JavaScript, images, and PDFs with the appropriate parser. Use a JSON parser for an actual JSON response rather than treating it as HTML.
- Parse the document tree. Build an HTML/XML tree with a tolerant parser, then apply CSS selectors or XPath.
- Extract semantic graphs. Read every JSON-LD script, Microdata item, and RDFa property before falling back to presentation-oriented selectors.
- Find rendered or embedded data. Inspect inline JSON, JavaScript state objects, and network JSON endpoints. If the required fields appear only after execution or interaction, use a headless browser or hosted rendering service.
- Normalize. Convert dates to one timezone-aware representation, numbers to typed values, URLs to absolute URLs, and repeated entities to stable IDs.
- Validate and preserve provenance. Check required properties, syntax, expected types, duplicates, and conflicts with visible text. For every field, retain its source URL, selector or JSON path, original value, normalized value, and parser version.
Choose the extraction technique by where the data lives
| Technique | Best use | Strength | Typical risk |
|---|---|---|---|
| JSON-LD | Publisher-declared entities and relationships | Easy to parse as a graph and less tied to visual layout | Missing, stale, duplicated, or contradictory values |
| Microdata | Properties attached to HTML elements | Direct relationship between markup and visible elements | Nested items and repeated properties complicate traversal |
| RDFa | Semantic attributes in ordinary HTML | Rich subject, predicate, and object relationships | Less consistent authoring and more involved parsing |
| CSS selectors | Stable classes, IDs, tags, and attributes | Readable and convenient | Break when presentation markup changes |
| XPath | Structural relationships and precise text nodes | Can navigate ancestors, siblings, and conditions | Verbose expressions can become brittle |
| Headless browser | JavaScript-injected DOM or interaction-gated content | Sees the page after rendering like a user agent | Higher startup cost and more failure modes |
Prefer semantic formats for entities when they are complete and agree with what a reader sees. Use CSS or XPath for fields absent from the semantic graph, not as a reason to ignore a richer graph that is already present.
Python: extract JSON-LD, Microdata, and CSS fallbacks
The following script uses requests and beautifulsoup4. It captures the raw response metadata, extracts JSON-LD blocks, collects basic Microdata properties, and applies a CSS fallback. Install dependencies with python -m pip install requests beautifulsoup4.
#1 Best Overall
import json
from datetime import datetime, timezone
from urllib.parse import urljoin
import requests
from bs4 import BeautifulSoup
def text(node):
return " ".join(node.get_text(" ", strip=True).split()) if node else None
def parse_json_ld(soup):
values = []
for script in soup.select('script[type="application/ld+json"]'):
raw = script.string or script.get_text()
try:
values.append(json.loads(raw))
except json.JSONDecodeError as exc:
values.append({"_error": "invalid JSON-LD", "detail": str(exc), "raw": raw})
return values
def parse_microdata(soup, page_url):
records = []
for item in soup.select('[itemscope]'):
record = {"@type": item.get("itemtype"), "properties": {}}
for prop in item.select('[itemprop]'):
name = prop.get("itemprop")
if prop.has_attr("content"):
value = prop["content"]
elif prop.name in ("meta",):
value = prop.get("content")
elif prop.name in ("time",):
value = prop.get("datetime") or text(prop)
elif prop.name in ("a", "link"):
value = urljoin(page_url, prop.get("href", ""))
elif prop.name in ("img", "audio", "video", "source"):
value = urljoin(page_url, prop.get("src", ""))
else:
value = text(prop)
record["properties"].setdefault(name, []).append(value)
records.append(record)
return records
def extract(url):
retrieved = datetime.now(timezone.utc).isoformat()
response = requests.get(url, timeout=30, headers={"User-Agent": "structured-data-extractor/1.0"})
content_type = response.headers.get("content-type", "")
result = {
"source_url": response.url,
"retrieved_at": retrieved,
"status": response.status_code,
"content_type": content_type,
"headers": {"etag": response.headers.get("etag"), "last_modified": response.headers.get("last-modified")},
"json_ld": [], "microdata": [], "fallback": {}, "validation_errors": []
}
response.raise_for_status()
if "html" not in content_type.lower() and "xml" not in content_type.lower():
raise ValueError(f"Expected HTML/XML, received {content_type}")
soup = BeautifulSoup(response.content, "html.parser")
result["json_ld"] = parse_json_ld(soup)
result["microdata"] = parse_microdata(soup, response.url)
title = soup.select_one("h1") or soup.select_one("title")
result["fallback"]["title"] = text(title)
result["fallback"]["canonical"] = (urljoin(response.url, soup.select_one('link[rel="canonical"]')["href"])
if soup.select_one('link[rel="canonical"][href]') else response.url)
for record in result["json_ld"]:
if isinstance(record, dict) and record.get("_error"):
result["validation_errors"].append(record["detail"])
return result
if __name__ == "__main__":
print(json.dumps(extract("https://example.com"), indent=2, ensure_ascii=False))
The script deliberately keeps malformed JSON-LD as an error record instead of silently dropping it. In production, add a schema-specific validator and required-field checks for your target record type.
CSS selectors and XPath in practice
CSS selectors
Use selectors that describe stable meaning: article h1, [data-product-id], or time[datetime]. Avoid generated class names and deeply nested paths. Select all matches when a property can repeat, and retain each match’s selector in provenance.
XPath
XPath is useful when the value is defined by a relationship rather than a class. For example, an XPath expression can locate the heading following a label, select an ancestor containing a price, or target a specific text node. Keep expressions short and test them against fixtures representing every supported template.
Free tools Windows power users keep installed
One-click scans. No signup required.
When BeautifulSoup is enough and when to use lxml
BeautifulSoup offers convenient traversal and tolerant parsing of imperfect HTML, with a performance trade-off for large crawls. lxml provides a fast HTML/XML parser and an ElementTree-style API, including native XPath support. Choose based on document size, XPath needs, and your team’s existing code rather than assuming one parser is universally better.
JSON-LD, Microdata, and RDFa: avoid common traps
JSON-LD
Read every application/ld+json block. A block may be an object, an array, or a graph under @graph; it may also contain multiple entities of different types. Resolve relative URLs, preserve @id values, and do not assume the first object is the page’s primary entity.
Microdata
Start at each top-level itemscope, follow itemprop values, and recursively process nested itemscope elements. Values may come from content, datetime, href, src, or visible text. A property can legitimately occur more than once.
RDFa
RDFa expresses subjects, predicates, and objects through attributes such as about, property, typeof, and resource. Treat extraction as a graph-building task, because a single visible element can contribute a relationship to a subject defined by an ancestor.
Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchPC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Reconcile semantics with visible text
Semantic markup is publisher-provided data, not proof that the value is correct. Compare key fields such as title, price, availability, author, and dates with visible page text. Flag disagreement for review rather than choosing whichever value is easier to parse.
JavaScript-rendered pages and network data
If the initial HTML lacks a field, inspect the page’s inline state, script tags, and network requests. An API response may contain cleaner JSON than the rendered DOM. Respect authentication, robots policies, rate limits, and terms that apply to the site and endpoint.
When rendering is required, use a headless browser to navigate, wait for a meaningful selector or network-idle condition, perform necessary clicks, and then extract the resulting DOM. Record the browser version, viewport, locale, timezone, cookies, and wait condition because each can change the output. A browser should be the final layer, not the first tool for every page: it costs more resources and introduces timeouts, bot checks, consent dialogs, and nondeterministic timing.
Rank #3
Normalization, validation, and provenance
- Dates: parse ISO and locale-specific forms, attach an explicit timezone, and store the original string.
- Numbers: remove presentation separators carefully, preserve currency and units, and reject ambiguous decimal formats.
- URLs: resolve relative references against the final response URL and retain fragments only when they carry meaning.
- Entities: use stable IDs such as
@idwhen supplied; otherwise derive a deterministic key from the source and identifying fields. - Validation: check required properties, data types, allowed values, duplicate entities, malformed markup, and semantic-versus-visible conflicts.
- Provenance: store source URL, retrieval time, selector or JSON path, original value, normalized value, parser version, and validation errors for each field.
Emit a typed record plus errors, rather than returning a dictionary whose missing and invalid values are indistinguishable. Add regression fixtures for representative page templates and monitor extraction completeness over time.
Quick wins for a faster PC:
Repair Windows errors before they cause bigger problemsFix Now →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Clear out junk files and repair common Windows errorsFree Scan →Performance, reliability, and cost decisions
- Reuse HTTP connections, set explicit connect and read timeouts, and retry only transient failures with backoff.
- Cache responses when permitted and send conditional requests using validators such as ETag or Last-Modified.
- Parse once and share the tree among selectors; do not re-download a page for each field.
- Limit concurrency to what the site and your network can handle, and honor applicable crawl policies.
- Use direct JSON endpoints when they are stable and authorized; use a browser only for fields unavailable in the response.
- Measure completeness, validation-error rate, median and tail latency, browser-start failures, and change frequency. Do not infer accuracy or speed from a single page.
Troubleshooting
The request succeeds but fields are empty
Inspect the saved response and content type. The data may be in JSON rather than HTML, hidden in a script, or injected after load. Find the network request or render the page before changing selectors.
JSON-LD parsing fails
Publishers sometimes include comments, trailing commas, HTML entities, or multiple JSON objects in one script. Record the raw block, report a validation error, and use a standards-compliant parser; never evaluate it as executable JavaScript.
Values are duplicated
Pages commonly expose the same entity in JSON-LD, Microdata, and visible markup. Deduplicate using a stable ID and retain every source so disagreements remain visible.
A selector broke after a redesign
Replace generated or positional selectors with semantic attributes, add a fixture for the new template, and alert when required fields fall below your completeness threshold.
A browser times out or sees a challenge
Check DNS, TLS, redirects, viewport, consent dialogs, and wait conditions. Reduce unnecessary resources, use a realistic but policy-compliant configuration, and stop retrying persistent bot checks. A successful browser launch does not guarantee that the target data was delivered.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Or skip the browser setup
ScreenshotNeo is useful when you need a rendered page artifact while your extractor handles the structured fields. One GET request returns a PNG, JPEG, WebP, or PDF. Before capture it accepts cookie or consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each cleanup step can be disabled. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing status.
With an API key, the basic call is:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
See the ScreenshotNeo API documentation for options such as full-page lazy-image capture, CSS-element capture, custom JavaScript, waits, request blocking, cookies, headers, user agents, timezone and geolocation, resizing, caching TTLs, signed links, asynchronous webhooks, bulk capture of up to 100 URLs per call, and usage reporting. It also provides an MCP server with take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients.
Python
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
Node.js
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
Every feature is available on every plan: 1,000 shots per month are free with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.
FAQ
Should I extract JSON-LD or scrape visible HTML?
Extract JSON-LD, Microdata, and RDFa first for semantic entities, then use visible HTML as a validated fallback. Compare both when correctness matters.
Can CSS selectors replace a headless browser?
No. Selectors operate on the tree you provide. If JavaScript creates the required nodes after the initial response, render the page or call an authorized data endpoint first.
Best Value
How do I make an extractor maintainable?
Keep raw responses, field-level provenance, typed output, validation errors, representative fixtures, and completeness monitoring. These make template changes diagnosable instead of mysterious.
Frequently Asked Questions
Should I extract JSON-LD or scrape visible HTML?
Extract JSON-LD, Microdata, and RDFa first for semantic entities, then use visible HTML as a validated fallback. Compare both when correctness matters.
Do these 3 things before closing this tab:
1Repair Windows errors before they cause bigger problems2Fix the driver behind crashes, sound loss and screen glitches3Clear out junk files and repair common Windows errorsCan CSS selectors replace a headless browser?
No. Selectors operate on the tree you provide. If JavaScript creates the required nodes after the initial response, render the page or call an authorized data endpoint first.
How do I make an extractor maintainable?
Keep raw responses, field-level provenance, typed output, validation errors, representative fixtures, and completeness monitoring. These make template changes diagnosable instead of mysterious.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

