Free tools Windows power users keep installed
One-click scans. No signup required.
Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
To scrape Schema.org Microdata, parse the page as an HTML tree, find elements with itemscope, read each scope’s itemtype and itemid, then collect its itemprop elements recursively. Preserve nested items, repeated properties, machine-readable attributes such as content and href, and properties referenced through itemref. The Python extractor below produces a structured JSON-compatible object and keeps enough source detail to debug the result.
What Schema.org Microdata is (and is not)
Schema.org is a vocabulary of types, such as Movie, Person and Product, and properties such as name and director. Microdata is one HTML syntax for expressing that vocabulary. It is different from JSON-LD, which is normally embedded in a script element, and RDFa, which uses its own attributes.
These attributes define the basic model:
| Attribute | Meaning |
|---|---|
itemscope |
Starts an item and establishes a boundary for its properties. |
itemtype |
Gives the item’s type URL, for example a Movie type. |
itemprop |
Names one or more properties belonging to the nearest item. |
itemid |
Supplies an identifier when the vocabulary supports one. |
itemref |
Lists element IDs whose properties should also be read for the item. |
A nested property is itself an item. In a movie example, a Movie item can have a director property whose element also has itemscope and a Person itemtype. The director’s name belongs to the Person object, not directly to the Movie.
Before you fetch a page
- Respect the site’s terms, robots rules, authentication requirements and rate limits.
- Keep the original response body, URL and headers with your extraction output so failures can be reproduced.
- Decide whether you need the server response or the post-JavaScript DOM. A normal HTTP request cannot see markup inserted only after a browser runs scripts.
- Use an HTML parser rather than regular expressions; browsers repair malformed HTML and Microdata boundaries are tree-based.
Install the Python dependencies
python -m pip install requests beautifulsoup4
The standard-library JSON encoder is sufficient for output. The parser below uses Beautiful Soup’s HTML parser; in production, you can substitute a standards-oriented parser if malformed markup is common.
#1 Best Overall
Complete Microdata extractor in Python
from __future__ import annotations
import json
import sys
from typing import Any
from urllib.parse import urljoin
import requests
from bs4 import BeautifulSoup, Tag
VALUE_ATTRIBUTES = {
"meta": "content",
"audio": "src",
"embed": "src",
"iframe": "src",
"img": "src",
"source": "src",
"track": "src",
"video": "src",
"a": "href",
"area": "href",
"link": "href",
"object": "data",
"data": "value",
"time": "datetime",
}
def element_value(el: Tag, base_url: str | None) -> Any:
"""Return the machine-readable value when the element defines one."""
attr = VALUE_ATTRIBUTES.get(el.name)
if attr and el.has_attr(attr):
value = el.get(attr)
if isinstance(value, str) and attr in {"href", "src", "data"} and base_url:
return urljoin(base_url, value)
return value
return el.get_text(" ", strip=True)
def item_type(el: Tag) -> list[str]:
value = el.get("itemtype", [])
return value if isinstance(value, list) else value.split()
def property_elements(root: Tag) -> list[Tag]:
"""Return direct properties, plus valid itemref targets, without crossing nested items."""
found: list[Tag] = []
seen: set[int] = set()
def walk(container: Tag) -> None:
for child in container.find_all(recursive=False):
if not isinstance(child, Tag):
continue
if child.has_attr("itemprop"):
marker = id(child)
if marker not in seen:
seen.add(marker)
found.append(child)
# A nested itemscope is a value of the current property (if it has itemprop),
# so its descendants belong to that nested item, not this one.
if child.has_attr("itemscope"):
continue
walk(child)
walk(root)
for ref in root.get("itemref", []):
target = root soup.find(id=ref) if False else None
return found
def extract_item(root: Tag, soup: BeautifulSoup, base_url: str | None, active: set[int] | None = None) -> dict[str, Any]:
active = set() if active is None else active
marker = id(root)
if marker in active:
return {"cycle": True}
active.add(marker)
result: dict[str, Any] = {"type": item_type(root), "properties": {}}
if root.has_attr("itemid"):
result["id"] = root.get("itemid")
elements: list[Tag] = []
seen: set[int] = set()
def collect(container: Tag) -> None:
for child in container.find_all(recursive=False):
if not isinstance(child, Tag):
continue
if child.has_attr("itemprop") and id(child) not in seen:
seen.add(id(child)); elements.append(child)
if child.has_attr("itemscope"):
continue
collect(child)
collect(root)
for ref in root.get("itemref", []):
target = soup.find(id=ref)
if target is not None:
if target.has_attr("itemprop") and id(target) not in seen:
seen.add(id(target)); elements.append(target)
elif not target.has_attr("itemscope"):
collect(target)
for el in elements:
value: Any
if el.has_attr("itemscope"):
value = extract_item(el, soup, base_url, active)
else:
value = element_value(el, base_url)
for prop in el.get("itemprop", []):
result["properties"].setdefault(prop, []).append(value)
active.remove(marker)
return result
def extract_document(html: str, page_url: str | None = None) -> list[dict[str, Any]]:
soup = BeautifulSoup(html, "html.parser")
roots = []
for el in soup.find_all(attrs={"itemscope": True}):
# A nested itemscope with itemprop is a property value, not a top-level root.
parent_item = el.find_parent(attrs={"itemscope": True})
if parent_item is None:
roots.append(el)
return [extract_item(root, soup, page_url) for root in roots]
def main() -> None:
url = sys.argv[1]
response = requests.get(url, timeout=30, headers={"User-Agent": "microdata-extractor/1.0"})
response.raise_for_status()
print(json.dumps({"url": response.url, "items": extract_document(response.text, response.url)}, indent=2, ensure_ascii=False))
if __name__ == "__main__":
main()
Save it as microdata.py and run python microdata.py https://example.com/page. The output has a top-level items array. Every property remains an array, so repeated values are not silently overwritten; a nested property is an object with its own type, optional id and properties.
One line in the helper is intentionally not used: the actual itemref handling occurs in extract_item, where the shared Beautiful Soup document is available. This keeps referenced IDs resolvable even when they are outside the item’s descendant subtree.
How the traversal works
Find item roots
Top-level scopes are itemscope elements that have no ancestor itemscope. Nested scopes are retained as property values rather than emitted as unrelated top-level records.
Collect direct properties
The collector walks descendants until it reaches another itemscope. That boundary prevents a nested Person’s properties from leaking into its parent Movie. If a nested scope also carries itemprop, the scope element itself is the parent’s value.
Rank #2
Follow itemref
For each token in itemref, the extractor looks up the matching ID in the same parsed document and collects its property elements. The ID must resolve to an element in that tree. A set of element identities prevents duplicate values when ordinary traversal and a reference reach the same node.
Choose the value correctly
Text is not always the data value. The implementation reads content from meta, URLs from link-like attributes, datetime from time, and falls back to visible text. It also resolves relative URLs against the final response URL. Keep the source element in a richer implementation if you need to distinguish display text from the machine value.
Inspecting a real markup pattern
<div itemscope itemtype="https://schema.org/Movie">
<h1 itemprop="name">Example film</h1>
<div itemprop="director" itemscope itemtype="https://schema.org/Person">
<span itemprop="name">Example director</span>
</div>
</div>
The resulting shape is conceptually:
{
"type": ["https://schema.org/Movie"],
"properties": {
"name": ["Example film"],
"director": [{
"type": ["https://schema.org/Person"],
"properties": {"name": ["Example director"]}
}]
}
}
Do not flatten this structure unless your application deliberately chooses a lossy representation.
Quick wins for a faster PC:
Scan for outdated or missing drivers - takes under a minuteDriver Scan →Clear out junk files and repair common Windows errorsFree Scan →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →When a plain HTTP fetch misses the data
Some sites insert structured data after JavaScript executes. Compare the saved response with the browser’s final DOM. If the response has no itemscope but the rendered page does, use a browser automation workflow or an API that renders the page. Also check whether the site uses JSON-LD instead of Microdata; an extractor designed only for Microdata should report “no Microdata found” rather than treating JSON-LD as equivalent input.
Or skip the browser setup
If you only need a clean rendered page before parsing it, ScreenshotNeo can capture the URL through one request. It accepts cookie and consent banners as a visitor and removes more than 60 known consent platforms, newsletter popups and chat widgets before capture. Bot checks, blank pages, timeouts, failed loads and cache hits are not billed, and response headers identify the page verdict and billing status. Its MCP server provides take_screenshot, get_page_info and capture_pdf tools for Claude, Cursor and other MCP clients.
See the ScreenshotNeo API documentation for all options. A capture call looks like this:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
The Free plan includes 1,000 screenshots each month with no card. Paid plans start at $5 for 3,000 shots; every feature is available on every plan. Create a free ScreenshotNeo account to begin.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Validation, correctness and Google Search
Extraction answers “what did my parser find in this input?” It does not prove that the HTML is valid, that a crawler received the same markup, or that a page qualifies for a search rich result. Use Schema Markup Validator to inspect Microdata structure and Google’s Rich Results Test for Google feature eligibility. Google documents Microdata, RDFa and JSON-LD as supported formats, while generally recommending JSON-LD when a site’s setup allows it because it is easier to maintain at scale. A valid extraction therefore remains useful even when the page is not eligible for a rich result.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Troubleshooting
No items returned
Check the downloaded HTML, not only the browser view. The page may use JSON-LD, inject Microdata later, require authentication, or serve different markup to bots. Save the response and inspect for itemscope.
Properties are missing
Look for itemref, duplicate IDs, or a nested itemscope that your traversal incorrectly crossed. Confirm that the property element is in the same document tree and that its itemprop token is spelled correctly.
The value is wrong
Inspect attributes before text. Dates commonly belong in datetime, links in href, and metadata in content. Preserve both the selected attribute and visible text if your downstream process needs both.
Duplicate properties appear
Repeated properties are legal, but accidental duplicates can come from following an itemref target and its normal descendant path. Deduplicate by element identity, as the example does, while retaining genuinely separate elements.
Requests fail or are blocked
Use a sensible timeout, identify your client, retry only transient failures with backoff, and do not bypass access controls. A browser-rendering service may be appropriate when the page requires JavaScript, but it does not remove your obligation to follow the site’s rules.
Best Value
Operational checklist
- Store the final URL, status code, response headers and raw HTML.
- Parse with an HTML-aware parser.
- Emit every top-level item and preserve repeated properties.
- Keep nested item boundaries and follow
itemref. - Resolve machine-readable attributes and relative URLs.
- Test pages with malformed markup, missing types, duplicate IDs and cyclic-looking references.
- Separate extraction results from validator and search-eligibility results.
FAQ
Can I scrape Microdata with XPath alone?
XPath can select attributes and elements, but you must still implement item boundaries, nested scopes and itemref semantics. A tree parser plus explicit traversal is easier to audit.
Should an extractor reject an item without itemtype?
No. Keep the item with an empty type list and let the caller decide whether an untyped item is useful. Missing type information is different from missing the item entirely.
Recommended Free Tools
Does Microdata extraction reveal what Google will display?
No. Search appearance depends on Google’s documentation, crawling and feature-specific requirements. Extraction and eligibility are separate checks.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

