Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

To scrape Schema.org Microdata, parse the page as an HTML tree, find elements with itemscope, read each scope’s itemtype and itemid, then collect its itemprop elements recursively. Preserve nested items, repeated properties, machine-readable attributes such as content and href, and properties referenced through itemref. The Python extractor below produces a structured JSON-compatible object and keeps enough source detail to debug the result.

What Schema.org Microdata is (and is not)

Schema.org is a vocabulary of types, such as Movie, Person and Product, and properties such as name and director. Microdata is one HTML syntax for expressing that vocabulary. It is different from JSON-LD, which is normally embedded in a script element, and RDFa, which uses its own attributes.

These attributes define the basic model:

Attribute Meaning
itemscope Starts an item and establishes a boundary for its properties.
itemtype Gives the item’s type URL, for example a Movie type.
itemprop Names one or more properties belonging to the nearest item.
itemid Supplies an identifier when the vocabulary supports one.
itemref Lists element IDs whose properties should also be read for the item.

A nested property is itself an item. In a movie example, a Movie item can have a director property whose element also has itemscope and a Person itemtype. The director’s name belongs to the Person object, not directly to the Movie.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Before you fetch a page

  • Respect the site’s terms, robots rules, authentication requirements and rate limits.
  • Keep the original response body, URL and headers with your extraction output so failures can be reproduced.
  • Decide whether you need the server response or the post-JavaScript DOM. A normal HTTP request cannot see markup inserted only after a browser runs scripts.
  • Use an HTML parser rather than regular expressions; browsers repair malformed HTML and Microdata boundaries are tree-based.

Install the Python dependencies

python -m pip install requests beautifulsoup4

The standard-library JSON encoder is sufficient for output. The parser below uses Beautiful Soup’s HTML parser; in production, you can substitute a standards-oriented parser if malformed markup is common.

Complete Microdata extractor in Python

from __future__ import annotations

import json
import sys
from typing import Any
from urllib.parse import urljoin

import requests
from bs4 import BeautifulSoup, Tag

VALUE_ATTRIBUTES = {
    "meta": "content",
    "audio": "src",
    "embed": "src",
    "iframe": "src",
    "img": "src",
    "source": "src",
    "track": "src",
    "video": "src",
    "a": "href",
    "area": "href",
    "link": "href",
    "object": "data",
    "data": "value",
    "time": "datetime",
}


def element_value(el: Tag, base_url: str | None) -> Any:
    """Return the machine-readable value when the element defines one."""
    attr = VALUE_ATTRIBUTES.get(el.name)
    if attr and el.has_attr(attr):
        value = el.get(attr)
        if isinstance(value, str) and attr in {"href", "src", "data"} and base_url:
            return urljoin(base_url, value)
        return value
    return el.get_text(" ", strip=True)


def item_type(el: Tag) -> list[str]:
    value = el.get("itemtype", [])
    return value if isinstance(value, list) else value.split()


def property_elements(root: Tag) -> list[Tag]:
    """Return direct properties, plus valid itemref targets, without crossing nested items."""
    found: list[Tag] = []
    seen: set[int] = set()

    def walk(container: Tag) -> None:
        for child in container.find_all(recursive=False):
            if not isinstance(child, Tag):
                continue
            if child.has_attr("itemprop"):
                marker = id(child)
                if marker not in seen:
                    seen.add(marker)
                    found.append(child)
            # A nested itemscope is a value of the current property (if it has itemprop),
            # so its descendants belong to that nested item, not this one.
            if child.has_attr("itemscope"):
                continue
            walk(child)

    walk(root)
    for ref in root.get("itemref", []):
        target = root soup.find(id=ref) if False else None
    return found


def extract_item(root: Tag, soup: BeautifulSoup, base_url: str | None, active: set[int] | None = None) -> dict[str, Any]:
    active = set() if active is None else active
    marker = id(root)
    if marker in active:
        return {"cycle": True}
    active.add(marker)
    result: dict[str, Any] = {"type": item_type(root), "properties": {}}
    if root.has_attr("itemid"):
        result["id"] = root.get("itemid")

    elements: list[Tag] = []
    seen: set[int] = set()

    def collect(container: Tag) -> None:
        for child in container.find_all(recursive=False):
            if not isinstance(child, Tag):
                continue
            if child.has_attr("itemprop") and id(child) not in seen:
                seen.add(id(child)); elements.append(child)
            if child.has_attr("itemscope"):
                continue
            collect(child)

    collect(root)
    for ref in root.get("itemref", []):
        target = soup.find(id=ref)
        if target is not None:
            if target.has_attr("itemprop") and id(target) not in seen:
                seen.add(id(target)); elements.append(target)
            elif not target.has_attr("itemscope"):
                collect(target)

    for el in elements:
        value: Any
        if el.has_attr("itemscope"):
            value = extract_item(el, soup, base_url, active)
        else:
            value = element_value(el, base_url)
        for prop in el.get("itemprop", []):
            result["properties"].setdefault(prop, []).append(value)
    active.remove(marker)
    return result


def extract_document(html: str, page_url: str | None = None) -> list[dict[str, Any]]:
    soup = BeautifulSoup(html, "html.parser")
    roots = []
    for el in soup.find_all(attrs={"itemscope": True}):
        # A nested itemscope with itemprop is a property value, not a top-level root.
        parent_item = el.find_parent(attrs={"itemscope": True})
        if parent_item is None:
            roots.append(el)
    return [extract_item(root, soup, page_url) for root in roots]


def main() -> None:
    url = sys.argv[1]
    response = requests.get(url, timeout=30, headers={"User-Agent": "microdata-extractor/1.0"})
    response.raise_for_status()
    print(json.dumps({"url": response.url, "items": extract_document(response.text, response.url)}, indent=2, ensure_ascii=False))

if __name__ == "__main__":
    main()

Save it as microdata.py and run python microdata.py https://example.com/page. The output has a top-level items array. Every property remains an array, so repeated values are not silently overwritten; a nested property is an object with its own type, optional id and properties.

One line in the helper is intentionally not used: the actual itemref handling occurs in extract_item, where the shared Beautiful Soup document is available. This keeps referenced IDs resolvable even when they are outside the item’s descendant subtree.

How the traversal works

Find item roots

Top-level scopes are itemscope elements that have no ancestor itemscope. Nested scopes are retained as property values rather than emitted as unrelated top-level records.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Collect direct properties

The collector walks descendants until it reaches another itemscope. That boundary prevents a nested Person’s properties from leaking into its parent Movie. If a nested scope also carries itemprop, the scope element itself is the parent’s value.

Follow itemref

For each token in itemref, the extractor looks up the matching ID in the same parsed document and collects its property elements. The ID must resolve to an element in that tree. A set of element identities prevents duplicate values when ordinary traversal and a reference reach the same node.

Choose the value correctly

Text is not always the data value. The implementation reads content from meta, URLs from link-like attributes, datetime from time, and falls back to visible text. It also resolves relative URLs against the final response URL. Keep the source element in a richer implementation if you need to distinguish display text from the machine value.

Inspecting a real markup pattern

<div itemscope itemtype="https://schema.org/Movie">
  <h1 itemprop="name">Example film</h1>
  <div itemprop="director" itemscope itemtype="https://schema.org/Person">
    <span itemprop="name">Example director</span>
  </div>
</div>

The resulting shape is conceptually:

{
  "type": ["https://schema.org/Movie"],
  "properties": {
    "name": ["Example film"],
    "director": [{
      "type": ["https://schema.org/Person"],
      "properties": {"name": ["Example director"]}
    }]
  }
}

Do not flatten this structure unless your application deliberately chooses a lossy representation.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

When a plain HTTP fetch misses the data

Some sites insert structured data after JavaScript executes. Compare the saved response with the browser’s final DOM. If the response has no itemscope but the rendered page does, use a browser automation workflow or an API that renders the page. Also check whether the site uses JSON-LD instead of Microdata; an extractor designed only for Microdata should report “no Microdata found” rather than treating JSON-LD as equivalent input.

Or skip the browser setup

If you only need a clean rendered page before parsing it, ScreenshotNeo can capture the URL through one request. It accepts cookie and consent banners as a visitor and removes more than 60 known consent platforms, newsletter popups and chat widgets before capture. Bot checks, blank pages, timeouts, failed loads and cache hits are not billed, and response headers identify the page verdict and billing status. Its MCP server provides take_screenshot, get_page_info and capture_pdf tools for Claude, Cursor and other MCP clients.

See the ScreenshotNeo API documentation for all options. A capture call looks like this:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

The Free plan includes 1,000 screenshots each month with no card. Paid plans start at $5 for 3,000 shots; every feature is available on every plan. Create a free ScreenshotNeo account to begin.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Validation, correctness and Google Search

Extraction answers “what did my parser find in this input?” It does not prove that the HTML is valid, that a crawler received the same markup, or that a page qualifies for a search rich result. Use Schema Markup Validator to inspect Microdata structure and Google’s Rich Results Test for Google feature eligibility. Google documents Microdata, RDFa and JSON-LD as supported formats, while generally recommending JSON-LD when a site’s setup allows it because it is easier to maintain at scale. A valid extraction therefore remains useful even when the page is not eligible for a rich result.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Troubleshooting

No items returned

Check the downloaded HTML, not only the browser view. The page may use JSON-LD, inject Microdata later, require authentication, or serve different markup to bots. Save the response and inspect for itemscope.

Properties are missing

Look for itemref, duplicate IDs, or a nested itemscope that your traversal incorrectly crossed. Confirm that the property element is in the same document tree and that its itemprop token is spelled correctly.

The value is wrong

Inspect attributes before text. Dates commonly belong in datetime, links in href, and metadata in content. Preserve both the selected attribute and visible text if your downstream process needs both.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Duplicate properties appear

Repeated properties are legal, but accidental duplicates can come from following an itemref target and its normal descendant path. Deduplicate by element identity, as the example does, while retaining genuinely separate elements.

Requests fail or are blocked

Use a sensible timeout, identify your client, retry only transient failures with backoff, and do not bypass access controls. A browser-rendering service may be appropriate when the page requires JavaScript, but it does not remove your obligation to follow the site’s rules.

Operational checklist

  • Store the final URL, status code, response headers and raw HTML.
  • Parse with an HTML-aware parser.
  • Emit every top-level item and preserve repeated properties.
  • Keep nested item boundaries and follow itemref.
  • Resolve machine-readable attributes and relative URLs.
  • Test pages with malformed markup, missing types, duplicate IDs and cyclic-looking references.
  • Separate extraction results from validator and search-eligibility results.

FAQ

Can I scrape Microdata with XPath alone?

XPath can select attributes and elements, but you must still implement item boundaries, nested scopes and itemref semantics. A tree parser plus explicit traversal is easier to audit.

Should an extractor reject an item without itemtype?

No. Keep the item with an empty type list and let the caller decide whether an untyped item is useful. Missing type information is different from missing the item entirely.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Does Microdata extraction reveal what Google will display?

No. Search appearance depends on Google’s documentation, crawling and feature-specific requirements. Extraction and eligibility are separate checks.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.