Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

To scrape microformats, fetch a page’s HTML, find microformats2 root classes such as h-card or h-entry, read their prefixed properties, and normalize the results into JSON. The key detail is that some values come from element attributes—such as an anchor’s href or an image’s src—rather than visible text. A robust scraper also preserves nested items and validates the fields it needs.

What microformats scraping extracts

Microformats are conventions for adding semantic, machine-readable information to ordinary HTML. A site can use the same markup for its visible page and for structured data that a scraper can consume. The Microformats project describes a parser as software that takes a URL or HTML input, understands it, and converts it to JSON. Microformats.io

For a scraper, the markup is the source of truth: it does not infer a person, product, or review from page layout alone. It looks for designated root classes and property classes, then interprets their values according to the microformats rules. Coverage therefore depends on whether the target publisher actually includes usable microformats.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Recognize the main microformats2 roots

A root class names the kind of item being described. The properties nested under it supply the item’s fields. Common roots include:

Root Typical subject Example properties or use
h-card Person or organization Name, URL, photo
h-entry Post or other entry Commonly used for published content
h-event Event Event-related information
h-product Product Can be embedded as the item in a review
h-recipe Recipe Name, ingredients, duration, yield, instructions
h-review Review Review name, item, author, publication date, rating, content

These labels are not interchangeable. Select the root that the markup actually declares; do not assume a page’s topic guarantees a particular vocabulary. The MDN overview describes h-card as a format for a person or organization and notes that open-source Microformats2 parsers are available for most languages.

Understand property prefixes and values

Within a root, property prefixes indicate how a value should be interpreted:

  • p- represents a plain-text property, such as p-name or p-ingredient.
  • u- represents a URL property, such as u-url or u-photo.
  • dt- represents a date/time property, such as dt-published or dt-duration.
  • e- represents an embedded HTML or content property, such as e-content or e-instructions.

Do not treat every property as the element’s text. For URL and media properties, parsing rules may give precedence to an element attribute: for example, an <a> uses href, an <img> uses src, and an <object> uses data. Follow the parsing guidance for the element and property instead of collecting visible labels indiscriminately. Microformats2 parsing

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Build a scraping workflow

  1. Fetch responsibly. Retrieve the target HTML in accordance with the site’s terms, robots rules, and rate limits. Keep the requested URL and retrieval timestamp with the result so downstream users can trace its origin.
  2. Parse the document with a Microformats2 parser. Prefer a parser that implements the conventions over a collection of ad hoc selectors. The project’s parser model covers URL input or an HTML fragment and JSON output. Microformats.io
  3. Find roots and properties. Identify roots such as h-card, h-entry, h-event, h-product, h-recipe, and h-review. Read the associated p-, u-, dt-, and e- properties using the parser’s value rules.
  4. Preserve nesting and repeated values. Keep embedded microformats as structured child items, and retain repeated fields as arrays when the parser returns them that way. Flattening everything to strings loses relationships.
  5. Normalize and validate. Map parser output into the shape your application expects, while preserving unknown fields where useful. Validate the fields required by your use case and handle missing, malformed, or conflicting values explicitly.

The project’s documented parsed shape includes an items array, an item type, and a properties object. Treat that as parser output, not as a guarantee that every page supplies every field. Microformats2 parsing

Example: parse an h-recipe

A recipe can mark its title with p-name, repeat p-ingredient for ingredients, use dt-duration for preparation time, p-yield for servings, and place instructions in e-instructions. A parser can represent it as an item with type h-recipe and a properties object; the repeated ingredients should remain repeatable rather than being forced into one string. h-recipe specification and example

There is also a classic hRecipe draft that documents a required recipe name (fn) and at least one ingredient, plus optional yield, instructions, duration, photo, author, publication, nutrition, and tags. That older vocabulary is useful for compatibility with existing markup; for new microformats2 work, use h-recipe conventions rather than assuming the older draft’s property names are identical. Classic hRecipe draft

Example: preserve an h-review and its nested item

An h-review can expose p-name, p-item, p-author, dt-published, p-rating, p-best, p-worst, e-content, p-category, and u-url. Its item may itself be an h-card, h-event, h-geo, h-product, h-recipe, or another h-item. Keep that nested item intact if you need to know what was reviewed rather than only the review’s text. h-review specification

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The h-review specification is marked as a draft and notes possible future convergence with h-entry. For long-lived scrapers, avoid making your data model depend on a single assumption about vocabulary boundaries; retain the parser’s type and build tolerant validation around fields your application actually needs. h-review specification

Convert parser output into useful JSON

A practical output record should retain the parsed item and enough provenance to interpret it later. For example, a normalized recipe record might include the source URL, retrieval time, item type, name, ingredients as a list, duration, yield, and instructions. A review record may include its own fields plus a nested parsed item. The exact JSON property names depend on your application; they are not a separate standard imposed by microformats.

  • Keep URL values as URLs, not anchor labels, when the markup’s URL-bearing element supplies the value.
  • Preserve date/time values and embedded content in forms your application can validate instead of silently converting them to guessed formats.
  • Represent absent properties as absent or null according to your API contract; do not fabricate defaults such as a zero rating.
  • Keep unknown properties or raw parser output when future vocabulary changes could matter.

Handle missing or inconsistent markup

Microformats scraping is publisher-dependent. A page may have no microformats, use only some properties, or contain malformed markup. In those cases, the parser cannot recover fields that were never provided. Decide whether your job should return an empty result, mark the record as incomplete, or apply a documented fallback such as another structured-data format or carefully scoped CSS selectors.

Do not confuse “no parsed item” with a valid item whose properties happen to be empty. Track parse status and validate fields by task: an indexer may accept an h-card without a photo, while a workflow that displays a contact card may require a usable name. The cited material establishes microformats’ conventions and parser model, but does not establish a universal benchmark against CSS selectors, JSON-LD, RDFa, or microdata for speed or accuracy.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Respect access, reliability, and operational limits

  • Check the target site’s terms, robots rules, and rate limits before fetching; use a conservative request rate and avoid unnecessary repeat downloads.
  • Record HTTP failures, timeouts, empty responses, and parse failures separately. A successful HTTP response does not mean the expected microformat exists.
  • Cache fetched HTML when permitted and appropriate, so retries do not create needless traffic. Re-parse stored HTML when your parser changes.
  • Test representative pages, including missing fields, repeated properties, nested items, and attribute-based URLs. Do not assume one page template describes an entire site.

There is no single microformats scraping cost or performance figure established here. Costs and throughput depend on how HTML is fetched, the target site’s behavior, your parser, and your storage and retry design; measure those in your own environment rather than applying a generic benchmark.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Capture a page when the HTML source is inconvenient

For microformats extraction, the HTML is usually the useful input because the conventions live in markup. A screenshot can help inspect what a rendered page looks like, but an image is not a substitute for parsing classes, attributes, nested items, or JSON output. If you need a rendered-page capture alongside your HTML workflow, ScreenshotNeo is a website screenshot API and MCP server; its clean captures remove known consent banners, newsletter popups, and chat widgets before capture, and only clean shots are billed.

Or skip the browser setup

If you need a screenshot of a rendered page in addition to your microformats parser, ScreenshotNeo takes a URL in one request. This does not replace retrieving and parsing HTML for microformats. API options and details are in the ScreenshotNeo documentation.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

ScreenshotNeo removes cookie banners, popups, and chat widgets before the shot; bot checks, blank pages, and failed loads are never billed. Its MCP server lets AI agents take screenshots, and the free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. Sign up for ScreenshotNeo free.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Troubleshooting

Symptom Likely cause What to do
No items in the parser output The document has no recognized microformats roots, or the fetched HTML differs from the page you expected. Inspect the retrieved HTML for root classes such as h-card or h-entry; verify redirects and the returned document.
A URL field contains a label instead of a link The scraper read visible text rather than the relevant URL attribute. Use a Microformats2 parser’s value rules and check href, src, or data where applicable.
Ingredient or category data is missing The publisher may omit the property, use nonconforming markup, or provide repeated values your code collapses incorrectly. Inspect the raw parsed property and source markup; support repeated values and represent absence without inventing data.
A nested author or product is flattened Your normalization step converted child items to strings. Preserve nested parser items as objects and validate their types separately.
A review fails strict schema validation The vocabulary is draft-marked and publishers may vary in their fields. Validate only required fields for your use case and tolerate recognized vocabulary evolution rather than rejecting all partial records.
HTTP fetch succeeds but extraction is empty The page loaded, but it may not publish the target microformat in its returned HTML. Distinguish fetch success from parse success; inspect the response before choosing a fallback extraction method.

Choosing a fallback when microformats are absent

Use microformats when the publisher exposes them and their fields fit your task. If they are absent, choose a fallback based on the target site and your required data rather than assuming another method is universally superior. CSS selectors can follow a known page layout; JSON-LD, RDFa, and microdata are other structured-data approaches. Compare options on publisher coverage, vocabulary stability, maintained parser libraries, nested entity handling, fidelity for dates and URLs, and behavior when markup is malformed or missing. The cited sources do not establish a universal speed or accuracy winner.

Frequently Asked Questions

Do microformats automatically make a page’s data available as JSON?

No. The markup encodes the conventions; a parser must read it and produce JSON. Pages without the relevant markup will not yield those fields.

Should I use the classic hRecipe format for a new scraper?

Use the h-recipe microformats2 vocabulary for new work. The classic hRecipe page is historical compatibility guidance.

Can I scrape microformats from a screenshot?

A screenshot can show the rendered appearance, but it does not reliably expose the semantic classes, attributes, and nested structure needed for microformats parsing.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.