Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
To scrape microformats, fetch a page’s HTML, find microformats2 root classes such as h-card or h-entry, read their prefixed properties, and normalize the results into JSON. The key detail is that some values come from element attributes—such as an anchor’s href or an image’s src—rather than visible text. A robust scraper also preserves nested items and validates the fields it needs.
What microformats scraping extracts
Microformats are conventions for adding semantic, machine-readable information to ordinary HTML. A site can use the same markup for its visible page and for structured data that a scraper can consume. The Microformats project describes a parser as software that takes a URL or HTML input, understands it, and converts it to JSON. Microformats.io
For a scraper, the markup is the source of truth: it does not infer a person, product, or review from page layout alone. It looks for designated root classes and property classes, then interprets their values according to the microformats rules. Coverage therefore depends on whether the target publisher actually includes usable microformats.
Recognize the main microformats2 roots
A root class names the kind of item being described. The properties nested under it supply the item’s fields. Common roots include:
#1 Best Overall
| Root | Typical subject | Example properties or use |
|---|---|---|
h-card |
Person or organization | Name, URL, photo |
h-entry |
Post or other entry | Commonly used for published content |
h-event |
Event | Event-related information |
h-product |
Product | Can be embedded as the item in a review |
h-recipe |
Recipe | Name, ingredients, duration, yield, instructions |
h-review |
Review | Review name, item, author, publication date, rating, content |
These labels are not interchangeable. Select the root that the markup actually declares; do not assume a page’s topic guarantees a particular vocabulary. The MDN overview describes h-card as a format for a person or organization and notes that open-source Microformats2 parsers are available for most languages.
Understand property prefixes and values
Within a root, property prefixes indicate how a value should be interpreted:
p-represents a plain-text property, such asp-nameorp-ingredient.u-represents a URL property, such asu-urloru-photo.dt-represents a date/time property, such asdt-publishedordt-duration.e-represents an embedded HTML or content property, such ase-contentore-instructions.
Do not treat every property as the element’s text. For URL and media properties, parsing rules may give precedence to an element attribute: for example, an <a> uses href, an <img> uses src, and an <object> uses data. Follow the parsing guidance for the element and property instead of collecting visible labels indiscriminately. Microformats2 parsing
Quick wins for a faster PC:
Clear out junk files and repair common Windows errorsFree Scan →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Repair Windows errors before they cause bigger problemsFix Now →Build a scraping workflow
- Fetch responsibly. Retrieve the target HTML in accordance with the site’s terms, robots rules, and rate limits. Keep the requested URL and retrieval timestamp with the result so downstream users can trace its origin.
- Parse the document with a Microformats2 parser. Prefer a parser that implements the conventions over a collection of ad hoc selectors. The project’s parser model covers URL input or an HTML fragment and JSON output. Microformats.io
- Find roots and properties. Identify roots such as
h-card,h-entry,h-event,h-product,h-recipe, andh-review. Read the associatedp-,u-,dt-, ande-properties using the parser’s value rules. - Preserve nesting and repeated values. Keep embedded microformats as structured child items, and retain repeated fields as arrays when the parser returns them that way. Flattening everything to strings loses relationships.
- Normalize and validate. Map parser output into the shape your application expects, while preserving unknown fields where useful. Validate the fields required by your use case and handle missing, malformed, or conflicting values explicitly.
The project’s documented parsed shape includes an items array, an item type, and a properties object. Treat that as parser output, not as a guarantee that every page supplies every field. Microformats2 parsing
Example: parse an h-recipe
A recipe can mark its title with p-name, repeat p-ingredient for ingredients, use dt-duration for preparation time, p-yield for servings, and place instructions in e-instructions. A parser can represent it as an item with type h-recipe and a properties object; the repeated ingredients should remain repeatable rather than being forced into one string. h-recipe specification and example
There is also a classic hRecipe draft that documents a required recipe name (fn) and at least one ingredient, plus optional yield, instructions, duration, photo, author, publication, nutrition, and tags. That older vocabulary is useful for compatibility with existing markup; for new microformats2 work, use h-recipe conventions rather than assuming the older draft’s property names are identical. Classic hRecipe draft
Rank #3
Example: preserve an h-review and its nested item
An h-review can expose p-name, p-item, p-author, dt-published, p-rating, p-best, p-worst, e-content, p-category, and u-url. Its item may itself be an h-card, h-event, h-geo, h-product, h-recipe, or another h-item. Keep that nested item intact if you need to know what was reviewed rather than only the review’s text. h-review specification
The h-review specification is marked as a draft and notes possible future convergence with h-entry. For long-lived scrapers, avoid making your data model depend on a single assumption about vocabulary boundaries; retain the parser’s type and build tolerant validation around fields your application actually needs. h-review specification
Convert parser output into useful JSON
A practical output record should retain the parsed item and enough provenance to interpret it later. For example, a normalized recipe record might include the source URL, retrieval time, item type, name, ingredients as a list, duration, yield, and instructions. A review record may include its own fields plus a nested parsed item. The exact JSON property names depend on your application; they are not a separate standard imposed by microformats.
- Keep URL values as URLs, not anchor labels, when the markup’s URL-bearing element supplies the value.
- Preserve date/time values and embedded content in forms your application can validate instead of silently converting them to guessed formats.
- Represent absent properties as absent or null according to your API contract; do not fabricate defaults such as a zero rating.
- Keep unknown properties or raw parser output when future vocabulary changes could matter.
Handle missing or inconsistent markup
Microformats scraping is publisher-dependent. A page may have no microformats, use only some properties, or contain malformed markup. In those cases, the parser cannot recover fields that were never provided. Decide whether your job should return an empty result, mark the record as incomplete, or apply a documented fallback such as another structured-data format or carefully scoped CSS selectors.
Do not confuse “no parsed item” with a valid item whose properties happen to be empty. Track parse status and validate fields by task: an indexer may accept an h-card without a photo, while a workflow that displays a contact card may require a usable name. The cited material establishes microformats’ conventions and parser model, but does not establish a universal benchmark against CSS selectors, JSON-LD, RDFa, or microdata for speed or accuracy.
Do these 3 things before closing this tab:
1Scan for outdated or missing drivers - takes under a minute2Clear out junk files and repair common Windows errors3Fix the driver behind crashes, sound loss and screen glitchesRespect access, reliability, and operational limits
- Check the target site’s terms, robots rules, and rate limits before fetching; use a conservative request rate and avoid unnecessary repeat downloads.
- Record HTTP failures, timeouts, empty responses, and parse failures separately. A successful HTTP response does not mean the expected microformat exists.
- Cache fetched HTML when permitted and appropriate, so retries do not create needless traffic. Re-parse stored HTML when your parser changes.
- Test representative pages, including missing fields, repeated properties, nested items, and attribute-based URLs. Do not assume one page template describes an entire site.
There is no single microformats scraping cost or performance figure established here. Costs and throughput depend on how HTML is fetched, the target site’s behavior, your parser, and your storage and retry design; measure those in your own environment rather than applying a generic benchmark.
Best Value
Capture a page when the HTML source is inconvenient
For microformats extraction, the HTML is usually the useful input because the conventions live in markup. A screenshot can help inspect what a rendered page looks like, but an image is not a substitute for parsing classes, attributes, nested items, or JSON output. If you need a rendered-page capture alongside your HTML workflow, ScreenshotNeo is a website screenshot API and MCP server; its clean captures remove known consent banners, newsletter popups, and chat widgets before capture, and only clean shots are billed.
Or skip the browser setup
If you need a screenshot of a rendered page in addition to your microformats parser, ScreenshotNeo takes a URL in one request. This does not replace retrieving and parsing HTML for microformats. API options and details are in the ScreenshotNeo documentation.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
ScreenshotNeo removes cookie banners, popups, and chat widgets before the shot; bot checks, blank pages, and failed loads are never billed. Its MCP server lets AI agents take screenshots, and the free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. Sign up for ScreenshotNeo free.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Troubleshooting
| Symptom | Likely cause | What to do |
|---|---|---|
| No items in the parser output | The document has no recognized microformats roots, or the fetched HTML differs from the page you expected. | Inspect the retrieved HTML for root classes such as h-card or h-entry; verify redirects and the returned document. |
| A URL field contains a label instead of a link | The scraper read visible text rather than the relevant URL attribute. | Use a Microformats2 parser’s value rules and check href, src, or data where applicable. |
| Ingredient or category data is missing | The publisher may omit the property, use nonconforming markup, or provide repeated values your code collapses incorrectly. | Inspect the raw parsed property and source markup; support repeated values and represent absence without inventing data. |
| A nested author or product is flattened | Your normalization step converted child items to strings. | Preserve nested parser items as objects and validate their types separately. |
| A review fails strict schema validation | The vocabulary is draft-marked and publishers may vary in their fields. | Validate only required fields for your use case and tolerate recognized vocabulary evolution rather than rejecting all partial records. |
| HTTP fetch succeeds but extraction is empty | The page loaded, but it may not publish the target microformat in its returned HTML. | Distinguish fetch success from parse success; inspect the response before choosing a fallback extraction method. |
Choosing a fallback when microformats are absent
Use microformats when the publisher exposes them and their fields fit your task. If they are absent, choose a fallback based on the target site and your required data rather than assuming another method is universally superior. CSS selectors can follow a known page layout; JSON-LD, RDFa, and microdata are other structured-data approaches. Compare options on publisher coverage, vocabulary stability, maintained parser libraries, nested entity handling, fidelity for dates and URLs, and behavior when markup is malformed or missing. The cited sources do not establish a universal speed or accuracy winner.
Frequently Asked Questions
Do microformats automatically make a page’s data available as JSON?
No. The markup encodes the conventions; a parser must read it and produce JSON. Pages without the relevant markup will not yield those fields.
Should I use the classic hRecipe format for a new scraper?
Use the h-recipe microformats2 vocabulary for new work. The classic hRecipe page is historical compatibility guidance.
Can I scrape microformats from a screenshot?
A screenshot can show the rendered appearance, but it does not reliably expose the semantic classes, attributes, and nested structure needed for microformats parsing.
The Tool Desk
Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

