October DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsPC HealthRecommendedCrashes, freezes, slowdowns? Check your PC nowSpot repairable issues before they interrupt work.Check PCOctober DealsAmazon USDeal season is back - check today's better picksAmazon US: current deals, useful picks and tech finds.See Picks×
Skip to content
JavaScript

How to Extract Structured Data From a Webpage as JSON (JSON-LD, Microdata, and RDFa)

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The dependable way to extract structured data is a staged pipeline: fetch the page, determine whether JavaScript is required, parse JSON-LD, traverse Microdata and RDFa, preserve graph relationships and provenance, then validate the combined result. A static HTTP client is faster and easier to reproduce; a browser-rendered DOM is necessary when scripts inject the markup after load.

1. Decide whether to fetch HTML or render a browser

Start by requesting the page with an HTTP client. Inspect the response body for script type="application/ld+json", Microdata attributes such as itemscope, or RDFa attributes such as typeof and property. If the required information is present in the response, a static parser is normally the best choice: it is quicker, cheaper in CPU and memory, and produces reproducible results.

Some sites build structured data only after JavaScript runs. In that case, use a browser automation environment, wait for the page to settle, and inspect the final DOM. Google documents that JSON-LD generated by JavaScript can be processed when it is available in the rendered DOM. Capture network responses too when an application obtains the structured payload from an API, because the response may be more complete than the visible markup.

What to record during acquisition

  • The requested URL and the final URL after redirects.
  • HTTP status, content type, character encoding, and retrieval time.
  • The original HTML or rendered DOM used for extraction.
  • Whether the result came from a static response or a browser render.

Do not assume that a successful HTTP status means the page contains the data. A consent wall, bot check, empty application shell, or timeout can all produce a technically valid response with no useful entities.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

2. Parse JSON-LD without losing its graph

JSON-LD is usually the easiest layer to extract because each block is contained in a script element. It is still a graph-capable format, not merely a flat record. Preserve @context, @type, @id, arrays, and @graph until you map the data to your application schema. Flattening too early can detach a Product from its Offer, or a WebPage from its Organization.

Minimal Python extractor

import json
import requests
from bs4 import BeautifulSoup

url = "https://example.com/page"
html = requests.get(url, timeout=20).text
soup = BeautifulSoup(html, "html.parser")
records = []

for node in soup.select('script[type="application/ld+json"]'):
    raw = node.string or node.get_text()
    try:
        records.append(json.loads(raw))
    except json.JSONDecodeError:
        # Keep malformed blocks for diagnostics; do not silently discard them.
        records.append({"_parse_error": True, "raw": raw})

result = {"url": url, "jsonld": records}
print(json.dumps(result, ensure_ascii=False, indent=2))

Install the two dependencies with pip install requests beautifulsoup4. The code deliberately keeps malformed blocks and their original text. Silently dropping them makes publisher errors impossible to diagnose.

JavaScript extraction from an HTML string

const response = await fetch("https://example.com/page");
const html = await response.text();
const doc = new DOMParser().parseFromString(html, "text/html");
const blocks = [...doc.querySelectorAll('script[type="application/ld+json"]')]
  .map(node => {
    try { return JSON.parse(node.textContent); }
    catch (error) { return { _parse_error: true, raw: node.textContent }; }
  });
console.log(JSON.stringify({ url: response.url, jsonld: blocks }, null, 2));

This browser-side pattern parses the HTML you already have. It does not execute the page’s application JavaScript. For client-generated data, run the URL in a real browser and query the post-render document.

Handling JSON-LD variations

  • A block can be one object, an array of objects, or an object containing @graph.
  • Values can be strings, numbers, arrays, nested objects, or references using @id.
  • Multiple blocks may describe the same entity. Keep them separate initially and deduplicate later using stable identifiers such as @id.
  • Do not remove the @context; it determines how terms are interpreted.

3. Extract Microdata from HTML attributes

Microdata stores entities directly on elements. An item begins with itemscope, its vocabulary is identified by itemtype, and properties are marked with itemprop. When present, itemid supplies a stable identifier. Nested item scopes represent nested entities rather than ordinary text fields.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #2
Sale
HTML and CSS: Design and Build Websites
  • HTML CSS Design and Build Web Sites
  • Comes with secure packaging
  • It can be a gift option

Property values come from the element appropriate to that HTML type: text content for ordinary elements, content on a meta element, href on a link or anchor, src on an image or media element, and the relevant date or value attributes on time and data elements. Resolve relative URLs against the document URL before storing them.

A traversal strategy

  1. Find top-level elements with itemscope that are not themselves an itemprop of another scope.
  2. Create an item record containing its itemtype, optional itemid, source element, and an empty property map.
  3. Walk descendants until a nested itemscope is reached. Add that complete child item to the parent’s property, then do not treat the child’s descendants as parent properties.
  4. For each itemprop, read the value-bearing attribute or text and retain the originating element.
  5. Allow repeated properties to become arrays; never overwrite the first value.

Keeping the source element and raw value makes it possible to explain exactly where a normalized field came from when two representations disagree.

4. Extract RDFa relationships

RDFa expresses subject–predicate–object relationships through attributes such as about, typeof, property, resource, href, and src. A page can establish a subject on one element and add properties on descendants, so RDFa extraction is a relationship traversal rather than a selector that returns isolated fields.

Practical RDFa rules

  • Use about to establish or change the current subject; resolve relative references against the page URL.
  • Use typeof to record the subject’s type or types.
  • Use property for predicates. The object normally comes from content, datetime, resource, href, src, or text, depending on the element.
  • When an element has both a nested subject and a property, attach the nested resource as the parent’s object and then process the nested subject’s own properties.
  • Preserve language, datatype, and relative URL information when your consumer needs RDF-level fidelity.

For broad coverage, use an RDFa-capable library rather than attempting to reduce every document to CSS selectors. The W3C RDFa model is designed to expose a uniform document query interface by type, subject, and property.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

5. Normalize into an auditable internal shape

After format-specific passes, convert records to a common envelope while retaining the original representation:

{
  "source_url": "https://example.com/page",
  "format": "json-ld",
  "type": "https://schema.org/Product",
  "id": "https://example.com/page#product",
  "properties": { "name": ["Example item"] },
  "raw": { "@context": "https://schema.org", "@type": "Product" },
  "source_element": "script[type=application/ld+json]"
}

Use the same envelope for Microdata and RDFa, changing format and the source description. Keep arrays even when a field currently has one value. Preserve nested entities and @graph nodes until the application mapping stage.

Duplicate and conflicting representations

Many pages publish JSON-LD plus Microdata or RDFa. They can describe the same entity, or they can disagree. Match records by @id, itemid, explicit URL, or a carefully chosen combination of type and page URL. Keep every source, then apply an explicit precedence rule for your product. For example, you might prefer an identified JSON-LD node for a field but retain the HTML value as an alternate. Never discard the conflict without recording it.

6. Validate the combined result

Send either the source URL or the extracted markup to Schema.org’s Markup Validator during development. It can extract JSON-LD, RDFa, and Microdata, combine them, summarize the graph, and expose syntax mistakes. Validation should happen after all format passes, not only on JSON-LD.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #4
Sale
Web Design with HTML, CSS, JavaScript and jQuery Set
  • Brand: Wiley
  • Set of 2 Volumes
  • A handy two-book set that uniquely combines related technologies Highly visual format and accessible language makes these books highly effective learning tools Perfect for beginning web designers and front-end developers
  • Check that JSON parses and that required brackets and quotes are present.
  • Check that item scopes close correctly and that nested properties belong to the intended item.
  • Check RDFa subjects, predicates, and object references after URL resolution.
  • Compare duplicate entities and flag incompatible types or values.
  • Store validator findings alongside the page URL and retrieval timestamp.

A validator confirms markup structure; it does not prove that a publisher’s claims are factually correct or that a field is suitable for your business logic.

7. Browser-rendered extraction when JavaScript injects data

Use a browser only when a static response lacks the fields you need. In automation, navigate to the URL, wait for a meaningful selector or network-idle condition, then read document.documentElement.outerHTML and run the same JSON-LD, Microdata, and RDFa passes. If the site renders progressively, wait for the specific product, article, or organization element rather than using an arbitrary long sleep.

Rendered-page checklist

  • Set a navigation timeout and capture a diagnostic screenshot or HTML dump on failure.
  • Record redirects, cookies, locale, timezone, and user-agent because they can change structured data.
  • Wait for the selector that proves the target entity exists.
  • Capture relevant network responses if the data is fetched from an API.
  • Use the final DOM, not only the initial response body.

Some pages block automation or present consent dialogs before the content appears. Treat those as acquisition failures, not as evidence that the page has no structured data.

8. Common failure modes and fixes

Symptom Likely cause Fix
No records found You only searched for JSON-LD, or the page uses client-side rendering. Add Microdata and RDFa passes; if the initial HTML is empty, render the page and inspect the final DOM.
Only part of an entity appears @graph or nested item scopes were flattened. Keep graph nodes and nested scopes intact until normalization.
Malformed JSON-LD disappears Parser exceptions were ignored. Store a parse-error record with the original block and report it.
URLs are inconsistent Relative references were copied without a base URL. Resolve @id, itemid, about, resource, href, and src against the final document URL.
Values disagree The page publishes multiple representations. Retain provenance, define precedence explicitly, and flag conflicts for review.
Browser run sees a blank page Consent, bot protection, a timeout, or an application error interrupted acquisition. Save the response and rendered diagnostics, handle consent where permitted, and retry with bounded timeouts rather than treating the result as an empty graph.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

9. Performance, reliability, and cost decisions

Static fetching should be your default for large crawls because it avoids browser startup and JavaScript execution. Add caching keyed by URL and relevant request context, and retain the response used to produce each record. Browser rendering gives better coverage but consumes more time and memory; reserve it for pages or fields proven to require it.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

For either mode, use bounded timeouts, retry only transient failures, and distinguish HTTP errors, bot checks, blank documents, parse errors, and genuine pages with no structured data. A per-page status such as success, render_required, blocked, timeout, or parse_error makes operations measurable without conflating different problems.

Or skip the browser setup

If your only reason to run a browser is obtaining a clean page image while diagnosing or documenting an extraction target, ScreenshotNeo provides a website screenshot API and MCP server. It accepts a URL and returns PNG, JPEG, WebP, or PDF; it can load lazy images, wait for a selector or network idle, supply cookies and headers, and run custom JavaScript. Before capture it accepts cookie or consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets, with each step independently configurable.

Only clean shots are billed. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and the response identifies the result with X-Page-Verdict and X-Billed headers. Its MCP server exposes take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients.

One-call cURL example

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

See the ScreenshotNeo API documentation for the complete option set, including viewport and device presets, full-page capture, CSS selectors, dark mode, retina scale, PDF settings, blocking rules, geolocation, timezone, resizing, caching TTLs, signed links, asynchronous webhooks, bulk capture of up to 100 URLs per call, usage reporting, and the OpenAPI specification.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Python

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Plans

Plan Allowance Price
Free 1,000 shots/month $0, no card
Starter 3,000 shots $5
Growth 15,000 shots $15
Pro 60,000 shots $39
Scale 250,000 shots $99
Business 1,000,000 shots $249

Every feature is included on every plan, and yearly billing gives two months free. You can start with 1,000 free screenshots a month without a card.

10. A production-ready checklist

  • Acquire the final document and record its retrieval context.
  • Parse every JSON-LD block, including arrays and @graph.
  • Traverse Microdata item scopes and value-bearing attributes.
  • Traverse RDFa subjects, predicates, and resources.
  • Resolve URLs and preserve raw fragments and source elements.
  • Deduplicate only after identifiers and provenance are retained.
  • Define conflict precedence instead of trusting one representation.
  • Render with a browser only when JavaScript is required.
  • Validate the combined graph and store diagnostics.

Frequently Asked Questions

Can a page expose structured data in more than one format?

Yes. JSON-LD, Microdata, and RDFa can coexist. Extract each independently, retain provenance, and reconcile duplicates explicitly.

Should I convert JSON-LD into a flat dictionary immediately?

No. Preserve @context, @id, arrays, nested objects, and @graph until your application mapping stage so entity relationships are not lost.

How do I know whether browser rendering is necessary?

Compare the fields in the initial HTTP response with the fields present after JavaScript runs. Render only when the required markup is injected or revealed client-side.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Read next

Recommended PC Tool
Recommended PC Tool
Crashes, No Sound, or Screen Glitches?Free driver scan
Windows Errors? Fix Them Before They SpreadFree repair scan

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.