Driver FixRecommendedSound, Wi-Fi or graphics acting up? Check drivers firstFind missing or outdated drivers fast.Check DriversOctober DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsWindows FixRecommendedWindows errors stealing your time? Find the fix fastScan stability, cleanup and performance issues.Fix Now×
Skip to content
APIs

How to Extract Structured JSON Data from Websites: APIs, JSON-LD, Network Calls, and DOM Fallbacks

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The most reliable way to extract structured JSON from a website is to use its official API. If no suitable API exists, download the HTML and parse embedded JSON or JSON-LD; for JavaScript-rendered pages, observe the browser’s network responses and replay the data endpoint when permitted. Use DOM extraction only as a fallback, then validate fields and preserve provenance so every record can be audited.

Choose the extraction path before writing a scraper

Different sources expose different contracts. Decide in this order:

  1. Official API: Best for stable field names, authentication, pagination, status codes, and documented limits.
  2. Embedded data: Inspect the initial HTML for ordinary JSON in script tags, <script type="application/ld+json"> blocks, Schema.org Microdata, or RDFa.
  3. Browser network data: For JavaScript applications, observe XHR and fetch responses to find the JSON payload used to render the page.
  4. DOM fallback: Select semantic elements only when no usable API or payload exists.

Do not assume that data visible on screen is present in the first HTML response. Conversely, do not launch a browser when a documented endpoint already returns the required records.

Start with an official API

Read the API documentation and record the endpoint version, authentication method, required fields, pagination parameters, rate limits, and error statuses. Treat the response as a contract rather than scraping whatever happens to be displayed.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Minimal Python request

import requests

url = "https://example.com/api/items"
headers = {"Authorization": "Bearer YOUR_TOKEN"}
params = {"page": 1, "limit": 100}

response = requests.get(url, headers=headers, params=params, timeout=30)
response.raise_for_status()
data = response.json()
print(data)

Check whether the API returns an object containing a records array, a bare array, or a pagination envelope. Continue until the documented next-page marker is exhausted, and deduplicate records using a stable identifier.

Validate the contract

  • Confirm the HTTP status and final URL after redirects.
  • Distinguish a missing property from an explicit null or empty array.
  • Validate required fields, types, date formats, and locale-specific numbers.
  • Log non-2xx responses and retain enough context to reproduce the request.

Extract JSON embedded in HTML

Many sites publish data in script elements so search engines and other clients can consume it. JSON-LD is a JSON-based format for Linked Data. Schema.org’s vocabulary supplies machine-readable types and properties and can be used with JSON-LD, Microdata, or RDFa.

Parse JSON-LD blocks with Python

import json
import requests
from bs4 import BeautifulSoup

page_url = "https://example.com/article"
r = requests.get(page_url, timeout=30,
                 headers={"User-Agent": "structured-data-client/1.0"})
r.raise_for_status()
soup = BeautifulSoup(r.text, "html.parser")

records = []
for node in soup.select('script[type="application/ld+json"]'):
    raw = node.string or node.get_text()
    try:
        value = json.loads(raw)
    except json.JSONDecodeError as exc:
        print(f"Skipping malformed JSON-LD: {exc}")
        continue
    if isinstance(value, list):
        records.extend(value)
    else:
        records.append(value)

for record in records:
    print(record)

A block may contain an object, an array, or an @graph array. Keep the original structure until you decide whether to flatten it. When linked-data semantics matter, use a JSON-LD 1.1 processor for expansion or compaction instead of treating every property as an ordinary scalar.

Normalize without discarding information

Map source types and properties into your application schema only after parsing. Retain unknown properties during this stage; dropping them early can silently remove fields added by the publisher. Store the source URL, retrieval time, extraction method, and a hash of the raw payload alongside normalized records.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Find JSON behind a JavaScript-rendered page

Client-side applications commonly fetch JSON after the initial document loads. A browser automation library can observe those requests. Playwright exposes request, response, requestfinished, and requestfailed events, which let you identify the response containing the desired record.

Capture response bodies with Playwright

from playwright.sync_api import sync_playwright

page_url = "https://example.com/catalog"

with sync_playwright() as p:
    browser = p.chromium.launch(headless=True)
    page = browser.new_page()
    candidates = []

    def inspect(response):
        content_type = response.headers.get("content-type", "")
        if "json" in content_type.lower():
            candidates.append(response)

    page.on("response", inspect)
    page.goto(page_url, wait_until="networkidle", timeout=60_000)

    for response in candidates:
        try:
            body = response.json()
        except Exception:
            continue
        print(response.url, body)
    browser.close()

Filter candidates by URL, HTTP method, response content type, or a distinctive field. Once you identify the endpoint, replay it directly if the site permits and the endpoint is stable. Direct replay is usually faster and less brittle than repeatedly scraping rendered text, but private endpoints can change without notice and may require browser cookies, CSRF tokens, or signed parameters.

When browser replay is not appropriate

  • The endpoint requires a short-lived token generated in the page.
  • Authentication or consent state is represented only by browser storage.
  • The site’s access rules prohibit automated requests.
  • The response is assembled from several calls and cannot be reconstructed safely.

In these cases, keep extraction inside the browser context, use permitted credentials, and limit request frequency.

Use DOM extraction only as a fallback

If no API or structured payload is available, select semantic elements and normalize their values. Prefer stable attributes such as data-id, itemprop, or accessible labels over positional selectors and generated class names.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
from bs4 import BeautifulSoup
from decimal import Decimal
import requests

url = "https://example.com/products"
r = requests.get(url, timeout=30)
r.raise_for_status()
soup = BeautifulSoup(r.text, "html.parser")

items = []
for card in soup.select("article.product-card"):
    name = card.select_one("[itemprop='name']")
    price = card.select_one("[itemprop='price']")
    if not name or not price:
        continue
    items.append({
        "name": " ".join(name.get_text(" ", strip=True).split()),
        "price": Decimal(price.get("content") or price.get_text(strip=True).replace(",", ""))
    })
print(items)

Record the selector set and add regression fixtures for representative pages. Presentation markup changes more often than a documented API, so a fixture-based test should fail loudly when a selector stops matching.

Handle JSON-LD, Microdata, and RDFa deliberately

JSON-LD

Process every application/ld+json block independently. Support objects, arrays, and @graph; preserve @id, @type, and context information until normalization is complete.

Microdata

Microdata expresses properties in HTML attributes such as itemscope, itemtype, and itemprop. Nested scopes can represent related entities, so build a tree or graph rather than collecting only text nodes.

RDFa

RDFa uses attributes such as typeof, property, and resource. Values may come from an element’s text, a URL attribute, or a content attribute. Normalize these sources according to the property’s expected type.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Schema.org terms are a vocabulary, not a guarantee that every publisher supplies every property. Treat absent fields as absent; do not manufacture defaults that look like source data.

Normalize and validate the output

Define an application schema before exporting records. A useful normalized record often includes a stable ID, source URL, retrieval timestamp, extraction method, and the domain fields your application needs.

from datetime import datetime, timezone

normalized = {
    "id": source.get("@id") or source.get("id"),
    "type": source.get("@type"),
    "name": source.get("name"),
    "source_url": page_url,
    "retrieved_at": datetime.now(timezone.utc).isoformat(),
    "method": "json-ld"
}

required = ("id", "name")
missing = [key for key in required if not normalized.get(key)]
if missing:
    raise ValueError(f"Missing required fields: {missing}")

Validate dates, numeric ranges, enumerated types, and locale-specific formats before writing JSON. Keep the raw response or a content-addressed copy when policy permits, and store its hash so a later audit can verify what was parsed.

Pagination, duplicates, and provenance

  • Follow every documented page, cursor, or continuation token.
  • Stop only when the API signals completion; a short page is not always the last page.
  • Deduplicate by a stable source identifier, not by display name.
  • Store retrieval time, URL, request parameters (excluding secrets), selector or endpoint, and parser version.
  • Record parse failures with the response status and a bounded sample of the offending payload.

Performance, reliability, and access rules

Prefer direct API or endpoint requests over full browser sessions because they use less CPU and memory. Reuse HTTP connections, set explicit timeouts, and apply bounded retries only to transient failures such as connection resets or 5xx responses. Do not blindly retry authentication errors, validation errors, or 404s.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Limit concurrency to what the site’s terms and rate limits allow. Cache immutable pages and use conditional requests where supported. A browser should wait for a meaningful condition—such as a response containing the required field—rather than an arbitrary long delay. Always identify your client honestly and avoid bypassing CAPTCHAs, bot checks, paywalls, or access controls.

Common failures and fixes

Symptom Likely cause Fix
JSON parser reports an unexpected character You downloaded an HTML error page or truncated response Check status, final URL, content type, and response length before parsing.
No JSON-LD blocks found Data is injected after load or uses Microdata/RDFa Inspect browser responses and scan semantic attributes.
Playwright sees no useful response Requests occur after an interaction or require authentication Perform the permitted click/login flow and capture request failures as well as responses.
Records are incomplete Pagination stopped early or a nested @graph was ignored Follow continuation tokens and explicitly process graph members.
Numbers or dates are wrong Locale formatting or timezone conversion Parse with the source locale and store timezone-aware timestamps.
DOM scraper suddenly returns empty fields Presentation markup changed Update selectors from a fixture diff and prefer semantic attributes.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup

ScreenshotNeo can provide a clean page capture when your workflow needs a rendered page rather than a custom scraper. It accepts cookie and consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be disabled. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers report the page verdict and billing status. Its MCP server includes take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients.

One request returns PNG, JPEG, WebP, or PDF. See the ScreenshotNeo API documentation for all options.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Plans include 1,000 screenshots per month free with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account to try it.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Which method should you choose?

Method Contract stability Rendered-content coverage Runtime cost Main risk
Official API Highest when versioned and documented Data exposed by the API Lowest Authentication, quotas, or missing fields
Embedded JSON-LD Moderate Publisher-selected metadata Low Multiple blocks, graphs, or stale markup
Network replay Moderate to low Usually broad Medium Private endpoints and changing tokens
DOM parsing Lowest Visible semantic content Low to medium Markup and selector changes
Full browser automation Depends on page behavior Highest for permitted flows Highest Timing, memory, authentication, and access controls

Use the least complex method that supplies the fields you can validate. Escalate from API to embedded data, network observation, and finally DOM or browser automation only when the preceding option cannot meet the requirement.

Frequently Asked Questions

Should I parse JSON-LD or call an API?

Call the official API when it provides the fields you need; use JSON-LD when the page publishes useful metadata but no suitable API is available.

Can I scrape a site’s private JSON endpoint?

Only when your access is authorized and the endpoint’s terms permit automated use. Private endpoints can change without notice and may require browser state.

How do I know whether JSON is complete?

Check pagination markers, required fields, duplicate IDs, HTTP status, and parser errors, then retain retrieval metadata and a raw-payload hash.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Why does a page show data that is absent from its HTML?

JavaScript may fetch it after load. Observe Playwright request and response events to locate the JSON response, then replay it only when permitted.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Read next

Recommended PC Tool
Recommended PC Tool
Outdated Drivers Are Slowing You DownFree scan - exact matches
Windows Errors? Fix Them Before They SpreadFree repair scan

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.