Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minuteWindows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallThe most reliable way to extract structured JSON from a website is to use its official API. If no suitable API exists, download the HTML and parse embedded JSON or JSON-LD; for JavaScript-rendered pages, observe the browser’s network responses and replay the data endpoint when permitted. Use DOM extraction only as a fallback, then validate fields and preserve provenance so every record can be audited.
Choose the extraction path before writing a scraper
Different sources expose different contracts. Decide in this order:
- Official API: Best for stable field names, authentication, pagination, status codes, and documented limits.
- Embedded data: Inspect the initial HTML for ordinary JSON in script tags,
<script type="application/ld+json">blocks, Schema.org Microdata, or RDFa. - Browser network data: For JavaScript applications, observe XHR and fetch responses to find the JSON payload used to render the page.
- DOM fallback: Select semantic elements only when no usable API or payload exists.
Do not assume that data visible on screen is present in the first HTML response. Conversely, do not launch a browser when a documented endpoint already returns the required records.
Start with an official API
Read the API documentation and record the endpoint version, authentication method, required fields, pagination parameters, rate limits, and error statuses. Treat the response as a contract rather than scraping whatever happens to be displayed.
#1 Best Overall
Minimal Python request
import requests
url = "https://example.com/api/items"
headers = {"Authorization": "Bearer YOUR_TOKEN"}
params = {"page": 1, "limit": 100}
response = requests.get(url, headers=headers, params=params, timeout=30)
response.raise_for_status()
data = response.json()
print(data)
Check whether the API returns an object containing a records array, a bare array, or a pagination envelope. Continue until the documented next-page marker is exhausted, and deduplicate records using a stable identifier.
Validate the contract
- Confirm the HTTP status and final URL after redirects.
- Distinguish a missing property from an explicit
nullor empty array. - Validate required fields, types, date formats, and locale-specific numbers.
- Log non-2xx responses and retain enough context to reproduce the request.
Extract JSON embedded in HTML
Many sites publish data in script elements so search engines and other clients can consume it. JSON-LD is a JSON-based format for Linked Data. Schema.org’s vocabulary supplies machine-readable types and properties and can be used with JSON-LD, Microdata, or RDFa.
Parse JSON-LD blocks with Python
import json
import requests
from bs4 import BeautifulSoup
page_url = "https://example.com/article"
r = requests.get(page_url, timeout=30,
headers={"User-Agent": "structured-data-client/1.0"})
r.raise_for_status()
soup = BeautifulSoup(r.text, "html.parser")
records = []
for node in soup.select('script[type="application/ld+json"]'):
raw = node.string or node.get_text()
try:
value = json.loads(raw)
except json.JSONDecodeError as exc:
print(f"Skipping malformed JSON-LD: {exc}")
continue
if isinstance(value, list):
records.extend(value)
else:
records.append(value)
for record in records:
print(record)
A block may contain an object, an array, or an @graph array. Keep the original structure until you decide whether to flatten it. When linked-data semantics matter, use a JSON-LD 1.1 processor for expansion or compaction instead of treating every property as an ordinary scalar.
Normalize without discarding information
Map source types and properties into your application schema only after parsing. Retain unknown properties during this stage; dropping them early can silently remove fields added by the publisher. Store the source URL, retrieval time, extraction method, and a hash of the raw payload alongside normalized records.
Do these 3 things before closing this tab:
1Scan for outdated or missing drivers - takes under a minute2Repair Windows errors before they cause bigger problems3Fix the driver behind crashes, sound loss and screen glitchesFind JSON behind a JavaScript-rendered page
Client-side applications commonly fetch JSON after the initial document loads. A browser automation library can observe those requests. Playwright exposes request, response, requestfinished, and requestfailed events, which let you identify the response containing the desired record.
Capture response bodies with Playwright
from playwright.sync_api import sync_playwright
page_url = "https://example.com/catalog"
with sync_playwright() as p:
browser = p.chromium.launch(headless=True)
page = browser.new_page()
candidates = []
def inspect(response):
content_type = response.headers.get("content-type", "")
if "json" in content_type.lower():
candidates.append(response)
page.on("response", inspect)
page.goto(page_url, wait_until="networkidle", timeout=60_000)
for response in candidates:
try:
body = response.json()
except Exception:
continue
print(response.url, body)
browser.close()
Filter candidates by URL, HTTP method, response content type, or a distinctive field. Once you identify the endpoint, replay it directly if the site permits and the endpoint is stable. Direct replay is usually faster and less brittle than repeatedly scraping rendered text, but private endpoints can change without notice and may require browser cookies, CSRF tokens, or signed parameters.
When browser replay is not appropriate
- The endpoint requires a short-lived token generated in the page.
- Authentication or consent state is represented only by browser storage.
- The site’s access rules prohibit automated requests.
- The response is assembled from several calls and cannot be reconstructed safely.
In these cases, keep extraction inside the browser context, use permitted credentials, and limit request frequency.
Use DOM extraction only as a fallback
If no API or structured payload is available, select semantic elements and normalize their values. Prefer stable attributes such as data-id, itemprop, or accessible labels over positional selectors and generated class names.
from bs4 import BeautifulSoup
from decimal import Decimal
import requests
url = "https://example.com/products"
r = requests.get(url, timeout=30)
r.raise_for_status()
soup = BeautifulSoup(r.text, "html.parser")
items = []
for card in soup.select("article.product-card"):
name = card.select_one("[itemprop='name']")
price = card.select_one("[itemprop='price']")
if not name or not price:
continue
items.append({
"name": " ".join(name.get_text(" ", strip=True).split()),
"price": Decimal(price.get("content") or price.get_text(strip=True).replace(",", ""))
})
print(items)
Record the selector set and add regression fixtures for representative pages. Presentation markup changes more often than a documented API, so a fixture-based test should fail loudly when a selector stops matching.
Handle JSON-LD, Microdata, and RDFa deliberately
JSON-LD
Process every application/ld+json block independently. Support objects, arrays, and @graph; preserve @id, @type, and context information until normalization is complete.
Rank #3
Microdata
Microdata expresses properties in HTML attributes such as itemscope, itemtype, and itemprop. Nested scopes can represent related entities, so build a tree or graph rather than collecting only text nodes.
RDFa
RDFa uses attributes such as typeof, property, and resource. Values may come from an element’s text, a URL attribute, or a content attribute. Normalize these sources according to the property’s expected type.
Recommended Free Tools
Schema.org terms are a vocabulary, not a guarantee that every publisher supplies every property. Treat absent fields as absent; do not manufacture defaults that look like source data.
Normalize and validate the output
Define an application schema before exporting records. A useful normalized record often includes a stable ID, source URL, retrieval timestamp, extraction method, and the domain fields your application needs.
from datetime import datetime, timezone
normalized = {
"id": source.get("@id") or source.get("id"),
"type": source.get("@type"),
"name": source.get("name"),
"source_url": page_url,
"retrieved_at": datetime.now(timezone.utc).isoformat(),
"method": "json-ld"
}
required = ("id", "name")
missing = [key for key in required if not normalized.get(key)]
if missing:
raise ValueError(f"Missing required fields: {missing}")
Validate dates, numeric ranges, enumerated types, and locale-specific formats before writing JSON. Keep the raw response or a content-addressed copy when policy permits, and store its hash so a later audit can verify what was parsed.
Pagination, duplicates, and provenance
- Follow every documented page, cursor, or continuation token.
- Stop only when the API signals completion; a short page is not always the last page.
- Deduplicate by a stable source identifier, not by display name.
- Store retrieval time, URL, request parameters (excluding secrets), selector or endpoint, and parser version.
- Record parse failures with the response status and a bounded sample of the offending payload.
Performance, reliability, and access rules
Prefer direct API or endpoint requests over full browser sessions because they use less CPU and memory. Reuse HTTP connections, set explicit timeouts, and apply bounded retries only to transient failures such as connection resets or 5xx responses. Do not blindly retry authentication errors, validation errors, or 404s.
Limit concurrency to what the site’s terms and rate limits allow. Cache immutable pages and use conditional requests where supported. A browser should wait for a meaningful condition—such as a response containing the required field—rather than an arbitrary long delay. Always identify your client honestly and avoid bypassing CAPTCHAs, bot checks, paywalls, or access controls.
Common failures and fixes
| Symptom | Likely cause | Fix |
|---|---|---|
| JSON parser reports an unexpected character | You downloaded an HTML error page or truncated response | Check status, final URL, content type, and response length before parsing. |
| No JSON-LD blocks found | Data is injected after load or uses Microdata/RDFa | Inspect browser responses and scan semantic attributes. |
| Playwright sees no useful response | Requests occur after an interaction or require authentication | Perform the permitted click/login flow and capture request failures as well as responses. |
| Records are incomplete | Pagination stopped early or a nested @graph was ignored |
Follow continuation tokens and explicitly process graph members. |
| Numbers or dates are wrong | Locale formatting or timezone conversion | Parse with the source locale and store timezone-aware timestamps. |
| DOM scraper suddenly returns empty fields | Presentation markup changed | Update selectors from a fixture diff and prefer semantic attributes. |
Or skip the browser setup
ScreenshotNeo can provide a clean page capture when your workflow needs a rendered page rather than a custom scraper. It accepts cookie and consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be disabled. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers report the page verdict and billing status. Its MCP server includes take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients.
One request returns PNG, JPEG, WebP, or PDF. See the ScreenshotNeo API documentation for all options.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
Plans include 1,000 screenshots per month free with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account to try it.
Which method should you choose?
| Method | Contract stability | Rendered-content coverage | Runtime cost | Main risk |
|---|---|---|---|---|
| Official API | Highest when versioned and documented | Data exposed by the API | Lowest | Authentication, quotas, or missing fields |
| Embedded JSON-LD | Moderate | Publisher-selected metadata | Low | Multiple blocks, graphs, or stale markup |
| Network replay | Moderate to low | Usually broad | Medium | Private endpoints and changing tokens |
| DOM parsing | Lowest | Visible semantic content | Low to medium | Markup and selector changes |
| Full browser automation | Depends on page behavior | Highest for permitted flows | Highest | Timing, memory, authentication, and access controls |
Use the least complex method that supplies the fields you can validate. Escalate from API to embedded data, network observation, and finally DOM or browser automation only when the preceding option cannot meet the requirement.
Frequently Asked Questions
Should I parse JSON-LD or call an API?
Call the official API when it provides the fields you need; use JSON-LD when the page publishes useful metadata but no suitable API is available.
Can I scrape a site’s private JSON endpoint?
Only when your access is authorized and the endpoint’s terms permit automated use. Private endpoints can change without notice and may require browser state.
How do I know whether JSON is complete?
Check pagination markers, required fields, duplicate IDs, HTTP status, and parser errors, then retain retrieval metadata and a raw-payload hash.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Why does a page show data that is absent from its HTML?
JavaScript may fetch it after load. Observe Playwright request and response events to locate the JSON response, then replay it only when permitted.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




