Quick wins for a faster PC:
Scan for outdated or missing drivers - takes under a minuteDriver Scan →Repair Windows errors before they cause bigger problemsFix Now →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
Reliable web extraction starts with a response contract, not a prompt. Define the exact fields, types, requiredness, null rules, evidence, and pagination your application needs; then choose selectors for stable page structures or schema-guided extraction for variable content. Finally, validate both the JSON shape and the source facts before data reaches downstream systems.
Start with a response contract
A response contract is the agreement between an extraction API and the software consuming its result. Write it before choosing a provider or model.
Define names, types, and requiredness
- Use stable, descriptive field names such as
title,price,currency, andsource_url. - Declare primitive types explicitly: string, integer, number, boolean, or a date/time representation.
- Mark fields as required only when the application cannot operate without them. Decide whether an unavailable value is represented by
null, an empty array, or an omitted optional property. - Model repeated data as arrays and related data as nested objects rather than encoded strings.
- Add descriptions for ambiguous fields, units, allowed values, and formatting rules.
Where the API supports strict JSON Schema, require the properties you need and reject accidental keys with additionalProperties: false. OpenAI’s structured-output examples demonstrate this pattern, while Cloudflare describes a schema as the expected output structure. See OpenAI’s structured outputs guide and Cloudflare’s JSON endpoint documentation.
Example contract for a product page
{
"type": "object",
"properties": {
"name": {"type": "string"},
"price": {"type": "number"},
"currency": {"type": "string"},
"availability": {"type": ["string", "null"]},
"features": {"type": "array", "items": {"type": "string"}},
"source_url": {"type": "string", "format": "uri"}
},
"required": ["name", "price", "currency", "features", "source_url"],
"additionalProperties": false
}
This contract makes an absent availability value explicit while preventing a provider from silently adding fields your consumer does not understand.
#1 Best Overall
Choose selectors or semantic extraction
The extraction method should follow how predictable the source is. CSS-rule extraction and prompt- or schema-guided extraction solve different problems and fail differently.
Use CSS selectors for a known structure
Selectors are deterministic when the target fields consistently occupy known DOM elements: for example, h1.product-name for a title or [data-price] for a price. They are fast and easy to test against fixtures, but a redesign, changed class name, or different template can make them return empty or incorrect values. Context.dev distinguishes a CSS-rule Scrape endpoint from its research-oriented Answers endpoint and warns that selectors may need updates when a site changes: Context.dev’s data extraction API documentation.
Use prompts and schemas for variable content
Semantic extraction is useful when the same concept appears in different wording or locations, or when the task requires interpretation across sources. Cloudflare’s endpoint accepts a prompt, a JSON Schema response_format, or both, and returns extracted data as JSON. A prompt states what to look for; the schema constrains the returned shape. A schema alone does not prove that the page supports every value.
The Tool Desk
Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Do not confuse an example JSON object with JSON Schema. Context.dev describes its json_format as an example shape and instructs applications to validate the resulting json_content; it is not equivalent to a standards-based schema with type and required-property enforcement.
Decision table
| Condition | Preferred method | Main failure mode |
|---|---|---|
| Stable DOM and known fields | CSS selectors or provider scrape rules | Breakage after a template or class-name change |
| Variable wording or layouts | Prompt plus JSON Schema | Misinterpretation or unsupported values despite valid JSON |
| Research across multiple sources | Semantic/research endpoint with source URLs | Incomplete coverage or weak attribution |
| Regulated or auditable workflow | Either method plus evidence fields and validation | Structurally valid values that lack source support |
Make evidence part of the returned data
For decisions, compliance, or editorial review, return traceability alongside extracted values. Include the page URL, retrieval timestamp, and—where the provider supplies it—source snippets or citations. Cloudflare supports extraction from a URL or supplied HTML and structured JSON output. Context.dev’s research-oriented documentation describes source URLs for its Answers endpoint. Store that evidence at the application layer when the API does not include it automatically.
A useful pattern is an object containing value, source_url, and evidence. Keep evidence separate from the normalized value so downstream code does not have to parse prose.
Render the page before extracting
Valid JSON can still contain no useful data if the browser captured the page before JavaScript rendered it. Cloudflare warns that JavaScript-heavy pages may be read before scripts finish and recommends waiting for networkidle0, networkidle2, or a known content selector: Cloudflare’s rendering guidance.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
- Identify a selector that appears only after the required content is rendered, such as a results container.
- Prefer a network-idle wait when the page loads data through several requests; use a selector wait when background polling never becomes idle.
- Set a timeout appropriate to the site and record timeout versus empty-result outcomes separately.
- Retry transient navigation failures with a bounded policy. Do not retry a deterministic selector miss forever.
- Remember that a configurable user agent does not bypass bot protection, according to Cloudflare’s documentation.
When a result is empty, distinguish at least four cases: the page genuinely has no matching records, the selector or prompt was wrong, rendering had not completed, or access was blocked. Preserve that status in logs rather than converting every case to an empty array.
Rank #3
Validate shape, values, and provenance
Schema validation is the first gate, not the final one. After parsing the response:
- Check required properties and reject unknown keys when strictness is part of the contract.
- Verify numeric ranges, currency codes, date formats, URLs, and enumerated values.
- Check that arrays contain the expected item type and that nested objects are complete.
- Apply explicit missing-value rules: distinguish
null, an empty list, and an omitted optional field. - Confirm that each important value has page evidence or a source URL.
- Flag contradictions, stale timestamps, and values that cannot be supported by the captured content.
Context.dev explicitly places validation of json_content in the client application. Treat structural conformance and factual verification as separate checks.
Consume result, error, and pagination fields deliberately
Many extraction workflows fail after the extraction itself because the client reads the wrong response path or silently stops at the first page. Map the success payload and documented error payload in configuration. AWS Glue’s connection documentation illustrates configurable result and error paths, while its pagination settings cover cursor- and offset-based APIs: AWS Glue Connection Type API.
Free tools Windows power users keep installed
One-click scans. No signup required.
Cursor pagination
Read the provider’s next-cursor field, send it unchanged on the next request, and stop when it is absent or null. Protect against a cursor that repeats.
Offset or page-number pagination
Track the current offset or page and the provider’s page size. Continue until the returned count is below the requested limit or the API signals that no next page exists. Record the effective limit because providers may cap it.
ScrAPIr notes that a client without pagination details may retrieve only a default first page. Its historical evaluation found a longest-text heuristic surfaced a human-readable error message 87.5% of the time, with a 95% confidence interval of ±14.78%, across 40 randomly selected APIs; that small 2017 result is about one error-parsing heuristic, not general API reliability: the ScrAPIr paper.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Account for provider and model limits
Feature names are not portable across endpoints. Amazon Bedrock documents structured outputs across several APIs but states that the Anthropic Messages API on bedrock-mantle does not support the format parameter. Bedrock also documents a citation incompatibility for Anthropic structured outputs. Check support for the exact model and API endpoint you deploy: Amazon Bedrock structured-output documentation.
Windows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallOutdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchKeep provider capability flags in configuration rather than assuming that a schema, citations, browser rendering, or pagination works identically everywhere.
Best Value
A practical implementation sequence
- Specify the consumer. Write down which service, database, queue, or analyst will read each field.
- Draft and version the schema. Define types, required properties, null semantics, arrays, nested objects, and evidence fields.
- Classify the source. Choose selectors for stable templates; choose semantic extraction when layout or wording varies.
- Configure rendering. Wait for network idle or a known selector, set a timeout, and capture blocked or failed states.
- Request the data. Supply the URL or HTML, prompt, and response format supported by the chosen endpoint.
- Validate twice. Run a JSON Schema validator, then apply domain and provenance checks.
- Follow every page. Implement the documented cursor or offset model and record limits and counts.
- Monitor drift. Alert on rising empty results, selector misses, schema failures, unsupported citations, and pagination anomalies.
Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server that can provide a rendered visual or PDF for a page before you run your own extraction checks. It accepts consent banners as a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each cleanup step can be disabled. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing status. Its MCP server exposes take_screenshot, get_page_info, and capture_pdf to Claude, Cursor, and other MCP clients.
One GET request returns PNG, JPEG, WebP, or PDF. See the ScreenshotNeo documentation for all options.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
The free plan includes 1,000 screenshots per month with no card. Paid plans start at $5 for 3,000 shots; every feature is included on every plan. Sign up for ScreenshotNeo.
Frequently Asked Questions
How do I define what data to extract?
Define a versioned response contract: field names, types, required and optional properties, null behavior, arrays, nested objects, and evidence fields. Then express semantic goals in a prompt and enforce the shape with JSON Schema where supported.
How do I extract lists and attributes?
Use CSS selectors when list items and attributes have stable DOM markers. For variable layouts, request an array in the schema, describe each item and attribute, wait for rendered content, and validate every item after extraction.
How do I get typed JSON from a webpage?
Provide a prompt, a JSON Schema response format, or both to an endpoint that supports them, then run client-side schema and domain validation. Typed output controls structure; it does not verify the page’s facts.
Why is my result empty on a JavaScript-heavy page?
The capture may have occurred before scripts finished, the selector may be wrong, or access may have been blocked. Wait for network idle or a known content selector, classify timeout and blocked states separately, and check the provider’s rendering limits.
Recommended Free Tools
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

