Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

Reliable web extraction starts with a response contract, not a prompt. Define the exact fields, types, requiredness, null rules, evidence, and pagination your application needs; then choose selectors for stable page structures or schema-guided extraction for variable content. Finally, validate both the JSON shape and the source facts before data reaches downstream systems.

Start with a response contract

A response contract is the agreement between an extraction API and the software consuming its result. Write it before choosing a provider or model.

Define names, types, and requiredness

  • Use stable, descriptive field names such as title, price, currency, and source_url.
  • Declare primitive types explicitly: string, integer, number, boolean, or a date/time representation.
  • Mark fields as required only when the application cannot operate without them. Decide whether an unavailable value is represented by null, an empty array, or an omitted optional property.
  • Model repeated data as arrays and related data as nested objects rather than encoded strings.
  • Add descriptions for ambiguous fields, units, allowed values, and formatting rules.

Where the API supports strict JSON Schema, require the properties you need and reject accidental keys with additionalProperties: false. OpenAI’s structured-output examples demonstrate this pattern, while Cloudflare describes a schema as the expected output structure. See OpenAI’s structured outputs guide and Cloudflare’s JSON endpoint documentation.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Example contract for a product page

{
  "type": "object",
  "properties": {
    "name": {"type": "string"},
    "price": {"type": "number"},
    "currency": {"type": "string"},
    "availability": {"type": ["string", "null"]},
    "features": {"type": "array", "items": {"type": "string"}},
    "source_url": {"type": "string", "format": "uri"}
  },
  "required": ["name", "price", "currency", "features", "source_url"],
  "additionalProperties": false
}

This contract makes an absent availability value explicit while preventing a provider from silently adding fields your consumer does not understand.

Choose selectors or semantic extraction

The extraction method should follow how predictable the source is. CSS-rule extraction and prompt- or schema-guided extraction solve different problems and fail differently.

Use CSS selectors for a known structure

Selectors are deterministic when the target fields consistently occupy known DOM elements: for example, h1.product-name for a title or [data-price] for a price. They are fast and easy to test against fixtures, but a redesign, changed class name, or different template can make them return empty or incorrect values. Context.dev distinguishes a CSS-rule Scrape endpoint from its research-oriented Answers endpoint and warns that selectors may need updates when a site changes: Context.dev’s data extraction API documentation.

Use prompts and schemas for variable content

Semantic extraction is useful when the same concept appears in different wording or locations, or when the task requires interpretation across sources. Cloudflare’s endpoint accepts a prompt, a JSON Schema response_format, or both, and returns extracted data as JSON. A prompt states what to look for; the schema constrains the returned shape. A schema alone does not prove that the page supports every value.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Do not confuse an example JSON object with JSON Schema. Context.dev describes its json_format as an example shape and instructs applications to validate the resulting json_content; it is not equivalent to a standards-based schema with type and required-property enforcement.

Decision table

Condition Preferred method Main failure mode
Stable DOM and known fields CSS selectors or provider scrape rules Breakage after a template or class-name change
Variable wording or layouts Prompt plus JSON Schema Misinterpretation or unsupported values despite valid JSON
Research across multiple sources Semantic/research endpoint with source URLs Incomplete coverage or weak attribution
Regulated or auditable workflow Either method plus evidence fields and validation Structurally valid values that lack source support

Make evidence part of the returned data

For decisions, compliance, or editorial review, return traceability alongside extracted values. Include the page URL, retrieval timestamp, and—where the provider supplies it—source snippets or citations. Cloudflare supports extraction from a URL or supplied HTML and structured JSON output. Context.dev’s research-oriented documentation describes source URLs for its Answers endpoint. Store that evidence at the application layer when the API does not include it automatically.

A useful pattern is an object containing value, source_url, and evidence. Keep evidence separate from the normalized value so downstream code does not have to parse prose.

Render the page before extracting

Valid JSON can still contain no useful data if the browser captured the page before JavaScript rendered it. Cloudflare warns that JavaScript-heavy pages may be read before scripts finish and recommends waiting for networkidle0, networkidle2, or a known content selector: Cloudflare’s rendering guidance.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  1. Identify a selector that appears only after the required content is rendered, such as a results container.
  2. Prefer a network-idle wait when the page loads data through several requests; use a selector wait when background polling never becomes idle.
  3. Set a timeout appropriate to the site and record timeout versus empty-result outcomes separately.
  4. Retry transient navigation failures with a bounded policy. Do not retry a deterministic selector miss forever.
  5. Remember that a configurable user agent does not bypass bot protection, according to Cloudflare’s documentation.

When a result is empty, distinguish at least four cases: the page genuinely has no matching records, the selector or prompt was wrong, rendering had not completed, or access was blocked. Preserve that status in logs rather than converting every case to an empty array.

Validate shape, values, and provenance

Schema validation is the first gate, not the final one. After parsing the response:

  • Check required properties and reject unknown keys when strictness is part of the contract.
  • Verify numeric ranges, currency codes, date formats, URLs, and enumerated values.
  • Check that arrays contain the expected item type and that nested objects are complete.
  • Apply explicit missing-value rules: distinguish null, an empty list, and an omitted optional field.
  • Confirm that each important value has page evidence or a source URL.
  • Flag contradictions, stale timestamps, and values that cannot be supported by the captured content.

Context.dev explicitly places validation of json_content in the client application. Treat structural conformance and factual verification as separate checks.

Consume result, error, and pagination fields deliberately

Many extraction workflows fail after the extraction itself because the client reads the wrong response path or silently stops at the first page. Map the success payload and documented error payload in configuration. AWS Glue’s connection documentation illustrates configurable result and error paths, while its pagination settings cover cursor- and offset-based APIs: AWS Glue Connection Type API.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Cursor pagination

Read the provider’s next-cursor field, send it unchanged on the next request, and stop when it is absent or null. Protect against a cursor that repeats.

Offset or page-number pagination

Track the current offset or page and the provider’s page size. Continue until the returned count is below the requested limit or the API signals that no next page exists. Record the effective limit because providers may cap it.

ScrAPIr notes that a client without pagination details may retrieve only a default first page. Its historical evaluation found a longest-text heuristic surfaced a human-readable error message 87.5% of the time, with a 95% confidence interval of ±14.78%, across 40 randomly selected APIs; that small 2017 result is about one error-parsing heuristic, not general API reliability: the ScrAPIr paper.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Account for provider and model limits

Feature names are not portable across endpoints. Amazon Bedrock documents structured outputs across several APIs but states that the Anthropic Messages API on bedrock-mantle does not support the format parameter. Bedrock also documents a citation incompatibility for Anthropic structured outputs. Check support for the exact model and API endpoint you deploy: Amazon Bedrock structured-output documentation.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Keep provider capability flags in configuration rather than assuming that a schema, citations, browser rendering, or pagination works identically everywhere.

A practical implementation sequence

  1. Specify the consumer. Write down which service, database, queue, or analyst will read each field.
  2. Draft and version the schema. Define types, required properties, null semantics, arrays, nested objects, and evidence fields.
  3. Classify the source. Choose selectors for stable templates; choose semantic extraction when layout or wording varies.
  4. Configure rendering. Wait for network idle or a known selector, set a timeout, and capture blocked or failed states.
  5. Request the data. Supply the URL or HTML, prompt, and response format supported by the chosen endpoint.
  6. Validate twice. Run a JSON Schema validator, then apply domain and provenance checks.
  7. Follow every page. Implement the documented cursor or offset model and record limits and counts.
  8. Monitor drift. Alert on rising empty results, selector misses, schema failures, unsupported citations, and pagination anomalies.

Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server that can provide a rendered visual or PDF for a page before you run your own extraction checks. It accepts consent banners as a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each cleanup step can be disabled. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing status. Its MCP server exposes take_screenshot, get_page_info, and capture_pdf to Claude, Cursor, and other MCP clients.

One GET request returns PNG, JPEG, WebP, or PDF. See the ScreenshotNeo documentation for all options.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

The free plan includes 1,000 screenshots per month with no card. Paid plans start at $5 for 3,000 shots; every feature is included on every plan. Sign up for ScreenshotNeo.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Frequently Asked Questions

How do I define what data to extract?

Define a versioned response contract: field names, types, required and optional properties, null behavior, arrays, nested objects, and evidence fields. Then express semantic goals in a prompt and enforce the shape with JSON Schema where supported.

How do I extract lists and attributes?

Use CSS selectors when list items and attributes have stable DOM markers. For variable layouts, request an array in the schema, describe each item and attribute, wait for rendered content, and validate every item after extraction.

How do I get typed JSON from a webpage?

Provide a prompt, a JSON Schema response format, or both to an endpoint that supports them, then run client-side schema and domain validation. Typed output controls structure; it does not verify the page’s facts.

Why is my result empty on a JavaScript-heavy page?

The capture may have occurred before scripts finished, the selector may be wrong, or access may have been blocked. Wait for network idle or a known content selector, classify timeout and blocked states separately, and check the provider’s rendering limits.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.