Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

Use an extraction API when you need page content as predictable fields rather than raw HTML. The usual process is to define a schema, choose direct fetching, browser rendering, crawling or a page-type extractor, submit the URL and credentials, then validate the response and retain its source URL. Hosted services differ in how they discover pages, execute JavaScript, represent missing values and expose jobs, so test representative pages before building a production pipeline.

What “structured data” means

Structured output has named fields and values in a predictable representation such as JSON. Instead of asking your application to interpret an entire HTML document, you can request fields such as title, author, published_at, price and currency. A schema can also state types and whether a field may be absent.

For example, an article record might be:

{
  "source_url": "https://example.com/story",
  "title": "Example headline",
  "author": "A. Writer",
  "published_at": "2026-09-20",
  "tags": ["technology"],
  "image_url": null
}

The source_url and retrieval time are useful provenance. Keep them with the extracted record so a person can revisit the page when a value looks wrong.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Choose the extraction approach

Approach Use it when Questions to answer first
Direct page extraction One known URL is accessible without site-wide discovery. Does the service read static HTML only, or can it render a browser? Can fields be named and typed?
Schema-driven extraction Your application requires a stable, custom JSON shape. How are absent fields represented? Is the source evidence or provenance returned?
Crawler or hosted scraper Data is spread across many pages or must run in batches or on a schedule. How are links discovered, limits enforced, jobs retried and datasets exported?
Page-type extractor Pages fit a supported class such as article or product. Which types are supported, and how are classification or extraction failures reported?

Context.dev documents site crawling into a JSON Schema you define (documentation). Scrapy.io documents scraper discovery, synchronous and asynchronous jobs, polling, dataset export and recurring schedules (API documentation). Firecrawl describes extraction from one or multiple URLs with prompts and/or schemas (project documentation). Diffbot provides typed page extractors, including several page classes (Extract API documentation).

A practical workflow

1. Define the contract

Write the fields your downstream code actually consumes. Specify string, number, Boolean, date or array types; permitted nulls; units and normalization rules. Decide whether a missing value should be null, an omitted key or a validation error. Do not request dozens of fields “just in case”: each additional field creates another opportunity for ambiguity.

2. Select representative URLs

Include normal pages, pages with missing fields, a recently changed page and a page whose content is loaded by JavaScript. Compare returned values with the rendered page before expanding the job. Documentation describes capabilities, not independent accuracy rates, so this sample is your quality check.

3. Decide whether JavaScript is required

Static HTML may contain the data you need. If the initial response is only an application shell, select a browser-rendered mode if the provider offers one. Monocrawl documents direct static fetching separately from an explicitly requested browser mode and notes that non-direct modes are deployment-gated and off by default; that behavior is vendor-specific, not a rule for every API (Monocrawl documentation).

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

4. Choose one page or a crawl

A single URL call is simpler and easier to retry. A multi-page task needs discovery as well as extraction: define allowed domains, path rules, maximum pages, pagination handling and whether links may leave the starting section. Context.dev says its crawler prioritizes relevant internal links; Scrapy.io documents discovery and run/job endpoints.

5. Submit the request

An extraction request normally contains an API credential, URL or crawl seed, schema or scraper selection, and crawl settings. Use the exact endpoint, authentication header and parameter names in the provider’s current documentation. Store credentials in environment variables or a secret manager, never in source control. Hosted APIs described here expose different contracts, so there is no single cross-provider request URL.

6. Validate before persistence

Parse the response as JSON, check required keys and types, validate dates and numeric ranges, and reject or quarantine records that fail. Record the provider’s status, any error or classification signal, the original URL and retrieval timestamp. Keep raw responses when your retention policy permits; they make debugging a changed selector or page layout possible.

7. Operate the job

For asynchronous crawls, submit once, poll the documented job status, then export the dataset. Apply bounded retries with backoff for transient transport errors, but do not endlessly retry an access denial or a deterministic schema failure. For recurring collection, record a run identifier and reconcile duplicates using a stable key such as canonical URL plus publication identifier.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Getting website data as JSON

The response shape depends on the provider. A schema-oriented service may return your named fields directly; a scraper platform may return a dataset whose items contain those fields; a page-type extractor may add classification metadata. Normalize all three into an internal envelope:

{
  "source_url": "…",
  "retrieved_at": "…",
  "status": "ok",
  "data": { "title": "…" },
  "errors": []
}

Keep extraction errors separate from legitimate empty values. “No price shown” is different from “the extractor timed out.” If the service supplies evidence such as the source text or selector, retain it beside the field for review.

Crawling, batching and scheduling decisions

  • Discovery: determine how links are found and whether robots, canonical links, sitemaps or explicit URL lists are used.
  • Limits: set page, depth, request-rate and payload limits before a crawl can expand unexpectedly.
  • Jobs: use asynchronous execution when a crawl may outlive an HTTP request; persist job IDs and poll status.
  • Exports: verify whether the dataset is JSON, JSON Lines, CSV or another format and how pagination works.
  • Schedules: define what counts as a new or changed record and how failed runs are surfaced.

Scrapy.io documents these job, dataset and recurring-schedule concepts. A crawler is not automatically better than a direct call: for a handful of known URLs, discovery adds latency and cost without helping.

Reliability, performance and cost

Do not infer comparative speed, accuracy or price from vendor feature lists. The available documentation does not provide independent head-to-head measurements or a common pricing comparison. Measure the fields that matter to your project: valid-record rate, missing-field rate, latency, browser-render percentage, retry rate and total request cost.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Improve throughput by sending only required fields, limiting crawl scope, reusing cached results where policy permits and separating discovery from extraction. Protect the target and your budget with concurrency limits. Browser rendering generally consumes more resources than static fetching, so enable it only for pages that need it. Cache by URL and relevant options, and invalidate when the page’s update cadence requires fresh data.

Common failures and fixes

Empty or null fields

Likely causes: the field is absent, content is JavaScript-generated, or the schema name does not match the page. Fix: inspect the page, confirm the exact field contract, try the provider’s browser mode, and distinguish a legitimate null from an extraction error.

HTML shell instead of content

Cause: data is rendered after load. Fix: use a documented browser-rendered mode, wait for the relevant selector or network idle condition if supported, and retest on a representative URL.

Crawl finds too many or too few pages

Cause: link rules, pagination or domain boundaries are wrong. Fix: constrain allowed paths and depth, define pagination behavior, set a page limit, and inspect discovered URLs before enabling a large run.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Job remains pending or fails

Cause: asynchronous work needs polling, a transient provider issue occurred, or credentials and limits are invalid. Fix: poll the documented status endpoint, use bounded backoff, capture the returned error, and verify credentials and quota before retrying.

Values changed after deployment

Cause: the site layout or content changed. Fix: retain provenance and sample responses, add schema validation and alerts for unusual null rates, then update the schema or scraper deliberately.

Access is denied

Review the site’s terms, access rules and applicable law for your jurisdiction and use case. The available documentation does not establish a universal legal rule. Do not attempt to bypass authentication, CAPTCHAs or technical restrictions without authorization.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup

If your task starts with collecting clean visual evidence of pages rather than parsing fields, ScreenshotNeo provides a website screenshot API and MCP server. It accepts one GET request and returns PNG, JPEG, WebP or PDF. Before capture it can accept cookie banners and remove more than 60 known consent platforms, newsletter popups and chat widgets; each step can be disabled. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads and cache hits are not billed, and response headers identify the page verdict and billing result.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

It also supports full-page shots with lazy images loaded, CSS-selector element capture, dark mode, device presets or custom viewports, retina scale, PDF paper and page settings, custom CSS and JavaScript, click-before-capture, waits, request blocking, headers, cookies, user agents, authorization, timezone, geolocation, transparent backgrounds, resizing, configurable-TTL caching, signed image links, asynchronous webhooks, bulk capture of up to 100 URLs per call, usage data and an OpenAPI specification. Its MCP tools are take_screenshot, get_page_info and capture_pdf.

Use the documented API details at ScreenshotNeo docs. cURL:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`HTTP ${res.status}`);
require('fs').writeFileSync('shot.webp', Buffer.from(await res.arrayBuffer()));

The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 shots. Sign up free.

Further learning

Hands-On Web Scraping with Python includes a section on extraction with web APIs (PDF). Check the edition and availability yourself; the document does not establish a current print edition or marketplace listing.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Frequently Asked Questions

Should I use an API or write my own scraper?

Use an API when hosted rendering, crawling, retries or scheduling save more engineering effort than they cost. Build in-house when you need specialized control, already operate the browser and crawl infrastructure, or cannot send data to a hosted service.

How can I tell whether an extraction result is trustworthy?

Compare it with representative source pages, enforce types and required fields, retain provenance, and monitor missing-field and error rates over time.

What is the difference between extraction and crawling?

Extraction turns one or more known pages into fields. Crawling discovers pages and then extracts them, adding scope, limits and job-management concerns.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.