Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
You can extract web data by describing the records and fields you need in plain language, then asking an extractor to return them in a defined structure such as a JSON Schema. For dependable results, use a browser-capable extractor when the page relies on JavaScript, validate the returned values, and retain each record’s source URL and extraction time. Natural-language instructions explain the task; a schema makes the output contract explicit.
What natural-language web extraction does—and does not do
Instead of writing selectors for every field, you describe the target content: for example, “Extract every product card and return its name, price, availability, and product URL.” An extraction service interprets the request against a webpage and returns structured records. Some services accept a JSON Schema as well as a prompt; Cloudflare documents its /json endpoint as extracting structured data from a webpage and accepting either a prompt or a JSON Schema.
This changes how you specify the task, not the need to verify the result. A natural-language request alone does not guarantee that every item was found, that values are correct, or that pagination was covered. Use a schema to constrain shape and types, and validation plus human review to check meaning and completeness.
Choose the right extraction approach
| Approach | Best fit | Trade-off to plan for |
|---|---|---|
| Prompt plus JSON Schema API | A quick, structured extraction from a page | Requires provider access and careful validation |
| Browser agent plus schema | Interactive pages or pages whose content appears after JavaScript runs | More moving parts and potentially higher runtime cost |
| Deterministic selectors | A stable, known layout with repeating rows or cards | Markup or layout changes can break selectors |
| Multi-page crawler | Catalogs, directories, and paginated sites | Requires crawl boundaries, deduplication, and rate-limit controls |
Cloudflare documents product, listing, and article-metadata use cases for its Browser Run JSON endpoint. Refyne documents single-page extraction, multi-page crawling, and JSON, JSONL, or YAML output. Magnitude’s BrowserAgent combines natural-language instructions with a Zod schema. Twin Browser documents extraction from a live rendered page using a field list, map, or JSON Schema, as well as a zero-LLM selector path when selectors are already known. Choose by the work you need done: rendering, schema support, crawling, selector control, output formats, and operating cost are distinct considerations.
#1 Best Overall
Define the record before writing the prompt
Start by deciding what one output object represents: one product, job, article, listing, or another repeated item. Then name every field and its type. Specify what to do when a value is absent, ambiguous, or displayed in a different currency or unit. If the output will enter code, a spreadsheet, or a database, write a schema before scaling the job.
For example, this schema describes an array of product records. It makes fields mandatory while allowing a missing page value to be represented as JSON null; the extractor should not guess a missing price or rating.
{
"type": "array",
"items": {
"type": "object",
"properties": {
"name": { "type": "string" },
"brand": { "type": ["string", "null"] },
"price": { "type": ["number", "null"] },
"currency": { "type": ["string", "null"] },
"availability": { "type": ["string", "null"] },
"rating": { "type": ["number", "null"] },
"review_count": { "type": ["integer", "null"] },
"product_url": { "type": "string" }
},
"required": [
"name", "brand", "price", "currency", "availability",
"rating", "review_count", "product_url"
],
"additionalProperties": false
}
}
In the schema, required means each key must be present; it does not mean the page must contain a value. A nullable type lets a record contain the key with null when the value is absent. Treat price as a number only if your downstream system also stores currency and the extraction rules specify how displayed formatting is handled.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Write a prompt that limits interpretation
Give the extractor the page, the repeated item to find, field definitions, inclusion rules, and an explicit treatment of missing values. For product cards, a practical starting instruction is:
Open the supplied page and extract one record for each product card.
Fields:
- name: string; the product name shown on the card
- brand: string or null; null if not shown
- price: number or null; use the displayed price, without a currency symbol
- currency: string or null; preserve the currency shown
- availability: string or null; preserve the page's wording
- rating: number or null; use the visible rating
- review_count: integer or null; use the visible count
- product_url: string; the product link from the card
Rules:
- Include only products visibly listed on the page.
- Ignore sponsored blocks.
- Do not infer absent values; use null.
- Return the source URL for each record.
- Return an array matching the supplied JSON Schema.
Keep field definitions operational. “Price” is vague if the page shows both a sale price and a former price; say which one to capture. “Rating” may mean a score, a count, or both, so define each separately. Tell the extractor to preserve units and currency rather than silently converting them.
Use a browser when the page needs one
A simple page fetch may not contain the content a person sees. JavaScript can populate cards after initial load, and interactive pages may require a click, consent dismissal, or navigation. In those cases, choose a browser-capable extractor that reads the rendered page and can perform the required interactions. Browser-agent workflows pair instructions with rendered-page access; where the layout is stable and selectors are known, selector-based extraction can be more deterministic.
Rank #3
For an interactive task, include the navigation sequence and a stopping rule in the instruction. For example: “Open the results page, dismiss the consent dialog if it blocks the results, select Next until no Next control remains, and stop. Extract only visible job listings.” Explicitly bound the task so an agent does not follow unrelated links indefinitely.
Natural-language extraction tools may expose different interfaces and schemas. Refyne documents extraction across one or multiple pages, while Twin Browser documents a live rendered-page workflow and a selector-based path. Confirm the chosen provider’s current input format, output format, and access requirements before wiring it into a pipeline.
Run a small extraction, then validate it
- Test one representative page. Include a normal page and, if relevant, a page with missing fields or an unusual card. Check whether the extractor sees the rendered content you expect.
- Parse the response. Reject malformed JSON and validate each record against the schema. Check required keys, value types, URL shape, and permitted nulls.
- Review content, not just shape. Confirm that a sample of names, prices, units, availability, and links matches what is visibly on the page. A schema can ensure that a price is numeric; it cannot prove the number is the right price.
- Check coverage and duplication. Count visible items against returned records, inspect duplicate URLs or names, and check whether pagination or lazy-loaded content was missed.
- Preserve provenance. Store the source URL, retrieval timestamp, schema version, and extraction prompt with each batch. Keep a small human-reviewed sample so later changes in the page or prompt can be diagnosed.
Chrome Developers’ built-in AI guidance recommends using a JSON Schema for predictable results and warns against relying on natural-language instructions such as “output only JSON” alone. That is a useful distinction: the prompt describes the task, while the schema constrains the output form.
Scale to multiple pages carefully
Once a single-page sample is sound, add pagination, retries, rate-limit handling, and duplicate detection. Define where the crawl starts and ends, which links count as next-page navigation, and how to stop. Store a stable item identifier when the source provides one; otherwise, choose a deduplication key deliberately, such as a canonical product URL rather than a display name that might change.
Test pagination on a small batch before running a large crawl. Keep the extraction prompt and schema version attached to results: changing either can alter what is collected or how nulls are represented. Follow the provider’s limits and the target site’s access rules; an extractor’s ability to load a page does not establish permission to collect or reuse its content.
Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchWindows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallOr skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server, not a natural-language data extractor. It can provide a clean page capture as an input to a separate review or extraction workflow; it does not return structured records from the page. Its API makes one GET request to capture a URL. For example, save a capture of a product listing page as WebP:
Best Value
ScreenshotNeo API documentation
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
The Python equivalent is:
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
open("shot.webp", "wb").write(r.content)
And Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
Replace the example URL with the page you intend to capture. ScreenshotNeo accepts cookie or consent banners as a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each cleanup step can be turned off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers identify the page verdict and billing status. Its MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients.
ScreenshotNeo offers 1,000 screenshots a month free without a card; paid plans start at $5 for 3,000. See ScreenshotNeo and its documentation for the capture workflow. Sign up for the free plan to try a capture workflow without a card.
Common problems and fixes
- Records are missing cards visible in the browser. The extractor may be reading the initial HTML rather than rendered content, or lazy loading may not have completed. Use a browser-capable workflow, wait for the relevant content, and compare the record count with the page.
- Output contains prose or invalid JSON. Do not depend on a prompt that merely says “return JSON.” Provide a JSON Schema where supported, parse the response, and reject output that fails validation.
- Fields have the right type but wrong meaning. Tighten field definitions, distinguish similar values such as sale and list price, and state whether to preserve page wording or normalize it. Review a sample against visible source content.
- Absent values are invented or omitted inconsistently. Specify
nullfor unavailable values and make the field nullable and required in the schema. State explicitly that the extractor must not infer missing information. - Pagination yields duplicates or gaps. Define the next-page action and a stopping rule. Track visited page URLs, deduplicate records with a stable key, and test a short run before increasing the crawl size.
- Results change between runs. Pages can change, rendering can vary, and prompts or schemas can be edited. Preserve the source URL, retrieval time, prompt, and schema version so a discrepancy can be traced.
Reliability, performance, and cost
There is no common accuracy percentage or universal success rate established by the reviewed official documentation. Results depend on page rendering, layout consistency, prompt specificity, schema design, access restrictions, and validation. Treat extracted fields as data to check, not proof that every value was captured correctly.
The Tool Desk
Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Operationally, rendered-browser work and multi-page crawling introduce more steps than extracting a static, stable page. Run a small sample first, bound the crawl, and account for provider access and runtime costs before scaling. A known, stable layout may favor selectors; a changing or interactive layout may justify a browser agent, but its additional flexibility does not remove the need for verification.
Frequently Asked Questions
Can I extract data from a page without writing CSS selectors?
Yes. A natural-language prompt can describe the records and fields. For predictable downstream output, pair it with a JSON Schema when the service supports one.
Does a schema guarantee the extracted values are accurate?
No. It constrains output structure and types. You still need to check values against the page and verify coverage.
When should I use selectors instead of an AI prompt?
Selectors are a good fit when the layout is stable and you know the relevant elements. A browser agent is useful when rendering or interaction is needed, but page changes can affect either approach.
Free tools Windows power users keep installed
One-click scans. No signup required.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

