Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

Use a cascade, not a model-first scraper. Fetch the page reliably, read schema.org JSON-LD or framework hydration data, probe a reachable product-data request, and try deterministic selector relocation before asking an LLM to infer anything. When those layers cannot cover the fields, have a model generate a selector map once, validate it on the template, and run that map as ordinary code. This reduces calls and latency while making semantic errors visible.

This is an engineering pattern rather than a guarantee that every store exposes complete structured data. A parser cannot repair a blocked, incomplete or unrendered page; fetching, JavaScript execution and access challenges belong underneath the extraction cascade.

What “zero-shot web scraping” means here

In this context, zero-shot scraping means extracting product attributes without hand-writing a custom parser for every retailer or training a store-specific model. An LLM may infer fields from an unfamiliar page, but the reliable design is to reserve that inference for the last fallback. The local LLM belongs at the bottom of the cascade, as the fallback you use last.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Do not confuse this HTML-focused workflow with image-based zero-shot attribute generation. The NAACL 2025 paper Visual Zero-Shot E-Commerce Product Attribute Value Extraction describes a cross-modal method using product images, OCR and a prompt-based language model. That is a different problem from reading a retailer’s HTML and embedded data.

The extraction cascade

Layer What you inspect Strength Typical failure
1. Embedded data JSON-LD and serialized framework state Typed values that are less sensitive to CSS class renames Missing fields, stale state or malformed markup
2. Internal API Fetch/XHR or GraphQL requests returning product data Structured response without selector maintenance Endpoint is private, unstable or store-specific
3. Selector relocation Existing selector plus nearby structural fingerprints Fast, deterministic repair of superficial markup drift Genuine template restructure
4. Reusable LLM map One representative page used to generate selectors Model cost is paid once; map can be reviewed and versioned Wrong semantics or a template change that validation misses

First make sure the page was actually fetched

A 403, 429, CAPTCHA, JavaScript challenge or tiny “enable JavaScript” stub is an access or rendering failure, not a bad CSS selector. Use a browser-capable fetcher when the product is assembled client-side, and record status code, final URL, response size and whether the expected root element appeared. If an access layer blocks the request, change the fetching strategy before changing parsers.

For a site you control or are authorized to crawl, inspect one page in a real browser, save the post-render HTML, and compare it with the raw HTTP response. Keep fetching and parsing as separate modules so a parser test can run against a fixed fixture.

Stage 1: inspect JSON-LD and hydration state

Find schema.org Product objects

Start with every <script type="application/ld+json"> element. A page can contain a single Product, an array, or an @graph; some sites emit several offers or variants. Check the type, field presence and value types instead of assuming the first object is complete. Product, offers, aggregateRating and brand often have different nesting and may be absent.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The Python example below uses extruct to read embedded metadata from saved HTML. It deliberately returns all Product-shaped objects so a later validator can choose the record matching the visible SKU or URL.

from extruct import extract
from w3lib.html import get_base_url


def product_objects(html: str, page_url: str):
    data = extract(html, base_url=get_base_url(html, page_url), syntaxes=["json-ld"])
    products = []
    for item in data.get("json-ld", []):
        candidates = item if isinstance(item, list) else [item]
        for obj in candidates:
            if not isinstance(obj, dict):
                continue
            graph = obj.get("@graph")
            nodes = graph if isinstance(graph, list) else [obj]
            for node in nodes:
                if isinstance(node, dict):
                    kind = node.get("@type")
                    kinds = kind if isinstance(kind, list) else [kind]
                    if "Product" in kinds:
                        products.append(node)
    return products

with open("product.html", encoding="utf-8") as f:
    records = product_objects(f.read(), "https://shop.example/item/123")
print(records)

Inspect framework hydration

Search for serialized state such as __NEXT_DATA__, __NUXT_DATA__ and __remixContext. Parse the JSON, then map candidate paths to your output fields. Hydration data can contain the complete product object, but it can also contain only routing or recommendation state. Verify that price, currency, availability, SKU and variant information are present and current.

Stage 2: probe the site’s own data requests

Open browser developer tools, select the Network panel, filter to Fetch/XHR, reload a product page and change a variant. Identify the request whose response contains the product record. Record its method, URL, query or body, cookies and headers, then replay the smallest request that works. Treat this as manual, store-specific discovery: an internal product endpoint may not exist, may require a session, or may change without notice.

The available sandbox example exposes a cart endpoint rather than a product API, so do not assume that finding one JSON route proves a general product service. Add contract checks for HTTP status, content type, required fields and variant identity. Respect the retailer’s terms, robots policy and rate limits.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Stage 3: relocate selectors when markup drifts

If an established selector stops matching after a class rename or small move, search nearby structure rather than immediately invoking a model. A fingerprint can combine tag name, stable attributes, a label such as “price,” and the relative position of a known heading. Score candidates, extract the value, normalize currency and compare it with an independent signal such as JSON-LD or a visible test fixture.

In the target article’s simulated sandbox, price relocation succeeded on 12 of 12 pages in 78 ms with zero tokens. That is an article-reported run, not a production guarantee. A real page restructure, a changed variant component or a different template may not be repairable by fingerprints.

Stage 4: generate a reusable selector map with an LLM

Ask for selectors, not final answers

Give a local model one representative, authorized HTML page and a strict output schema. Ask for selectors or JSON paths for the fields you need, plus a short rationale and confidence. Do not ask it to return the product values for every page; that turns each crawl into a fresh, slow inference task.

Return only JSON matching this schema:
{
  "title": "CSS selector",
  "price": "CSS selector",
  "currency": "CSS selector or null",
  "sku": "CSS selector or null",
  "availability": "CSS selector or null",
  "rating_value": "CSS selector or null"
}
Rules:
- Prefer attributes and relationships that repeat across this template.
- Never infer a numeric rating from the number of visible star icons.
- Use null when the field is not present.
- Do not return extracted values or executable code.

Validate across the template

Run the proposed map against several URLs from the same template before enabling it. Check that selectors match exactly one intended element, required fields are non-empty, numeric values parse, currency is plausible, and SKU or URL identity agrees with the page. Compare ratings with the source attribute or numeric class value; a model can see five decorative stars while the class encodes 3.7. Reject or regenerate the map when semantic checks fail, then store the accepted map in version control with the sample HTML and validation results.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Execute deterministically

Once validated, use a normal HTML parser for every page. Cache the map by site and template fingerprint, invalidate it when validation failure crosses a threshold, and send only the new representative page to the model. This makes model use auditable and prevents silent drift.

A practical Python pipeline

The following outline shows the control flow. The functions that fetch rendered HTML, call an internal API and relocate a selector are site-specific; keep their interfaces explicit.

def extract_product(url, html, old_map=None):
    embedded = product_objects(html, url)
    record = choose_complete_product(embedded)
    if record and covers_required_fields(record):
        return {"source": "json-ld", "data": normalize(record)}

    api_record = try_site_api(url)
    if api_record and covers_required_fields(api_record):
        return {"source": "internal-api", "data": normalize(api_record)}

    if old_map:
        values = run_selectors(html, old_map)
        if validate(values, html):
            return {"source": "selector-map", "data": values}
        relocated = relocate_selectors(html, old_map)
        if relocated and validate(relocated, html):
            save_map(relocated)
            return {"source": "relocated-map", "data": run_selectors(html, relocated)}

    new_map = ask_llm_for_selector_map(html)
    if not validate_map_on_sample(new_map):
        raise ValueError("No validated map for this template")
    save_map(new_map)
    values = run_selectors(html, new_map)
    if not validate(values, html):
        raise ValueError("Semantic validation failed")
    return {"source": "llm-generated-map", "data": values}

Keep a per-page audit record containing the fetch verdict, template identifier, extraction layer, validation results and normalized output. That record lets you distinguish a blocked page from a parser regression.

What the published measurements do—and do not—show

Results vary by store, template and field definition. Use the figures below as context, not promises for your crawl.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Study or run Reported result How to interpret it
Brosch, Brumm, Krieger and Scheffler (2025), 3,000 food pages from three shops LLM-generated extraction functions averaged 96.48% accuracy, 1.61 percentage points below direct extraction, with 95.82% fewer LLM calls Curated dataset and preprint; generation runs varied, so benchmark your own fields
WebLists benchmark (Bohra et al., 2025), 200 enterprise tasks 3% recall for LLMs with search and 31% for state-of-the-art web agents Shows that broad web extraction remains difficult outside a controlled template
WebLists’ BardeenAgent 66% overall recall and three-times lower cost per output row Reported by the benchmark authors for their proposed agent
Target article’s 12-page sandbox Direct extraction returned 87 of 96 fields (90.6%) and took 14–55 seconds per page; mean 30.1 seconds Small, local-hardware sample; rating errors came from reading star icons
Target article’s two-store cold run 65 products used one model call; the second run used zero because the cached map validated Example crawl economics, not a universal call rate

The article also reports Product markup on more than 3.3 million hosts across about 280 million URLs in an October 2024 Web Data Commons extraction. Web Data Commons documents that its class-specific corpus covers only a subset of pages and can contain duplicate annotations, so prevalence does not imply that an arbitrary live store exposes complete, current Product data.

Benchmark your own scraper

  • Field coverage: measure required fields separately; a perfect title score can hide missing variant prices.
  • Semantic correctness: compare normalized values with trusted fixtures, including currency, decimal separators, ratings and availability.
  • Drift resilience: test class renames, moved nodes and a genuinely new template.
  • Operations: record model calls, tokens, latency, fetch/render time, retries and access failures.
  • Reuse: measure how many pages run from a validated map before regeneration.

Sample every major template and include pages with variants, sale pricing, missing ratings and out-of-stock states. Re-run the suite after a redesign rather than trusting a single successful page.

Common failures and fixes

403, 429, CAPTCHA or a JavaScript challenge

Fix the fetch layer: use an authorized browser-rendered session, reduce concurrency, preserve required cookies and obey site policies. Do not tune selectors against an error page.

JSON-LD exists but fields are missing

Merge only after identity checks. Use hydration state or the internal request for variants, and keep the source of each field in your audit record.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The selector map matches multiple nodes

Add an ancestor constraint, stable data attribute or proximity rule. If the template truly changed, regenerate and validate instead of picking the first match.

Price or rating is plausible but wrong

Check hidden sale and list-price nodes, locale parsing and numeric attributes. Validate against the selected variant and never equate star count with rating value.

Map works on one page only

You likely sampled a different template or personalized state. Cluster pages by structure, maintain one map per cluster and test representative URLs from each.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server for developers. It can render a URL before you parse the result, while removing cookie-consent banners, newsletter popups and chat widgets; each cleanup step can be disabled. Bot checks, blank pages, timeouts, failed loads and cache hits are not billed, and response headers identify the page verdict and billing status. It is fetching infrastructure, not a substitute for validating product fields in JSON-LD or an API.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

One GET request returns PNG, JPEG, WebP or PDF. The API accepts full-page capture with lazy images, CSS-selector element capture, device and viewport settings, custom CSS or JavaScript, waits, request blocking, headers, cookies, user agents, geolocation, caching, signed links, asynchronous jobs, bulk capture and more. Its MCP server exposes take_screenshot, get_page_info and capture_pdf for Claude, Cursor and other MCP clients.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`${res.status} ${res.statusText}`);
const fs = await import('node:fs/promises');
await fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer()));

See the parameter reference in the ScreenshotNeo documentation. Plans are Free (1,000 shots/month, no card), Starter $5 for 3,000, Growth $15 for 15,000, Pro $39 for 60,000, Scale $99 for 250,000 and Business $249 for 1,000,000; yearly billing gives two months free, and every feature is on every plan. Start with the free ScreenshotNeo account.

FAQ

Can this approach extract data from product images?

Not reliably from HTML alone. Image attributes require a separate vision workflow, OCR and validation; do not treat that as the same zero-shot cascade.

How many pages should validate a selector map?

There is no universal count. Include every major template and state you intend to crawl, then set a failure threshold that triggers regeneration.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Should I keep both JSON-LD and visible-text outputs?

Yes. Keeping source provenance lets you detect disagreements instead of silently choosing whichever value parses first.

Is an internal endpoint an official public API?

Usually not. Confirm authorization, terms, required headers and stability before depending on a browser-observed request.

Frequently Asked Questions

Can this approach extract data from product images?

Not reliably from HTML alone. Image attributes require a separate vision workflow, OCR and validation; do not treat that as the same zero-shot cascade.

How many pages should validate a selector map?

There is no universal count. Include every major template and state you intend to crawl, then set a failure threshold that triggers regeneration.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Should I keep both JSON-LD and visible-text outputs?

Yes. Keeping source provenance lets you detect disagreements instead of silently choosing whichever value parses first.

Is an internal endpoint an official public API?

Usually not. Confirm authorization, terms, required headers and stability before depending on a browser-observed request.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.