Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

Use a content endpoint for one page, a scrape endpoint for selected elements, and a crawl endpoint for linked pages. Turn on browser rendering when JavaScript builds the page, then wait for networkidle0, networkidle2, or a selector that proves the data is ready. For typed records, request JSON with a prompt or schema, validate every field against the source page, and retain the source URL for auditing.

Choose the API shape that matches your output

“Extract the HTML” and “crawl the site” are different jobs. Picking the narrowest endpoint reduces rendering time, parsing work, and accidental collection of pages you do not need.

Need Best pattern What you receive
One page’s rendered markup Content endpoint HTML for the page after browser JavaScript runs; Cloudflare documents that its content endpoint includes the head section as well as the rendered body. Cloudflare content endpoint
Specific fields or elements Scrape endpoint with CSS selectors Structured details for matching elements, including inner HTML and dimensions
Many related pages Crawl job An asynchronous job that discovers child pages, subject to depth, page-limit, source, and include/exclude rules. Cloudflare crawl endpoint
Typed records such as products or articles JSON extraction Fields shaped by a prompt and, where supported, a JSON response format or schema

Do not ask a crawler for a whole DOM when you need three fields, and do not build a one-off selector script when you need to discover an entire documentation section.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Decide whether static HTML is enough

Use a direct HTTP request first

Fetch the page without a browser when the values you need are present in the server response. This is faster and transfers fewer resources. Inspect the raw response for the title, product data, pagination links, or embedded JSON. If the data is there, parse it directly instead of paying the cost of JavaScript execution.

Scrapy’s guidance makes the same trade-off: reproducing the underlying data request can provide “structured, complete data with minimum parsing time and network transfer.” For an application that loads data from a JSON request, reproduce that request rather than rendering every visual component.

Switch to browser rendering for client-side pages

Single-page applications often return a nearly empty shell and populate the DOM later. A normal page-load event only means navigation reached a browser milestone; it does not prove that the API response has been rendered. Cloudflare documents a render: false option for static crawling and rendered mode by default in its crawl API. Use rendered mode when the required content appears only after scripts run.

Wait for the data, not merely navigation

  • networkidle0: wait until there are no active network connections. It is the strictest choice and can delay pages with analytics or streaming requests.
  • networkidle2: allow up to two active connections. It is often a practical compromise for pages that keep background requests open.
  • waitForSelector: wait for a stable element such as [data-testid="results"] or main article. This is usually the most meaningful condition when you know what “ready” looks like.

Prefer a selector that represents the actual data, with a timeout as a safety limit. If a page can legitimately contain zero results, wait for a container or an explicit empty-state element rather than a result row that may never appear.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Extract one page as rendered HTML

Cloudflare’s content request is a POST to the account endpoint with the target URL and an API token. The following examples use environment variables so credentials do not end up in shell history or source control.

cURL

export ACCOUNT_ID="your-account-id"
export CF_API_TOKEN="your-api-token"

curl -X POST "https://api.cloudflare.com/client/v4/accounts/$ACCOUNT_ID/browser-run/content" 
  -H "Authorization: Bearer $CF_API_TOKEN" 
  -H "Content-Type: application/json" 
  --data '{"url":"https://example.com"}'

Save the JSON response, then read the returned HTML field according to the response schema for your account. Keep the requested URL alongside the result and record the fetch time.

Python

import os
import requests

account_id = os.environ["ACCOUNT_ID"]
token = os.environ["CF_API_TOKEN"]
endpoint = f"https://api.cloudflare.com/client/v4/accounts/{account_id}/browser-run/content"

response = requests.post(
    endpoint,
    headers={
        "Authorization": f"Bearer {token}",
        "Content-Type": "application/json",
    },
    json={"url": "https://example.com"},
    timeout=90,
)
response.raise_for_status()
payload = response.json()
print(payload)

Node.js

const accountId = process.env.ACCOUNT_ID;
const token = process.env.CF_API_TOKEN;
const endpoint = `https://api.cloudflare.com/client/v4/accounts/${accountId}/browser-run/content`;

const res = await fetch(endpoint, {
  method: 'POST',
  headers: {
    'Authorization': `Bearer ${token}`,
    'Content-Type': 'application/json'
  },
  body: JSON.stringify({ url: 'https://example.com' })
});

if (!res.ok) throw new Error(`${res.status} ${await res.text()}`);
console.log(await res.json());

For repeatable extraction, store the raw response before cleaning it. That gives you a replayable audit record when a selector or JSON field changes.

Extract selected elements with CSS selectors

Use a scrape endpoint when the page is large but your output is small. Selectors such as article h1, .price, or [data-product-id] can return element text, attributes, inner HTML, and geometry instead of forcing you to parse an entire document. Cloudflare describes its scrape endpoint as returning structured details for specific elements, including dimensions and inner HTML.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Make selectors resilient

  • Prefer semantic elements, stable IDs, or data attributes over classes generated by a design system.
  • Return a small set of selectors and include the page URL with every record.
  • Expect zero matches, multiple matches, and missing attributes; treat each as a normal outcome that your code validates.
  • Keep a fixture of representative pages so a redesign fails a test instead of silently producing empty fields.

If your provider returns inner HTML, parse it with an HTML parser rather than regular expressions. Decode entities, normalize whitespace deliberately, and preserve the original fragment when later review may matter.

Extract JSON from a JavaScript-heavy page

Ask for a schema-shaped response

Cloudflare exposes JSON options with a prompt and response-format or schema controls. XCrawl also documents JSON output with a prompt and optional JSON schema. A useful schema makes the crawler’s contract explicit:

{
  "type": "object",
  "required": ["name", "price", "currency"],
  "properties": {
    "name": {"type": "string"},
    "price": {"type": "number"},
    "currency": {"type": "string"},
    "availability": {"type": ["string", "null"]}
  },
  "additionalProperties": false
}

Phrase the extraction prompt narrowly: identify the product name, numeric price, currency, and availability shown on the page; return null when a field is absent; do not infer values from related pages. The exact JSON-option wrapper differs by API, so submit these controls using the provider’s current request schema.

Validate before writing to your database

  1. Parse the response as JSON and reject malformed output.
  2. Validate required keys, types, allowed values, and numeric ranges.
  3. Compare the extracted values with the source HTML or a selected-element result.
  4. Store the source URL, retrieval timestamp, and raw response with the normalized record.
  5. Send ambiguous records to a review queue instead of silently coercing them.

Schema guidance improves consistency; it does not make the result authoritative. Prices, stock labels, and dates can be absent, duplicated, or rendered differently for different visitors.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Crawl a whole site without losing control

A crawl starts with one URL and follows child pages. Cloudflare’s crawl API returns a job that you check separately rather than keeping one request open until every page finishes. Configure these controls:

  • depth: how many link levels to follow from the starting URL.
  • limit: the maximum number of pages to collect.
  • source: discover URLs from sitemaps, links, or all.
  • include/exclude patterns: constrain paths such as /docs/ and exclude account, search, or logout URLs.
  • formats: request html, markdown, or json.
  • rendering options: choose static or browser-rendered fetching and set waits appropriate to the site.
  • JSON options: provide a prompt and response format or schema when requesting structured fields.

Start with a minimal crawl job

export ACCOUNT_ID="your-account-id"
export CF_API_TOKEN="your-api-token"

curl -X POST "https://api.cloudflare.com/client/v4/accounts/$ACCOUNT_ID/browser-rendering/crawl" 
  -H "Authorization: Bearer $CF_API_TOKEN" 
  -H "Content-Type: application/json" 
  --data '{"url":"https://example.com"}'

Use the job identifier and status procedure returned by the API to monitor completion. For a production crawl, add a conservative depth and page limit first, confirm the discovered URL set, then increase limits. Keep include and exclude rules under version control so a later run is explainable.

Choose discovery sources deliberately

Sitemaps are predictable for documentation and commerce catalogs. Link discovery follows what the site exposes in its rendered navigation. all can find more pages but also increases the chance of faceted URLs, calendars, and duplicate tracking parameters. Normalize URLs and de-duplicate them before downstream processing.

Why a scraper returns empty HTML

The page is an SPA shell

Symptom: you receive a root element with no products, comments, or article body. Fix: enable rendering and wait for networkidle0, networkidle2, or a known content selector.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Your wait condition is wrong

Symptom: the page is sometimes complete and sometimes partial. Fix: wait for the element that your extractor consumes, increase the timeout within the provider’s limits, and avoid relying only on a fixed sleep.

The data is in an underlying request

Symptom: browser HTML is expensive but the page calls a stable JSON endpoint. Fix: identify and reproduce that request directly, subject to authentication, terms, rate limits, and publisher controls.

A bot check or access boundary intervenes

Symptom: you capture a challenge page, login form, or denial instead of content. Fix: do not treat the challenge as valid data. Check permissions, cookies, authentication requirements, and the site’s rules. Cloudflare notes that changing the user agent does not bypass Browser Run bot identification.

Selectors changed

Symptom: the response is non-empty but fields are null. Fix: log match counts, retain sample HTML, test fallback selectors, and alert when required fields disappear.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Compliance, reliability, and cost controls

Respect publisher and legal boundaries

Check robots.txt, terms of service, authentication boundaries, rate limits, and applicable law before collecting data. Cloudflare’s crawl API exposes contentUse and crawlPurposes controls for publisher Content-Signal directives. Those controls do not establish a universal legal rule for every jurisdiction; document your own purpose and review process.

Make runs reproducible

  • Record URL, timestamp, rendering mode, wait condition, selector or schema version, HTTP status, and page verdict.
  • Cache immutable pages and use a chosen TTL for changing pages.
  • Retry transient network failures with exponential backoff, but do not retry authorization failures indefinitely.
  • Set a per-page timeout and a crawl-wide budget so one hanging resource cannot consume the entire job.
  • Hash raw HTML or normalized JSON to detect unchanged pages.

Estimate spend before scaling

Pricing, quotas, concurrency, timeout maxima, and cache behavior vary by vendor and plan. Verify the current plan documentation before committing to a volume. Measure pages discovered, rendered pages, retries, and bytes transferred; a low page count can still be expensive when every page launches a browser and loads many third-party resources.

Or skip the browser setup

If your goal is a clean visual capture rather than HTML or field extraction, ScreenshotNeo provides a single-call screenshot API. It accepts consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and each response reports the page verdict and billing status.

It also offers an MCP server for Claude, Cursor, and other MCP clients, with take_screenshot, get_page_info, and capture_pdf tools. Every plan includes the features; 1,000 screenshots per month are free with no card, Starter is $5 for 3,000, and paid plans start at $5.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

One-call example

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

See the ScreenshotNeo API documentation for options such as full-page capture, element selectors, custom JavaScript, waiting for a selector or network idle, cookies and headers, PDFs, signed links, asynchronous jobs, and bulk capture. Create a free ScreenshotNeo account to get 1,000 screenshots each month without adding a card.

Practical decision checklist

  1. Can the server response provide the data? Use a direct request.
  2. Do scripts build the DOM? Enable rendering and wait for a meaningful selector.
  3. Do you need a few fields? Use selector extraction and validate match counts.
  4. Do you need typed records? Supply a strict prompt and schema, then compare with source content.
  5. Do you need linked pages? Start a bounded crawl with depth, limit, source, and include/exclude rules.
  6. Could a challenge, login wall, or consent layer contaminate results? Detect and classify it instead of storing it as content.
  7. Can another engineer explain the result later? Save the URL, settings, raw response, and extraction version.

Frequently Asked Questions

How can I tell whether a page is truly ready for extraction?

Wait for the element your parser needs, then verify that required fields have non-empty values. A completed navigation event alone is not proof that an SPA finished rendering.

Should I request HTML, Markdown, or JSON for an archive?

Keep raw rendered HTML when fidelity and later re-parsing matter. Use Markdown for readable text corpora, and JSON when you have a stable schema plus validation and source snapshots.

What should happen when a crawl finds a challenge or login page?

Classify it as an access failure, stop treating it as source content, and resolve authorization or site-policy requirements before retrying.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.