Quick wins for a faster PC:
Clear out junk files and repair common Windows errorsFree Scan →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Repair Windows errors before they cause bigger problemsFix Now →Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
Use a content endpoint for one page, a scrape endpoint for selected elements, and a crawl endpoint for linked pages. Turn on browser rendering when JavaScript builds the page, then wait for networkidle0, networkidle2, or a selector that proves the data is ready. For typed records, request JSON with a prompt or schema, validate every field against the source page, and retain the source URL for auditing.
Choose the API shape that matches your output
“Extract the HTML” and “crawl the site” are different jobs. Picking the narrowest endpoint reduces rendering time, parsing work, and accidental collection of pages you do not need.
| Need | Best pattern | What you receive |
|---|---|---|
| One page’s rendered markup | Content endpoint | HTML for the page after browser JavaScript runs; Cloudflare documents that its content endpoint includes the head section as well as the rendered body. Cloudflare content endpoint |
| Specific fields or elements | Scrape endpoint with CSS selectors | Structured details for matching elements, including inner HTML and dimensions |
| Many related pages | Crawl job | An asynchronous job that discovers child pages, subject to depth, page-limit, source, and include/exclude rules. Cloudflare crawl endpoint |
| Typed records such as products or articles | JSON extraction | Fields shaped by a prompt and, where supported, a JSON response format or schema |
Do not ask a crawler for a whole DOM when you need three fields, and do not build a one-off selector script when you need to discover an entire documentation section.
Free tools Windows power users keep installed
One-click scans. No signup required.
Decide whether static HTML is enough
Use a direct HTTP request first
Fetch the page without a browser when the values you need are present in the server response. This is faster and transfers fewer resources. Inspect the raw response for the title, product data, pagination links, or embedded JSON. If the data is there, parse it directly instead of paying the cost of JavaScript execution.
#1 Best Overall
Scrapy’s guidance makes the same trade-off: reproducing the underlying data request can provide “structured, complete data with minimum parsing time and network transfer.” For an application that loads data from a JSON request, reproduce that request rather than rendering every visual component.
Switch to browser rendering for client-side pages
Single-page applications often return a nearly empty shell and populate the DOM later. A normal page-load event only means navigation reached a browser milestone; it does not prove that the API response has been rendered. Cloudflare documents a render: false option for static crawling and rendered mode by default in its crawl API. Use rendered mode when the required content appears only after scripts run.
Wait for the data, not merely navigation
networkidle0: wait until there are no active network connections. It is the strictest choice and can delay pages with analytics or streaming requests.networkidle2: allow up to two active connections. It is often a practical compromise for pages that keep background requests open.waitForSelector: wait for a stable element such as[data-testid="results"]ormain article. This is usually the most meaningful condition when you know what “ready” looks like.
Prefer a selector that represents the actual data, with a timeout as a safety limit. If a page can legitimately contain zero results, wait for a container or an explicit empty-state element rather than a result row that may never appear.
Windows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallCrashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minuteExtract one page as rendered HTML
Cloudflare’s content request is a POST to the account endpoint with the target URL and an API token. The following examples use environment variables so credentials do not end up in shell history or source control.
cURL
export ACCOUNT_ID="your-account-id"
export CF_API_TOKEN="your-api-token"
curl -X POST "https://api.cloudflare.com/client/v4/accounts/$ACCOUNT_ID/browser-run/content"
-H "Authorization: Bearer $CF_API_TOKEN"
-H "Content-Type: application/json"
--data '{"url":"https://example.com"}'
Save the JSON response, then read the returned HTML field according to the response schema for your account. Keep the requested URL alongside the result and record the fetch time.
Python
import os
import requests
account_id = os.environ["ACCOUNT_ID"]
token = os.environ["CF_API_TOKEN"]
endpoint = f"https://api.cloudflare.com/client/v4/accounts/{account_id}/browser-run/content"
response = requests.post(
endpoint,
headers={
"Authorization": f"Bearer {token}",
"Content-Type": "application/json",
},
json={"url": "https://example.com"},
timeout=90,
)
response.raise_for_status()
payload = response.json()
print(payload)
Node.js
const accountId = process.env.ACCOUNT_ID;
const token = process.env.CF_API_TOKEN;
const endpoint = `https://api.cloudflare.com/client/v4/accounts/${accountId}/browser-run/content`;
const res = await fetch(endpoint, {
method: 'POST',
headers: {
'Authorization': `Bearer ${token}`,
'Content-Type': 'application/json'
},
body: JSON.stringify({ url: 'https://example.com' })
});
if (!res.ok) throw new Error(`${res.status} ${await res.text()}`);
console.log(await res.json());
For repeatable extraction, store the raw response before cleaning it. That gives you a replayable audit record when a selector or JSON field changes.
Extract selected elements with CSS selectors
Use a scrape endpoint when the page is large but your output is small. Selectors such as article h1, .price, or [data-product-id] can return element text, attributes, inner HTML, and geometry instead of forcing you to parse an entire document. Cloudflare describes its scrape endpoint as returning structured details for specific elements, including dimensions and inner HTML.
Do these 3 things before closing this tab:
1Scan for outdated or missing drivers - takes under a minute2Repair Windows errors before they cause bigger problems3Fix the driver behind crashes, sound loss and screen glitchesMake selectors resilient
- Prefer semantic elements, stable IDs, or data attributes over classes generated by a design system.
- Return a small set of selectors and include the page URL with every record.
- Expect zero matches, multiple matches, and missing attributes; treat each as a normal outcome that your code validates.
- Keep a fixture of representative pages so a redesign fails a test instead of silently producing empty fields.
If your provider returns inner HTML, parse it with an HTML parser rather than regular expressions. Decode entities, normalize whitespace deliberately, and preserve the original fragment when later review may matter.
Extract JSON from a JavaScript-heavy page
Ask for a schema-shaped response
Cloudflare exposes JSON options with a prompt and response-format or schema controls. XCrawl also documents JSON output with a prompt and optional JSON schema. A useful schema makes the crawler’s contract explicit:
{
"type": "object",
"required": ["name", "price", "currency"],
"properties": {
"name": {"type": "string"},
"price": {"type": "number"},
"currency": {"type": "string"},
"availability": {"type": ["string", "null"]}
},
"additionalProperties": false
}
Phrase the extraction prompt narrowly: identify the product name, numeric price, currency, and availability shown on the page; return null when a field is absent; do not infer values from related pages. The exact JSON-option wrapper differs by API, so submit these controls using the provider’s current request schema.
Rank #3
Validate before writing to your database
- Parse the response as JSON and reject malformed output.
- Validate required keys, types, allowed values, and numeric ranges.
- Compare the extracted values with the source HTML or a selected-element result.
- Store the source URL, retrieval timestamp, and raw response with the normalized record.
- Send ambiguous records to a review queue instead of silently coercing them.
Schema guidance improves consistency; it does not make the result authoritative. Prices, stock labels, and dates can be absent, duplicated, or rendered differently for different visitors.
Crawl a whole site without losing control
A crawl starts with one URL and follows child pages. Cloudflare’s crawl API returns a job that you check separately rather than keeping one request open until every page finishes. Configure these controls:
- depth: how many link levels to follow from the starting URL.
- limit: the maximum number of pages to collect.
- source: discover URLs from
sitemaps,links, orall. - include/exclude patterns: constrain paths such as
/docs/and exclude account, search, or logout URLs. - formats: request
html,markdown, orjson. - rendering options: choose static or browser-rendered fetching and set waits appropriate to the site.
- JSON options: provide a prompt and response format or schema when requesting structured fields.
Start with a minimal crawl job
export ACCOUNT_ID="your-account-id"
export CF_API_TOKEN="your-api-token"
curl -X POST "https://api.cloudflare.com/client/v4/accounts/$ACCOUNT_ID/browser-rendering/crawl"
-H "Authorization: Bearer $CF_API_TOKEN"
-H "Content-Type: application/json"
--data '{"url":"https://example.com"}'
Use the job identifier and status procedure returned by the API to monitor completion. For a production crawl, add a conservative depth and page limit first, confirm the discovered URL set, then increase limits. Keep include and exclude rules under version control so a later run is explainable.
Choose discovery sources deliberately
Sitemaps are predictable for documentation and commerce catalogs. Link discovery follows what the site exposes in its rendered navigation. all can find more pages but also increases the chance of faceted URLs, calendars, and duplicate tracking parameters. Normalize URLs and de-duplicate them before downstream processing.
Why a scraper returns empty HTML
The page is an SPA shell
Symptom: you receive a root element with no products, comments, or article body. Fix: enable rendering and wait for networkidle0, networkidle2, or a known content selector.
The Tool Desk
Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Your wait condition is wrong
Symptom: the page is sometimes complete and sometimes partial. Fix: wait for the element that your extractor consumes, increase the timeout within the provider’s limits, and avoid relying only on a fixed sleep.
The data is in an underlying request
Symptom: browser HTML is expensive but the page calls a stable JSON endpoint. Fix: identify and reproduce that request directly, subject to authentication, terms, rate limits, and publisher controls.
A bot check or access boundary intervenes
Symptom: you capture a challenge page, login form, or denial instead of content. Fix: do not treat the challenge as valid data. Check permissions, cookies, authentication requirements, and the site’s rules. Cloudflare notes that changing the user agent does not bypass Browser Run bot identification.
Selectors changed
Symptom: the response is non-empty but fields are null. Fix: log match counts, retain sample HTML, test fallback selectors, and alert when required fields disappear.
Recommended Free Tools
Compliance, reliability, and cost controls
Respect publisher and legal boundaries
Check robots.txt, terms of service, authentication boundaries, rate limits, and applicable law before collecting data. Cloudflare’s crawl API exposes contentUse and crawlPurposes controls for publisher Content-Signal directives. Those controls do not establish a universal legal rule for every jurisdiction; document your own purpose and review process.
Best Value
Make runs reproducible
- Record URL, timestamp, rendering mode, wait condition, selector or schema version, HTTP status, and page verdict.
- Cache immutable pages and use a chosen TTL for changing pages.
- Retry transient network failures with exponential backoff, but do not retry authorization failures indefinitely.
- Set a per-page timeout and a crawl-wide budget so one hanging resource cannot consume the entire job.
- Hash raw HTML or normalized JSON to detect unchanged pages.
Estimate spend before scaling
Pricing, quotas, concurrency, timeout maxima, and cache behavior vary by vendor and plan. Verify the current plan documentation before committing to a volume. Measure pages discovered, rendered pages, retries, and bytes transferred; a low page count can still be expensive when every page launches a browser and loads many third-party resources.
Or skip the browser setup
If your goal is a clean visual capture rather than HTML or field extraction, ScreenshotNeo provides a single-call screenshot API. It accepts consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and each response reports the page verdict and billing status.
It also offers an MCP server for Claude, Cursor, and other MCP clients, with take_screenshot, get_page_info, and capture_pdf tools. Every plan includes the features; 1,000 screenshots per month are free with no card, Starter is $5 for 3,000, and paid plans start at $5.
One-call example
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
See the ScreenshotNeo API documentation for options such as full-page capture, element selectors, custom JavaScript, waiting for a selector or network idle, cookies and headers, PDFs, signed links, asynchronous jobs, and bulk capture. Create a free ScreenshotNeo account to get 1,000 screenshots each month without adding a card.
Practical decision checklist
- Can the server response provide the data? Use a direct request.
- Do scripts build the DOM? Enable rendering and wait for a meaningful selector.
- Do you need a few fields? Use selector extraction and validate match counts.
- Do you need typed records? Supply a strict prompt and schema, then compare with source content.
- Do you need linked pages? Start a bounded crawl with depth, limit, source, and include/exclude rules.
- Could a challenge, login wall, or consent layer contaminate results? Detect and classify it instead of storing it as content.
- Can another engineer explain the result later? Save the URL, settings, raw response, and extraction version.
Frequently Asked Questions
How can I tell whether a page is truly ready for extraction?
Wait for the element your parser needs, then verify that required fields have non-empty values. A completed navigation event alone is not proof that an SPA finished rendering.
Should I request HTML, Markdown, or JSON for an archive?
Keep raw rendered HTML when fidelity and later re-parsing matter. Use Markdown for readable text corpora, and JSON when you have a stable schema plus validation and source snapshots.
What should happen when a crawl finds a challenge or login page?
Classify it as an access failure, stop treating it as source content, and resolve authorization or site-policy requirements before retrying.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

