Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
A web scraping API lets your application request a public web page and receive HTML, text, rendered output, or structured JSON over HTTPS. The reliable pattern is: choose an endpoint, keep the key on your server, send the target URL in the provider’s expected parameter or JSON body, set connect and read timeouts, check the HTTP status, validate the response before parsing, and persist pagination checkpoints. The examples below show that workflow with REST, Python, PHP, cURL, and Node.js, including JavaScript pages, authentication, retries, rate limits, and provider selection.
What a web scraping API does
Instead of building and operating your own browser fleet, proxy pool, queue, and parser, you call an HTTPS endpoint. The service fetches the target URL and returns a response. Depending on the provider and plan, that response can be raw HTML, cleaned text, Markdown, a screenshot, or fields extracted into JSON. Browser-based providers can execute page JavaScript before returning content, while dataset services may run a predefined crawler and deliver a file or asynchronous job result.
API access does not bypass a site’s terms, robots directives, login controls, or applicable law. Use examples only for pages and data you are authorized to access, and avoid collecting personal or confidential information without a lawful basis.
PC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchChoose the request model before writing code
| Model | Best fit | Trade-off |
|---|---|---|
| Synchronous REST request | One page or a small batch where you need the result immediately | Your request remains open while the provider fetches and renders the page |
| Asynchronous job | Slow pages, large crawls, or bulk exports | You must poll a job or receive a webhook and persist job state |
| Actor or workflow API | Multi-step crawlers, datasets, and reusable automation | More concepts than a single endpoint: runs, datasets, pagination, and resource limits |
| Prebuilt site dataset | Catalogs or domains covered by a provider’s maintained connector | Less control over fields and update timing |
Compare services on JavaScript execution, proxy and anti-bot options, output format, synchronous versus asynchronous behavior, pagination, rate limits, retries, geographic targeting, and pricing. Apify centers its REST API on JSON, Actors, datasets, clients, and documented limits. ScrapingBee focuses on rendered pages, rotating proxy tiers, screenshots, and structured extraction. Bright Data emphasizes prebuilt site datasets and asynchronous bulk jobs.
#1 Best Overall
Secure REST authentication and request construction
Keep credentials in a server-side secret manager or an environment variable such as SCRAPER_API_KEY. Do not put a key in browser JavaScript, a mobile app, a public repository, or a URL that will be logged. When supported, use an HTTP header:
Authorization: Bearer YOUR_API_KEY
Apify recommends header authentication as more secure than a URL token, and ScrapingBee documents bearer authentication as its recommended method; ScrapingBee marks query-string API keys deprecated. Follow the selected provider’s exact endpoint, parameter names, and authentication rules.
For a GET endpoint, URL-encode the target URL. For a POST endpoint, send a JSON body. Always set a connect timeout separately from a read timeout, treat non-2xx responses as errors, and record a request ID or provider job ID for diagnosis.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Python: a production-safe synchronous request
import json
import os
import time
import random
import requests
endpoint = "https://api.example.com/v1/scrape"
api_key = os.environ["SCRAPER_API_KEY"]
session = requests.Session()
for attempt in range(5):
response = session.get(
endpoint,
params={"url": "https://example.com"},
headers={
"Authorization": f"Bearer {api_key}",
"Accept": "application/json",
},
timeout=(10, 60),
)
if response.status_code == 429 or 500 <= response.status_code < 600:
if attempt == 4:
response.raise_for_status()
retry_after = response.headers.get("Retry-After")
delay = float(retry_after) if retry_after else (2 ** attempt) + random.random()
time.sleep(min(delay, 60))
continue
response.raise_for_status()
content_type = response.headers.get("Content-Type", "")
if "json" in content_type.lower():
data = response.json()
else:
data = {"body": response.text}
print(json.dumps(data, ensure_ascii=False))
break
requests exposes query parameters, headers, JSON decoding, status checks, TLS verification, sessions, and connection pooling. The session is useful for repeated calls; the bounded retry loop handles transient 429 and 5xx responses without retrying a permanent 4xx error. Adapt the response fields to the provider. Some services return an object containing rendered HTML; others return an item URL, dataset ID, or job ID.
POST providers and JavaScript pages
If the provider accepts a POST payload, preserve the same timeout and status handling:
Rank #2
payload = {
"url": "https://example.com/products",
"render_js": True,
"wait_for": ".product-card"
}
response = session.post(
endpoint,
json=payload,
headers={"Authorization": f"Bearer {api_key}", "Accept": "application/json"},
timeout=(10, 90),
)
response.raise_for_status()
data = response.json()
JavaScript rendering adds browser startup and page-execution time. Request it only when the data is absent from the initial HTML. A selector wait is usually more reliable than an arbitrary long sleep; use a delay when the application has no stable selector, and use network-idle only when the provider documents its behavior.
PHP: portable cURL integration
<?php
$target = 'https://example.com';
$query = http_build_query(['url' => $target]);
$ch = curl_init('https://api.example.com/v1/scrape?' . $query);
curl_setopt_array($ch, [
CURLOPT_RETURNTRANSFER => true,
CURLOPT_HTTPHEADER => [
'Authorization: Bearer ' . getenv('SCRAPER_API_KEY'),
'Accept: application/json',
],
CURLOPT_CONNECTTIMEOUT => 10,
CURLOPT_TIMEOUT => 60,
]);
$body = curl_exec($ch);
if ($body === false) {
throw new RuntimeException(curl_error($ch));
}
$status = curl_getinfo($ch, CURLINFO_RESPONSE_CODE);
curl_close($ch);
if ($status < 200 || $status >= 300) {
throw new RuntimeException("Scraping API returned HTTP $status");
}
$data = json_decode($body, true, 512, JSON_THROW_ON_ERROR);
var_dump($data);
For a JSON POST, replace the URL query with CURLOPT_POST => true, set CURLOPT_POSTFIELDS => json_encode($payload, JSON_THROW_ON_ERROR), and add Content-Type: application/json. An official client can simplify pagination or dataset downloads; Apify documents a PHP client option, while ScrapingBee publishes PHP cURL examples.
cURL for diagnosis and automation
curl --fail-with-body --connect-timeout 10 --max-time 60
-H "Authorization: Bearer $SCRAPER_API_KEY"
-H "Accept: application/json"
--get "https://api.example.com/v1/scrape"
--data-urlencode "url=https://example.com"
--fail-with-body keeps an error payload available while returning a failing exit status. Add -D headers.txt when you need to inspect rate-limit or request-ID headers.
Node.js equivalent
const endpoint = 'https://api.example.com/v1/scrape';
const target = new URLSearchParams({ url: 'https://example.com' });
const controller = new AbortController();
const timer = setTimeout(() => controller.abort(), 60000);
try {
const res = await fetch(`${endpoint}?${target}`, {
headers: {
Authorization: `Bearer ${process.env.SCRAPER_API_KEY}`,
Accept: 'application/json'
},
signal: controller.signal
});
const text = await res.text();
if (!res.ok) throw new Error(`HTTP ${res.status}: ${text}`);
const data = res.headers.get('content-type')?.includes('json') ? JSON.parse(text) : text;
console.log(data);
} finally {
clearTimeout(timer);
}
Pagination, checkpoints, and 429 handling
Providers expose pagination differently: a numeric page, offset and limit, a cursor token, a dataset item cursor, or a next URL. Never assume that page numbers are interchangeable. Persist the last successful cursor (and the target or job ID) before requesting the next page. On restart, resume from that checkpoint instead of duplicating earlier items.
For HTTP 429, stop adding concurrency, inspect Retry-After and provider-specific rate headers, then retry with exponential backoff and jitter. A practical sequence is approximately 1, 2, 4, 8, and 16 seconds, capped at a safe maximum and bounded by a total retry count. Apify documents a global limit of 250,000 requests per minute and a default per-resource limit of 60 requests per second in its API v2 reference; these are provider-specific values that can change, so use the limits returned by your account and endpoint rather than hard-coding them.
Rank #3
Deduplicate by a stable source ID or canonical URL. Store raw responses when permitted, plus extraction time, provider, status, and parser version. That makes parser changes and partial failures recoverable.
Recommended Free Tools
Output validation and parsing
- Check the status code before decoding.
- Inspect
Content-Type; a proxy error page may be HTML even when you expected JSON. - Validate required fields and distinguish an empty result from a failed extraction.
- Preserve the original HTML or text when downstream parsing is important and retention is lawful.
- Normalize dates, currencies, whitespace, and character encoding only after confirming the source format.
Do not treat a HTTP 200 response as proof that the page was successfully scraped. Bot challenges, consent walls, login pages, and application errors can all arrive with a successful transport status.
Provider selection by workload
| Requirement | Questions to ask |
|---|---|
| JavaScript application | Does the service run a real browser, support a wait condition, and expose console or navigation errors? |
| Anti-bot and geography | Are residential or premium proxies available, in which countries, and at what credit cost? |
| Structured extraction | Can you define fields, receive JSON, and detect missing fields rather than silently returning an empty object? |
| Large crawl | Is there an asynchronous queue, bulk endpoint, dataset export, webhook, or resumable cursor? |
| Budget | Is billing per request, browser time, proxy tier, credit, dataset item, or job? |
ScrapingBee’s documented credit examples list rotating proxy without JavaScript at 1 credit, rotating proxy with JavaScript at 5, premium proxy without JavaScript at 10, premium proxy with JavaScript at 25, and stealth proxy with JavaScript at 75; verify current pricing before committing. Bright Data’s dataset model can be more suitable than page-by-page requests when a maintained site connector matches your fields. Apify’s Actor and dataset model is useful when you need reusable workflows and pagination controls.
Common failures and fixes
401 or 403
Check the environment variable, bearer prefix, account permissions, and endpoint-specific scope. Rotate an exposed key; do not print it in logs.
400 or 422
Confirm the parameter name, URL encoding, required JSON fields, and supported rendering options. Compare the request with the provider’s current schema.
Do these 3 things before closing this tab:
1Scan for outdated or missing drivers - takes under a minute2Clear out junk files and repair common Windows errors3Fix the driver behind crashes, sound loss and screen glitchesRank #4
429
Reduce concurrency, honor rate headers, back off with jitter, and resume from a checkpoint rather than replaying a whole batch.
Timeout or empty HTML
Increase the read timeout only after checking whether the target is slow, blocked, or waiting for JavaScript. Add a documented selector wait, choose an asynchronous job, or disable rendering when the initial HTML already contains the data.
JSON decoding error
Log status, content type, and a redacted prefix of the body. Many failures are HTML challenge pages or gateway errors returned where JSON was expected.
Duplicate or missing records
Persist cursors atomically, use idempotent writes keyed by a source identifier, and record the exact request and parser version for each batch.
Or skip the browser setup
When your goal is a dependable screenshot rather than parsed fields, ScreenshotNeo provides a single-call website screenshot API. It accepts consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be disabled. Only clean shots are billed: bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers identify the page verdict and billing result. Its MCP server gives Claude, Cursor, and other MCP clients take_screenshot, get_page_info, and capture_pdf tools.
See the ScreenshotNeo API documentation for all options, including full-page lazy-image loading, CSS-selector element capture, dark mode, device presets, retina scale, PDF controls, custom CSS and JavaScript, clicks, waits, blocking rules, headers, cookies, user agent, authorization, timezone, geolocation, transparent backgrounds, resizing, TTL caching, signed image links, asynchronous webhooks, bulk capture of up to 100 URLs per call, usage data, and OpenAPI. Every feature is included on every plan. The Free plan includes 1,000 shots per month with no card; paid plans start at $5 for 3,000.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Create a free ScreenshotNeo account to get 1,000 screenshots a month without a card.
Operational checklist
- Use an environment variable or secret manager for the key.
- Set connect and read timeouts appropriate to rendering mode.
- Validate status, content type, and required fields.
- Implement bounded retries only for 429 and transient 5xx responses.
- Honor provider rate headers and persist pagination checkpoints.
- Measure request latency, empty-result rate, extraction errors, and billed units.
- Review target-site permissions, terms, robots rules, and retention obligations.
Frequently Asked Questions
Should I use GET or POST for a scraping API?
Use the method required by the provider. GET is convenient for a URL and a few options; POST is better for long payloads, extraction schemas, cookies, or job configuration.
When should a scrape become an asynchronous job?
Choose asynchronous execution when rendering is slow, the crawl is large, or you need bulk output and restartable processing instead of holding one HTTP request open.
Can a scraping API legally access any public page?
No. Public visibility does not remove terms of service, robots directives, authentication boundaries, privacy duties, or other applicable legal restrictions.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

