To modify a web scrape with an API, change both sides of the pipeline: the request you send and the code that interprets the response. Add the API’s documented authentication, URL, parameters, headers, body, rendering and session settings; then parse the returned JSON or HTML, normalize it into your application’s schema, follow pagination or cursors, validate records and save them.
The safest process is contract-first: identify whether you are calling a documented data endpoint, a rendered-page endpoint or a hosted scraper platform, then implement its exact response and rate-limit rules. The examples below show a reusable implementation, pagination, retries, asynchronous jobs and a browser-free screenshot option.
Start by identifying what kind of API you are adding
“Modify a scrape with an API” can mean three different changes. Decide which one applies before changing selectors or request code.
A documented data API
A data API normally returns structured JSON. You replace an HTML request and selector logic with an authenticated request, inspect the records array and map fields by name. Pagination, authentication and error objects are explicit parts of the contract, so this is usually the most stable option when the publisher provides it.
The Tool Desk
Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →#1 Best Overall
A rendered-page API
A rendered-page service fetches a target URL, optionally runs JavaScript and returns the resulting HTML or a page image. This is useful when content appears only after client-side rendering. The request may include a target URL, custom headers, cookies, a user agent, proxy or country settings and a JavaScript-rendering flag. Those parameters are provider-specific; do not assume that selectors or option names from an HTML scraper work unchanged.
A hosted scraper platform
A hosted platform can expose a tool catalog, a synchronous run endpoint, an asynchronous job endpoint, a status endpoint and a dataset-export endpoint. Scrapy.io documents a run, poll and dataset pattern, while other services may return one response per URL. You trade browser, proxy, CAPTCHA, scheduling and storage work for provider-specific quotas, schemas and billing rules.
Use the endpoint contract as your checklist
- Write down the endpoint and method. Record whether it is GET or POST, the required URL or body fields, content type and maximum request size.
- List authentication requirements. Prefer the documented
Authorization: Bearer ...header or API-key header. Never put a production secret in browser JavaScript, a public repository or a URL that third parties can copy. - Separate target inputs from scraper controls. A target URL, query string and POST body describe the site request. Rendering, proxy, country, session, timeout and wait settings describe how the scraper service should fetch it.
- Describe the response before writing a parser. Identify the records array, nested objects, status or error object, request identifier and any
nextlink, cursor, offset, limit or total value. - Record limits. Note quota, concurrency, maximum page size, timeout and retry guidance. A correct parser can still fail if it sends requests faster than the service permits.
Runnable baseline: request, parse and persist
The following examples use a placeholder endpoint. Replace API_URL, field names and parameters with the target service’s documented values. They implement bearer authentication, offset pagination, bounded retries and a stable output file.
cURL: inspect one response first
curl --fail-with-body --retry 2
-H 'Authorization: Bearer YOUR_API_TOKEN'
-H 'Accept: application/json'
'https://api.example.com/v1/items?offset=0&limit=100'
Save a representative response before writing transformation code. Check whether records are under items, data or another property, and preserve the raw response while you develop.
Quick wins for a faster PC:
Scan for outdated or missing drivers - takes under a minuteDriver Scan →Clear out junk files and repair common Windows errorsFree Scan →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Python: paginated scraper with validation
import json
import os
import time
from decimal import Decimal, InvalidOperation
import requests
API_URL = os.environ['API_URL']
TOKEN = os.environ['API_TOKEN']
PAGE_SIZE = 100
MAX_RETRIES = 4
session = requests.Session()
session.headers.update({
'Authorization': f'Bearer {TOKEN}',
'Accept': 'application/json',
})
def get_page(offset):
for attempt in range(MAX_RETRIES):
response = session.get(
API_URL,
params={'offset': offset, 'limit': PAGE_SIZE},
timeout=45,
)
if response.status_code == 429 or response.status_code >= 500:
if attempt == MAX_RETRIES - 1:
response.raise_for_status()
retry_after = response.headers.get('Retry-After')
delay = float(retry_after) if retry_after else min(2 ** attempt, 16)
time.sleep(delay)
continue
response.raise_for_status()
return response.json()
raise RuntimeError('unreachable')
def normalize(record):
item_id = record.get('id')
name = record.get('name')
if item_id is None or not isinstance(name, str) or not name.strip():
return None
price = record.get('price')
if price is not None:
try:
price = str(Decimal(str(price)))
except (InvalidOperation, ValueError):
return None
return {'id': str(item_id), 'name': name.strip(), 'price': price}
results = []
offset = 0
seen = set()
while True:
payload = get_page(offset)
records = payload.get('items', [])
if not records:
break
for record in records:
item = normalize(record)
if item and item['id'] not in seen:
seen.add(item['id'])
results.append(item)
total = payload.get('total')
offset += len(records)
if total is not None and offset >= total:
break
with open('items.json', 'w', encoding='utf-8') as output:
json.dump(results, output, ensure_ascii=False, indent=2)
Run it with API_URL=https://api.example.com/v1/items API_TOKEN=... python scrape.py. The parser rejects records without a stable ID or usable name, converts IDs to strings and keeps decimal values from being silently rounded by binary floating point.
Node.js: the same pattern with fetch
const API_URL = process.env.API_URL;
const TOKEN = process.env.API_TOKEN;
const pageSize = 100;
const records = [];
const seen = new Set();
async function getPage(offset) {
for (let attempt = 0; attempt < 4; attempt++) {
const url = new URL(API_URL);
url.searchParams.set('offset', offset);
url.searchParams.set('limit', pageSize);
const response = await fetch(url, {
headers: {
Authorization: `Bearer ${TOKEN}`,
Accept: 'application/json'
},
signal: AbortSignal.timeout(45000)
});
if (response.status === 429 || response.status >= 500) {
if (attempt === 3) throw new Error(`HTTP ${response.status}`);
const retryAfter = Number(response.headers.get('retry-after'));
const delay = Number.isFinite(retryAfter) ? retryAfter * 1000 : Math.min(2 ** attempt, 16) * 1000;
await new Promise(resolve => setTimeout(resolve, delay));
continue;
}
if (!response.ok) throw new Error(`HTTP ${response.status}: ${await response.text()}`);
return response.json();
}
}
let offset = 0;
while (true) {
const payload = await getPage(offset);
const page = Array.isArray(payload.items) ? payload.items : [];
if (page.length === 0) break;
for (const row of page) {
if (row?.id != null && typeof row.name === 'string' && row.name.trim()) {
const id = String(row.id);
if (!seen.has(id)) {
seen.add(id);
records.push({ id, name: row.name.trim(), price: row.price ?? null });
}
}
}
offset += page.length;
if (payload.total != null && offset >= payload.total) break;
}
console.log(JSON.stringify(records, null, 2));
Implement pagination deliberately
Offset and limit
Some APIs return items, total, offset and limit. Start at offset zero, request a bounded page size, advance by the number of records actually returned and stop when the page is empty or the offset reaches total. Advancing by the requested limit can skip records when the final page is short.
Rank #2
- HTML CSS Design and Build Web Sites
- Comes with secure packaging
- It can be a gift option
Cursor or next-link pagination
When the response contains a cursor or next URL, send that value exactly as returned. Do not manufacture a page number from a cursor; cursors can encode a sort position or expire. Keep a maximum-page guard and stop if the service repeats the same cursor, which prevents an accidental infinite loop.
Changing data during a crawl
Offset pagination can duplicate or miss records when rows are inserted or deleted while you crawl. Prefer a documented cursor or a stable sort key such as an update timestamp plus ID. If the API offers a snapshot token, retain it for the complete run.
Hosted asynchronous runs: create, poll and export
For long jobs, a hosted service may not return records immediately. The general sequence is:
- Call the tool or actor catalog if the platform requires a tool identifier.
- Create a run with the target URL, authentication and scraper options.
- Store the returned run ID and request ID.
- Poll the status endpoint at an increasing interval until it reports success or failure.
- Fetch the dataset or export endpoint, then apply the same validation and deduplication used for synchronous responses.
Keep polling bounded by a deadline. On failure, log the provider’s error object and run ID rather than retrying blindly; a non-idempotent run may create duplicate work. If the platform supports webhooks, verify the signature and make the handler idempotent before accepting completion notifications.
Transform, validate and store the result
- Normalize names. Map provider-specific names such as
product_titleandtitleinto one application field. - Normalize types. Parse dates with an explicit timezone, represent money as decimal values or integer minor units, and convert IDs to one consistent type.
- Reject malformed records. Send invalid rows to a quarantine file with the source URL and request ID instead of silently dropping them.
- Deduplicate. Use a stable key supplied by the API. If none exists, combine documented fields and record the limitation.
- Version your raw data. Store the original response, retrieval time, endpoint version and parser version so a schema change can be diagnosed and reprocessed.
Rate limits, retries and reliability
Read the service’s quota and concurrency documentation before choosing worker counts. The api.data.gov documentation says its participating services have a default limit of 1,000 requests per hour, with service-specific variation; exceeding a limit returns HTTP 429. Treat 429 as a command to slow down, not as a permanent parsing failure.
Retry only transient conditions: 429, connection resets, timeouts and selected 5xx responses. Honor Retry-After when present, otherwise use bounded exponential backoff with jitter. Do not retry most 400-series errors until the request is corrected. Use idempotency keys for APIs that support them.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Rank #3
- Brand: Wiley
- Set of 2 Volumes
- A handy two-book set that uniquely combines related technologies Highly visual format and accessible language makes these books highly effective learning tools Perfect for beginning web designers and front-end developers
ScraperAPI’s FAQ describes typical latency of roughly 4–12 seconds and says some requests can take up to 60 seconds. That is the vendor’s operational guidance, not an independent benchmark, so set timeouts and queue sizes for your own workload. WebScraping.AI documents an “80%+ success rate” claim for most websites; it is a provider claim, not a universal guarantee.
API scraping versus HTML scraping
| Approach | Best fit | Responsibilities you still own |
|---|---|---|
| Documented JSON API | Stable records, explicit authentication and predictable pagination | Credentials, schema mapping, pagination, quota handling and data permissions |
| Rendered-page API | JavaScript-created content or pages with no usable public data endpoint | Wait conditions, selectors, rendering cost, session state and page changes |
| Hosted scraper platform | Teams that want managed browsers, proxies, CAPTCHA handling, scheduling or exports | Provider schema, credits, concurrency, supported targets and platform-specific failures |
ScraperAPI documents JavaScript rendering and proxy options. WebScraping.AI documents JavaScript execution, custom headers and target-URL parameters. These controls help with dynamic pages, but they do not turn an undocumented or restricted site into an automatically permitted data source.
Common failures and fixes
401 or 403 responses
Confirm the header name, token format, account scope and target environment. Check that a proxy, country or user-agent option has not invalidated the credential. Never “fix” a 403 by exposing the secret in client-side code.
200 response with no records
Inspect the raw body. You may be reading data while the service uses items, receiving an error object with HTTP 200, or requesting a page before a JavaScript-rendered result is ready. Log the response content type and request ID.
Recommended Free Tools
429 Too Many Requests
Reduce concurrency, increase the delay between pages and honor rate-limit or Retry-After headers. Check whether the quota is per key, IP address, account or endpoint.
Repeated or missing pages
Verify that the cursor is passed unchanged and that offset is advanced by the actual page length. Use a stable sort or snapshot when records can change during the run, and stop if a cursor repeats.
Rank #4
Timeouts on dynamic pages
Increase the documented timeout only within the provider’s maximum, wait for a meaningful selector or network-idle condition and block unnecessary resources when supported. If the content is available through a documented JSON endpoint, use that instead of rendering the whole page.
Parser breaks after a schema change
Compare the stored raw response with the last known fixture, make optional fields nullable, fail loudly on removed required fields and deploy parser changes separately from crawl scheduling.
Legal and operational checks
An API’s technical accessibility does not establish permission to collect or republish a site’s content. Review the target’s terms, robots guidance, authentication policy, privacy obligations and applicable law. Minimize personal data, protect credentials, restrict access to raw responses and define a retention period. Test against saved fixtures that include empty pages, missing fields, 401, 403, 429 and 5xx responses before production; no live integration result should be assumed from sample code.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Or skip the browser setup:
ScreenshotNeo is a website screenshot API and MCP server for developers. One GET request returns a PNG, JPEG, WebP or PDF, with browser work handled by the service. Its clean-shot flow accepts cookie and consent banners like a visitor, then removes more than 60 known consent platforms, newsletter popups and chat widgets; each step can be switched off.
One-call example
See the ScreenshotNeo API documentation for all parameters.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Only clean shots are billed. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads and cache hits cost nothing, and the response identifies the result with X-Page-Verdict and X-Billed headers. An MCP server provides take_screenshot, get_page_info and capture_pdf tools to Claude, Cursor and other MCP clients.
Python
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
Node.js
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
ScreenshotNeo includes full-page capture with lazy images loaded, CSS-selector element capture, dark mode, 12 device presets plus custom viewports, retina scale, PDF paper sizes, margins, landscape mode and page ranges. You can also supply custom CSS or JavaScript, click an element before capture, hide selectors, wait for a selector, delay or network idle, block ads, trackers, requests or resource types, set headers, cookies, user agent, Authorization, timezone and geolocation, use a transparent background, resize images, choose a cache TTL, create signed public-image links, submit asynchronous jobs with signed webhooks, capture up to 100 URLs per bulk call, query usage and use the OpenAPI specification. Parameter names used by other screenshot APIs also work to ease migration.
Best Value
| Plan | Included screenshots | Price |
|---|---|---|
| Free | 1,000 per month | $0, no card |
| Starter | 3,000 | $5 |
| Growth | 15,000 | $15 |
| Pro | 60,000 | $39 |
| Scale | 250,000 | $99 |
| Business | 1,000,000 | $249 |
Yearly billing provides two months free, and every feature is available on every plan. Create a free ScreenshotNeo account to use 1,000 screenshots a month without adding a card; paid plans start at $5 for 3,000 screenshots.
FAQ
Should I parse JSON or rendered HTML?
Use a documented JSON endpoint when it contains the fields you need. Choose rendered HTML only when the data is created in the browser or no suitable data endpoint exists; rendering adds wait, session and page-change concerns.
How do I know whether a retry is safe?
Retry read-only GET requests and documented idempotent operations after transient failures. For writes or hosted runs that create jobs, use an idempotency key or inspect the run status before submitting again.
Free tools Windows power users keep installed
One-click scans. No signup required.
What should I log for every record?
Keep the retrieval timestamp, source URL, endpoint or API version, request or run ID, parser version and a stable record key. Those values let you trace a bad row without storing more personal data than necessary.
Frequently Asked Questions
Can an API replace every HTML scraper?
No. It replaces selector-based scraping only when the endpoint exposes the required data or the service can reliably render the page. Otherwise, you still need page-specific extraction logic.
Is a 429 an authentication error?
No. HTTP 429 means the service is throttling requests. Slow the client, honor rate-limit headers and verify the account quota.
Why keep raw API responses?
Raw fixtures make schema changes reproducible, allow parser updates without recrawling and provide evidence when a provider returns an error object with a successful HTTP status.
Do these 3 things before closing this tab:
1Fix the driver behind crashes, sound loss and screen glitches2Repair Windows errors before they cause bigger problems3Scan for outdated or missing drivers - takes under a minuteQuick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

