Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

To scrape a website with an API, first check whether the site offers an authorized data API. If it does, request the relevant endpoint and process its structured response. If not, use a managed scraping API to fetch HTML or render JavaScript pages. In either case, confirm you are allowed to collect the data, keep credentials on your server, validate every response, and control request volume.

Choose the right way to get the data

“Scraping with an API” can mean two different things: calling an API that the website itself exposes, or sending a page URL to a third-party scraping API that retrieves the page for you. The direct site API is usually cleaner when it is available and permitted: it can return structured data such as JSON instead of markup you must parse and maintain. Apify describes API scraping as finding a website’s API endpoints and fetching the data directly rather than parsing rendered HTML; complex endpoints may still require special headers, payloads, rate-limit handling, encoded-response processing, or GraphQL knowledge. Apify’s guide to scraping through APIs explains the approach.

Use a managed scraping API when you need a service to fetch pages, render JavaScript, handle proxy infrastructure, return structured extraction, or support a larger pipeline. These services do not make access restrictions disappear: you remain responsible for the target site’s rules and the data you collect.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Use a site’s direct API when its documentation covers the data you need and your use is authorized.
  • Use an HTML-fetching API when the desired data is present in the server-returned page markup.
  • Use browser rendering or an extractor when the data appears only after JavaScript runs, or when maintaining your own selectors would be fragile.

Check permission, robots.txt, and scope first

Before sending requests, review the site’s terms, API documentation, authentication requirements, and applicable privacy or data-use rules. Identify which pages and fields you need, how often you need them, and whether the site provides an official endpoint.

Also inspect the site’s robots.txt. The IETF’s Robots Exclusion Protocol standard, RFC 9309, published in September 2022, says that after a successful fetch, crawlers must follow parseable rules. It also makes the boundary clear: “These rules are not a form of access authorization.” RFC 9309 is crawler guidance, not permission to access protected data and not a replacement for authentication. If a site requires login, an API key, or another authorization, use it only when you have legitimate access. Stop if the site denies access or repeatedly blocks your requests.

Build a small, safe API request

Start with one URL or one documented endpoint. Store your API key in a server-side environment variable or secret manager; do not put it in browser JavaScript, a public repository, or a URL that may be logged. The following generic Python example calls a site endpoint. Replace the example URL and parameter names with those in the target site’s documentation. It intentionally does not guess a real website’s API schema.

import os
import requests

API_URL = "https://example.com/api/items"
API_KEY = os.environ["SITE_API_KEY"]

response = requests.get(
    API_URL,
    headers={"Authorization": f"Bearer {API_KEY}"},
    params={"page": 1},
    timeout=30,
)
response.raise_for_status()

content_type = response.headers.get("Content-Type", "")
if "application/json" not in content_type.lower():
    raise ValueError(f"Expected JSON, got {content_type!r}")

data = response.json()
if not isinstance(data, dict) or "items" not in data:
    raise ValueError("Response does not match the expected schema")

for item in data["items"]:
    print(item)

Set the secret before running the script, for example with export SITE_API_KEY='…' in a private shell. Use the authentication scheme and request parameters specified by the site; bearer authentication is only an example. If the endpoint is public and requires no credential, remove the authorization header rather than inventing one.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Check the response before storing it

  • Check the HTTP status code before parsing the response body. A non-2xx response may contain an error message rather than the expected data.
  • Check the content type and validate the response shape, required fields, and field types.
  • Follow documented pagination rather than assuming the first response contains every record.
  • Normalize values and record enough non-secret request metadata to reproduce a run, such as the endpoint, page number, and retrieval time.
  • Make writes idempotent where possible, so retrying a page does not create duplicate records.

Apify’s REST API documentation describes resource-oriented URLs, JSON responses, standard HTTP status codes, and bearer-token authentication. Its platform also provides official JavaScript and Python clients. Use a provider’s documented client when it fits your workflow; otherwise, use a well-tested HTTP library and validate its response as above. Apify API documentation and Apify integrations describe its API and integration options.

When to use a managed scraping API

A managed service accepts a target URL or job request and performs some of the retrieval work for you. The exact output and controls vary by provider, so confirm whether you receive raw HTML, rendered HTML, extracted fields, or a job result. Do not assume every service supports every site or bypasses every access control.

ScraperAPI

ScraperAPI documents a simple authenticated request that sends a URL and returns that page’s HTML. It also documents API, asynchronous, proxy, structured-data, and DataPipeline interfaces, along with JavaScript rendering and JSON-parsing controls. Choose the controls your target actually requires; JavaScript rendering can add work and should not be enabled when the page’s useful content is already available in the returned HTML. See the ScraperAPI documentation and request-control documentation.

Apify

Apify is suited to workflows organized around Actors and platform facilities such as storage, proxies, schedules, integrations, and monitoring. An Actor’s input and output depend on that Actor; inspect its documentation and access conditions before relying on it for a recurring collection. See Apify platform documentation.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Bright Data Web Scraper API

Bright Data documents prebuilt scrapers for more than 100 popular websites in its Web Scraper API reference, with URL or keyword inputs and JSON, NDJSON, or CSV output. The same reference describes bearer authentication and synchronous or asynchronous jobs; it says synchronous jobs are for smaller real-time requests and asynchronous jobs for larger batches. That is a documented product distinction, not an independent performance comparison. See the Bright Data Web Scraper API reference.

Compare capabilities, not just the word “API”

Option What the documentation describes Evaluate before choosing
Direct site API Site-defined endpoints and response formats; often structured data. Authorization, endpoint stability, schema, pagination, and rate limits.
ScraperAPI URL-to-HTML requests plus documented rendering, proxy, structured-data, and asynchronous controls. Whether you need rendering or extraction, request constraints, and your expected volume.
Apify REST API and Actors, with platform documentation for storage, proxies, schedules, integrations, and monitoring. Which Actor fits, how its input/output works, and what platform components your pipeline needs.
Bright Data Web Scraper API Documented prebuilt scrapers, multiple output formats, and sync or async jobs. Whether the relevant prebuilt scraper supports your target and the job mode fits your workflow.

These descriptions reflect the linked provider documentation, not a claim that one service is universally more accurate, faster, or cheaper. The cited documentation does not establish an independent benchmark or a comparable price for these options. Compare current plan terms and limits directly before committing.

Handle JavaScript-rendered pages deliberately

A page that loads in a browser may not contain its final data in the initial HTML. The browser can run scripts that fetch data later or build the visible interface client-side. First check whether the underlying data comes from a documented, authorized site endpoint. If so, calling that endpoint may be simpler than rendering the entire page. If there is no suitable permitted endpoint, choose a scraping service with documented JavaScript rendering or browser execution.

Rendering is not a universal fix. A page may still need cookies, authentication, a particular interaction, or time for a request to finish. Selectors can change as a site changes. Test a small sample and confirm that the fields you need appear in the returned result before scaling up. If the service provides a predefined extractor or dataset for the target, compare its output and maintenance requirements with a custom selector-based approach.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Make collection reliable without overloading the site

Run a bounded, observable process rather than launching an unrestrained set of requests. Respect documented limits and crawler instructions, and do not keep retrying authorization failures or blocks. For transient failures, use a limited retry policy with exponential backoff and jitter. For example, wait roughly 1, 2, and 4 seconds between a maximum of three retries, adding a small random delay; treat those values as starting settings to tune to the service’s published limits, not universal limits.

  • Bound concurrency: set a fixed number of simultaneous requests and lower it if errors or latency rise.
  • Cache repeat work: retain successful results for an appropriate period instead of fetching unchanged pages repeatedly.
  • Checkpoint pagination: save the last completed page or cursor so a failed run can resume without starting over.
  • Make retries safe: use idempotent writes and avoid retrying requests that create side effects unless the API supports idempotency.
  • Monitor data quality: track missing fields, duplicate records, schema changes, latency, and failure rates.
  • Protect secrets and personal data: avoid logging keys, tokens, or unnecessary sensitive page content.

For bulk or recurring jobs, choose a synchronous request when the result must be returned immediately and the workload is small enough for that mode. Use an asynchronous job flow when the provider documents it for larger batches; poll or receive completion according to the provider’s instructions, and persist job identifiers so you can recover after a client disconnect. Scheduling, storage, and monitoring may be built into a platform, as Apify documents, but confirm the selected Actor and plan support the parts your workflow needs.

Or skip the browser setup

If your goal is a clean screenshot or PDF rather than extracting structured records, a screenshot API is a different tool from a web-scraping API. ScreenshotNeo returns a screenshot or PDF from one GET request. It accepts cookie or consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be turned off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and responses identify the page verdict and billing status in headers. Its MCP server offers take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients. The free plan includes 1,000 shots per month without a card; paid plans start at $5 for 3,000 shots.

Example cURL request (replace the target URL and put your key in a secure environment before using it):

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

See the ScreenshotNeo API documentation for request parameters and response details. This captures a visual page; it does not extract page text into structured records. Sign up for 1,000 free screenshots a month with no card.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Troubleshoot common failures

401 or 403 response

The key may be missing, invalid, expired, or sent using the wrong authentication scheme; the endpoint may also deny your account or requested resource. Check the site’s or provider’s authentication documentation, verify the secret is available to the server process, and confirm your authorization. Do not try to evade a denial by changing identities or repeatedly resending the request.

429 or repeated blocking

You may be sending requests too quickly or exceeding an account or site limit. Reduce concurrency, respect any documented retry guidance, and use backoff for transient rate limits. If blocks persist, stop and check the site’s rules or ask the provider or site owner about permitted access.

200 response but no useful data

A successful HTTP status does not prove the expected records are present. The response could be an error payload, an empty result, a login page, or HTML where you expected JSON. Check the content type, response schema, required fields, and authentication state before saving data. For browser-rendered pages, confirm whether the data appears only after scripts or a user action.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Missing fields or broken selectors

The site may have changed its schema or markup, or the selected endpoint may not return every field. Compare a current response with your expected schema, log missing-field counts, and update selectors or endpoint handling only after verifying the new structure. Prefer a stable documented API over markup parsing when one is available and permitted.

Timeouts, partial runs, or duplicates

Use a finite timeout, a bounded retry policy for transient failures, and pagination checkpoints. Store records idempotently and retain non-secret job or cursor metadata. For long-running work, use an asynchronous job feature when the provider documents one for that workload, rather than holding a synchronous request open indefinitely.

Plan for maintenance and total cost

Direct APIs avoid much HTML selector maintenance but may have endpoint, schema, authentication, or pagination requirements of their own. HTML and browser-based scraping can require ongoing selector checks as a site changes. Managed services can reduce infrastructure work, but cost depends on the provider’s current plans, usage units, rendering or extraction options, and job volume. The provider documentation linked above does not establish a comparable total cost or independent accuracy and performance figures, so estimate from your real request pattern and check current terms before choosing.

Start by measuring your own workload: pages per run, runs per month, average retries, whether rendering is needed, expected output size, and how long results must be stored. Include engineering time for validation and recovery, not only the API bill. A small pilot with representative pages will expose schema gaps and workflow constraints before you schedule a full collection.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Frequently Asked Questions

Is scraping a website’s API better than parsing its HTML?

Usually, if the endpoint is documented or otherwise authorized and returns the fields you need. Structured responses avoid much of the selector maintenance required by HTML parsing, though the endpoint can still have authentication, pagination, or schema constraints.

Does robots.txt give permission to scrape a page?

No. RFC 9309 treats robots.txt as crawler instructions and explicitly says its rules are not access authorization. You still need legitimate permission and any required authentication.

Should I use a scraping API or a screenshot API?

Use a scraping API when you need page content or structured records. Use a screenshot API when the desired output is a visual image or PDF; a screenshot does not replace data extraction.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.