Hardware FixRecommendedDevice not working? Your driver may be the problemCheck updates for common hardware issues.Fix DriversOctober DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsClean PCRecommendedOne scan can reveal what keeps slowing WindowsLook for cleanup and repair opportunities.Run Scan×
Skip to content
Bright Data

Libraries and SDKs for Web Scraping APIs: A Practical 2026 Guide

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The easiest way to integrate a web-scraping API is usually a plain HTTPS request from the language your application already uses. Choose a vendor that can render JavaScript when necessary, then add a small client wrapper for authentication, retries, logging and schema validation. Hosted APIs remove proxy rotation, browser maintenance and much of the anti-bot work; the trade-off is usage-based pricing and dependence on a provider’s limits and extraction model.

This guide compares Oxylabs, Zyte, ScraperAPI and Bright Data, explains when an SDK or browser library is a better fit, and provides implementation patterns in cURL, Python and Node.js.

Start with the integration model, not the SDK brand

Most scraping APIs expose HTTPS endpoints, so an SDK is optional. A language-native library is useful when it adds retries, typed responses, pagination helpers, asynchronous jobs or framework integration. Before selecting one, answer four questions:

  • What does the target need? Static HTML can be fetched with a normal HTTP client. JavaScript applications, infinite scroll and client-side APIs need a renderer or headless browser.
  • What output do you need? Providers may return raw HTML, parsed fields, JSON, Markdown, documents or files. Structured extraction is convenient, but custom schemas and selectors give you more control.
  • How much access handling is required? Proxy rotation, geographic IPs, CAPTCHA or ban handling, retries and browser fingerprints differ substantially by vendor.
  • How will you measure cost? Billing can be per successful result, per request, by API credit, by bandwidth or by subscription. Rendering and premium domains often consume more than a basic request.

For a first proof of concept, use the smallest request that demonstrates access and extraction quality. Only then add concurrency, browser actions or bulk jobs.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What hosted scraping APIs take care of

A hosted service typically runs the request through its proxy network, retries transient failures and returns the page or extracted data to your application. This avoids operating a fleet of browsers and proxy sessions yourself. It also means your application must understand the provider’s definition of success: one vendor may bill a successful content entity, another a request or a variable number of credits.

Hosted access is particularly valuable when targets change anti-bot behavior or when you need several countries. It is less attractive when you require a highly customized browser workflow, strict on-premises processing or a cost model that is predictable per byte rather than per result.

Capabilities to verify before purchase

  • JavaScript execution and browser actions, including clicks, scrolling and waiting for selectors.
  • Proxy geography, session persistence and rotation controls.
  • CAPTCHA or ban handling, retry policy and timeout limits.
  • Raw response access in addition to parsed or structured fields.
  • Concurrency, rate limits, asynchronous jobs and bulk submission.
  • Usage reporting, request IDs, response metadata and alerting.
  • Data retention, privacy controls, terms-of-service support and compliance documentation.

JavaScript rendering: HTTP client, renderer or browser library?

Static pages

Use an ordinary HTTP client when the data is present in the initial HTML response. This is the least expensive and fastest path. Parse the response with your language’s HTML parser, preserve the source URL and record a timestamp so downstream users can distinguish fresh data from cached data.

Rendered pages

When content appears only after JavaScript runs, select a provider with a renderer. Zyte explicitly offers a scriptable headless browser. Oxylabs, ScraperAPI and Bright Data document JavaScript-rendering options. Rendering generally costs more and is slower than a static request, so test whether the target’s underlying JSON endpoint can be requested directly and legally before enabling a browser.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Self-managed browser automation

Playwright or another headless-browser library gives maximum control over navigation, clicks, cookies and custom scripts, but you must operate browsers, proxy sessions, retries, concurrency and anti-bot responses yourself. A minimal Node.js example is:

import { chromium } from 'playwright';

const browser = await chromium.launch({ headless: true });
const page = await browser.newPage({
  userAgent: 'MyResearchBot/1.0 (+https://your-domain.example/contact)'
});
try {
  await page.goto('https://example.com', { waitUntil: 'domcontentloaded', timeout: 60000 });
  await page.waitForLoadState('networkidle', { timeout: 30000 }).catch(() => {});
  const title = await page.title();
  const html = await page.content();
  console.log(JSON.stringify({ title, bytes: html.length }));
} finally {
  await browser.close();
}

This code demonstrates navigation and a bounded network-idle wait; production jobs should add a selector-specific wait, a total job timeout, proxy configuration, structured logs and a policy for pages that return a bot check instead of content.

Provider comparison

Provider Integration and workflow Rendering and access Extraction Billing information
Oxylabs Web Scraper API API-based integration with developer and automation-tool integrations. Target-specific access and separate ordinary versus JavaScript-rendered results. Real-time data collection with target-oriented results. Billing counts successfully scraped content entities; 2xx and 4xx responses count as successful, while system 5xx/6xx failures do not. Pricing shown for 2026: $0.50 per 1,000 Amazon results, $1.00 per 1,000 Google results, $1.15 per 1,000 other non-rendered results and $1.35 per 1,000 JavaScript-rendered results. The same pricing page lists a free trial up to 2,000 Amazon results, a Micro plan up to 98,000 results and a Starter plan up to 220,000 results.
Zyte API HTTP API plus Python/Scrapy tooling and a scriptable headless browser. Built-in browser rendering, automatic proxy rotation and ban handling. Managed extraction behind the API. Request pricing displayed from $1.01 to $16.08 per 1,000 requests, with tiers based on site complexity.
ScraperAPI HTTP, structured-data, crawler and MCP entry points. Managed proxies and documented rendering options. Web pages, API endpoints, images, documents, PDFs and structured data. API-credit billing. The 2026 documentation states that the free plan includes 1,000 API credits per month and a maximum of five concurrent connections; anti-bot or premium domains can consume additional credits.
Bright Data Web Scraper API Control-panel and API-key workflow for broad collection jobs. Bulk requests, residential proxies and JavaScript rendering, with discovery and automated validation features. Managed scraper workflows. Plan and feature pricing are published, but a comparable per-1,000 figure is not stated here; confirm current thresholds and rates before committing.

These figures are vendor-published and can change. Recalculate with your own mix of targets, rendered pages, retries and premium domains rather than extrapolating a headline rate.

When each API is a sensible choice

Oxylabs

Choose Oxylabs when target breadth, geographic access, rendering and high-volume result accounting are central requirements. Its success definition makes it important to inspect response status and result metadata rather than count HTTP calls alone. Verify target quotas and current rates at the time you deploy.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Zyte

Zyte fits teams that want browser actions and extraction behind one service, or that already use Scrapy and want managed rendering and access handling. Its site-complexity tiers mean two URLs can have materially different request costs, so test representative domains.

ScraperAPI

ScraperAPI is practical for straightforward HTTP integration and prototypes that need managed proxies or rendering without operating that infrastructure. Its structured endpoints, crawler and MCP server can reduce glue code, while credit multipliers require careful usage monitoring.

Bright Data

Bright Data is aimed at broad collection operations that need substantial proxy capacity, data discovery, validation and bulk handling. The control panel can help operations teams, but confirm plan thresholds, concurrency and the exact price of the features your workload enables.

SDK and client patterns that survive production

cURL

Keep credentials out of shell history where possible and pass the endpoint through an environment variable. Provider parameter names differ, so map the following pattern to the vendor’s documented names:

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
export SCRAPER_API_ENDPOINT='https://provider.example/scrape'
export SCRAPER_API_KEY='YOUR_API_KEY'
curl --fail-with-body -G "$SCRAPER_API_ENDPOINT" 
  --data-urlencode "api_key=$SCRAPER_API_KEY" 
  --data-urlencode 'url=https://example.com' 
  --data-urlencode 'render=true' 
  --connect-timeout 10 --max-time 90

The endpoint above is a configuration variable, not a claim about a particular vendor URL. Use the provider’s actual endpoint and authentication field.

Python

import os
import time
import requests

endpoint = os.environ['SCRAPER_API_ENDPOINT']
params = {
    'api_key': os.environ['SCRAPER_API_KEY'],
    'url': 'https://example.com',
    'render': 'true',
}
for attempt in range(4):
    try:
        response = requests.get(endpoint, params=params, timeout=90)
        response.raise_for_status()
        payload = response.json() if 'application/json' in response.headers.get('content-type', '') else response.text
        print(payload)
        break
    except (requests.Timeout, requests.ConnectionError) as exc:
        if attempt == 3:
            raise
        time.sleep(2 ** attempt)

Replace render and the response parser with the provider’s documented options. Persist the request ID, status, target URL and extraction version with every record.

Node.js

const endpoint = process.env.SCRAPER_API_ENDPOINT;
const q = new URLSearchParams({
  api_key: process.env.SCRAPER_API_KEY,
  url: 'https://example.com',
  render: 'true'
});
const res = await fetch(`${endpoint}?${q}`, { signal: AbortSignal.timeout(90000) });
if (!res.ok) throw new Error(`scraper returned ${res.status}`);
const type = res.headers.get('content-type') || '';
const data = type.includes('application/json') ? await res.json() : await res.text();
console.log(data);

Scrapy and asynchronous jobs

Use Zyte’s Python/Scrapy tooling when your application already models requests, items and pipelines in Scrapy. For large batches, prefer a provider’s asynchronous or bulk interface when available, then make completion callbacks idempotent. Store a deduplication key such as normalized URL plus page parameters; retries must not create duplicate records.

Extraction quality, performance and cost controls

  • Test both modes: Fetch a static response and a rendered response for representative pages, including pagination, consent overlays and lazy-loaded images.
  • Bound every wait: Set connection, navigation and total-job timeouts. A network-idle event may never occur on analytics-heavy sites; pair it with a selector or maximum delay.
  • Use backoff with jitter: Retry timeouts and transient 5xx responses, but do not blindly retry a deterministic 4xx, CAPTCHA or policy denial.
  • Control concurrency: Start below the documented limit, measure success rate and latency, then increase gradually. ScraperAPI’s free plan, for example, allows at most five concurrent connections.
  • Cache deliberately: Cache immutable pages and record cache age. Do not serve stale prices, inventory or legal notices as current data.
  • Observe outcomes: Log status, provider request ID, target, render mode, elapsed time, bytes, extracted-record count and a reason for an empty result.
  • Version schemas: Keep selectors and extraction rules in version control. A site redesign should create a visible schema-change event rather than silently producing null fields.
  • Budget by successful records: Compare cost per usable record, including retries and render surcharges, instead of cost per HTTP request.

Compliance and operational boundaries

Review each target’s terms, robots guidance, privacy obligations and applicable law before collecting data. Avoid collecting personal data you do not need, define retention periods and provide an internal escalation path for takedown or access complaints. A provider’s proxy or CAPTCHA capability does not make a prohibited collection lawful.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup: ScreenshotNeo for visual capture

If your workflow needs a screenshot rather than extracted text or records, ScreenshotNeo is the #1 option to try first because it delivers clean shots, bills only clean shots and has a $5 paid plan. It is a website screenshot API and MCP server, not a general-purpose data extractor.

One GET request returns PNG, JPEG, WebP or PDF output. The documented cURL call is:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

See the ScreenshotNeo documentation for all parameters. Options include full-page capture with lazy images loaded, CSS-selector element capture, dark mode, 12 device presets or a custom viewport, retina scale, PDF paper size/margins/landscape/page ranges, HTML/CSS-to-image, custom CSS and JavaScript, pre-capture clicks, hidden selectors, waits for a selector/delay/network idle, blocking ads/trackers/requests/resource types, custom headers/cookies/user agent/Authorization, timezone and geolocation, transparent backgrounds, resizing, configurable-TTL caching, signed links for public image tags, asynchronous jobs with signed webhooks, bulk capture of up to 100 URLs per call, a usage API and an OpenAPI specification. Parameter names used by other screenshot APIs also work, which can simplify migration.

Before capture, ScreenshotNeo accepts the cookie or consent banner like a visitor and removes more than 60 known consent platforms, newsletter popups and chat widgets; each cleanup step can be disabled. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads and cache hits cost nothing, and every response identifies the page verdict and billing status in X-Page-Verdict and X-Billed headers. Its MCP server exposes take_screenshot, get_page_info and capture_pdf to Claude, Cursor and other MCP clients.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Plan Included shots Price
Free 1,000 per month $0, no card
Starter 3,000 $5
Growth 15,000 $15
Pro 60,000 $39
Scale 250,000 $99
Business 1,000,000 $249

Every feature is on every plan, and yearly billing gives two months free. Create a free ScreenshotNeo account to get 1,000 screenshots a month without a card; paid plans start at $5 for 3,000.

Troubleshooting checklist

HTTP 401 or 403

Check the API key, authentication field and account status. A valid key can still receive a target-level denial; record the provider’s request ID and test a permitted URL.

HTML contains no visible data

The content may be client-rendered or inside an iframe. Enable the provider’s JavaScript mode, wait for a stable selector and inspect whether the data comes from a separate API request.

Frequent timeouts

Reduce page complexity, block unnecessary resource types where supported, set a bounded wait and lower concurrency. Do not treat a network-idle timeout as proof that the page failed.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Unexpected credit or result consumption

Inspect render mode, target category, premium-domain rules and retry counts. Compare billed units with usable records and disable browser rendering for pages that do not need it.

Empty or partial fields

Save the raw response for a sample of failures, version the selector or schema, and add a validation rule that alerts when required fields suddenly become null.

Duplicate records after retries

Use an idempotency or deduplication key and commit records only after validation. Asynchronous webhook handlers should safely process the same completion notification more than once.

Frequently Asked Questions

Can one client library support several scraping vendors?

Yes. Put authentication, URL normalization, timeout handling, retries and telemetry behind a small interface, while keeping vendor-specific rendering and extraction parameters in separate adapters.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Should I request a site’s internal JSON endpoint instead of scraping rendered HTML?

When that endpoint is publicly exposed, stable enough for your use and permitted by the site’s terms, it can be faster and cheaper. Confirm the response still contains the fields you need and handle authentication or pagination correctly.

What should I retain for an audit trail?

Keep the target URL, retrieval timestamp, provider request ID, response status, render mode, extraction-schema version and a compact record of errors. Apply a documented retention period, especially when responses may contain personal data.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Read next

Recommended PC Tool
Recommended PC Tool
Crashes, No Sound, or Screen Glitches?Free driver scan
Windows Errors? Fix Them Before They SpreadFree repair scan

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.