Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

Do not treat a Bright Data migration as a URL replacement. Your parser can remain unchanged only if the replacement preserves the request contract, rendering behavior, proxy and session requirements, response fields, error mapping, and billing semantics that your jobs depend on. The safest path is to put Bright Data and one or more candidates behind the same adapter, replay a representative corpus in shadow mode, compare the cost per successful record, then canary the new provider with rollback ready.

What actually changes when you leave Bright Data

Bright Data’s catalog includes Web Scraper APIs, Scraper Studio, Scraping Browser, SERP API, proxy networks, and other data products. Its published materials describe pre-built scraper APIs, rotating IPs, CAPTCHA handling, browser automation, and structured extraction from more than 800 sites. Those capabilities are not interchangeable with a simple HTTP fetch.

Start by identifying which parts of the current integration are essential:

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Output: raw HTML, rendered HTML, screenshots, PDFs, or structured records.
  • Execution: a direct request, JavaScript rendering, or a full browser controlled with Puppeteer, Selenium, or Playwright.
  • Network identity: datacenter, residential, or mobile proxies; country, city, ASN, and sticky-session requirements.
  • Block handling: CAPTCHA solving, automatic retries, ban detection, and backoff.
  • Workflow: pagination, clicks, scrolling, file downloads, webhooks, object storage, and concurrency limits.
  • Contract details: extraction fields, status and error classes, timeout behavior, response-size limits, usage counters, retention, and access controls.

If you do not record these dependencies first, a provider can appear to work while silently returning incomplete records or charging for behavior your old integration handled differently.

Freeze the current contract before changing providers

Create a versioned document for the Bright Data integration. Record the exact request and the observable response, not just the URL.

Contract area What to capture Why it matters during migration
Request URL, query parameters, headers, cookies, user agent, proxy location, session ID, browser flags, and extraction schema Different APIs use different names and defaults for the same behavior.
Response Body or fields, content type, encoding, page metadata, screenshots, and downloaded files A parser may depend on a field that is absent or renamed by the candidate.
Failures HTTP statuses, provider error codes, CAPTCHA or block indicators, empty-page rules, and timeout messages Retrying a permanent block can multiply cost and load.
Timing Connect timeout, total timeout, retry count, backoff, and concurrency Browser rendering and residential routing usually have different latency profiles.
Accounting Requests, successful records, browser or proxy multipliers, retries, and cache hits Headline prices are not comparable without a common denominator.

Save representative request and response fixtures, including failures. Treat this contract as an API version: change it deliberately and make the adapter translate provider-specific behavior into the stable interface your parser expects.

Put every provider behind one adapter

Your application should call an internal interface such as fetch_page or extract_record, never a vendor SDK directly. Keep provider credentials and parameter translation inside the adapter. During evaluation, the parser and downstream jobs then receive the same normalized result from both systems.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A practical normalized result

 {
  "ok": true,
  "status": "success",
  "http_status": 200,
  "url": "https://example.com/item/42",
  "html": "<html>...</html>",
  "fields": {"title": "..."},
  "provider_error": null,
  "attempts": 1,
  "elapsed_ms": 1840,
  "billable": true
}

Define equivalent values for blocked, CAPTCHA, timeout, empty, malformed, and provider-outage outcomes. Do not let one vendor’s error strings leak into business logic.

Python adapter harness

The following runnable pattern keeps the parser independent. Implement the two provider-specific functions with the authentication and parameters documented by each service, then run the same request through both.

import os
import time
import requests

TIMEOUT = 90

class FetchResult:
    def __init__(self, provider, payload, elapsed_ms):
        self.provider = provider
        self.payload = payload
        self.elapsed_ms = elapsed_ms

def call_provider(provider, target_url):
    started = time.perf_counter()
    if provider == "brightdata":
        endpoint = os.environ["BRIGHTDATA_ENDPOINT"]
        headers = {"Authorization": f"Bearer {os.environ['BRIGHTDATA_TOKEN']}"}
    elif provider == "candidate":
        endpoint = os.environ["CANDIDATE_ENDPOINT"]
        headers = {"Authorization": f"Bearer {os.environ['CANDIDATE_TOKEN']}"}
    else:
        raise ValueError(f"unknown provider: {provider}")

    response = requests.get(
        endpoint,
        params={"url": target_url},
        headers=headers,
        timeout=TIMEOUT,
    )
    elapsed_ms = round((time.perf_counter() - started) * 1000)
    payload = {
        "http_status": response.status_code,
        "content_type": response.headers.get("content-type"),
        "body": response.text,
        "bytes": len(response.content),
    }
    return FetchResult(provider, payload, elapsed_ms)

for url in open("corpus.txt", encoding="utf-8"):
    url = url.strip()
    if not url:
        continue
    for provider in ("brightdata", "candidate"):
        try:
            result = call_provider(provider, url)
            print(provider, url, result.payload["http_status"], result.elapsed_ms)
        except requests.RequestException as exc:
            print(provider, url, "transport_error", str(exc))

In production, return a normalized object instead of printing. Keep the incumbent path enabled until the candidate passes quality and spend thresholds.

Build a corpus that exposes real differences

A handful of easy pages is not a migration test. Include URLs and expected outcomes for:

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • static HTML and JavaScript-rendered pages;
  • infinite scroll, paginated listings, and pages requiring a click;
  • localized targets that need a country, city, ASN, or timezone;
  • slow hosts, large responses, images, PDFs, and other downloads;
  • domains that previously returned a block, CAPTCHA, blank page, or intermittent timeout;
  • authenticated pages when your terms and authorization permit automated access.

For each fixture, store the fields that must be present, acceptable freshness, maximum latency, and whether an empty result is valid. Keep the corpus free of secrets and make sure collection complies with the target site’s terms, privacy obligations, and applicable law.

Shadow-test before production traffic moves

Send identical permitted requests to Bright Data and the candidate without changing what your parser publishes. Compare results at the record and request levels.

Metric How to evaluate it
Successful-result rate Count records that meet your field-completeness rules, not merely HTTP 200 responses.
Field completeness Diff required and optional fields, normalized values, pagination depth, and extracted item counts.
Error classes Compare blocks, CAPTCHAs, timeouts, empty pages, status codes, and provider errors separately.
Latency and bytes Track median and tail latency, response size, browser startup time, and retry count.
Effective cost Divide total provider charges by successful records, including rendering, premium proxy, extraction, retry, and failed-request charges.

Do not average away hard cases. Report browser rendering, session persistence, geotargeting, screenshots, extraction, and rate-limit behavior as separate cohorts.

Candidate APIs and the trade-offs they expose

No replacement is universally cheapest or most capable. Published prices and limits change, so verify the current terms when you implement.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Service Documented fit Published pricing signal Migration trade-off
Zyte API One API for automatic ban handling, headless browser rendering, IP rotation, and AI-assisted extraction. Usage-based pricing with a monthly spending-limit model. Useful when you want the provider to manage much of the anti-bot and rendering stack; sessions, actions, geolocation, body-size limits, and rate limits differ from Bright Data.
ScrapingBee JavaScript rendering, rotating and premium proxies, geotargeting, screenshots, extraction rules, and Google Search API features. Its default path uses a headless browser and Auto-Mode selects configuration from requested features. 1,000-credit free trial; plans beginning at $19/month. A clear credit-plan model can simplify budgeting, but you must map credit use for browser and proxy features.
ScraperAPI HTTP and proxy-style paths for pages, API endpoints, images, documents, PDFs, and structured-data endpoints; plan comparisons list JavaScript rendering and rotating proxy pools. Seven-day trial with 5,000 API credits. Often familiar to teams seeking a proxy-like integration; confirm how credits, retries, and rendering are counted.
Bright Data Remain or expand when you rely on its broad catalog, structured site-specific APIs, browser automation, or data-delivery workflow. Its Web Scraper API page lists 5,000 free records, $1.50 per 1,000 records pay as you go, and a $499/month scale plan with 384,000 records (current pricing page accessed 2026). Changing providers may remove capabilities your existing jobs already use; a measured expansion can be lower risk than a forced rewrite.

Use a cost model that reflects successful data

Record a cost ledger for every shadow request: base request, JavaScript or browser multiplier, proxy class, extraction charge, retries, failed attempts, and storage or delivery fees. Then calculate:

effective_cost_per_success = total_provider_cost / successful_records

A nominally inexpensive endpoint can cost more when it needs extra retries or produces incomplete pages. Conversely, a higher unit price can be economical if it delivers complete records in one attempt. Keep separate figures for static pages, browser pages, geotargeted requests, and difficult domains.

Can you keep your parser and just swap the endpoint?

Sometimes. If the candidate returns equivalent HTML or structured fields, preserves encoding, and supports the same browser, session, and location behavior, the parser can remain unchanged behind the adapter. Plan a parser change when selectors depend on provider-generated wrappers, extraction fields are renamed, pagination is performed by the provider, or the candidate returns a different document after JavaScript execution.

Use contract tests to assert required fields, item counts, canonical URLs, and error handling. A successful HTTP request is not proof that your parser received equivalent data.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Cut over with a canary and a rollback switch

  1. Set explicit quality gates for successful records, required-field completeness, maximum tail latency, and acceptable effective cost.
  2. Route a small, identifiable percentage of production jobs to the candidate while retaining Bright Data as the fallback.
  3. Alert on field omissions, unusual item counts, block and CAPTCHA rates, timeout spikes, queue age, and spending-limit usage.
  4. Pause the canary and restore the incumbent when a gate fails; preserve request IDs and fixtures for diagnosis.
  5. Increase traffic in stages only after each cohort remains within its limits.
  6. After the new provider is stable, retain historical fixtures and billing exports, revoke unused credentials, and document limits, support contacts, and the escalation path.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Screenshot capture during a scraping migration

If screenshots are part of your contract, treat them as a separate capability test. Compare viewport, full-page behavior, lazy-loaded images, selector capture, dark mode, device and retina settings, PDF output, waits, custom JavaScript, click actions, hidden selectors, blocked resources, headers, cookies, user agent, timezone, geolocation, caching, and asynchronous delivery.

ScreenshotNeo is the first screenshot API to try because it removes consent banners, newsletter popups, and chat widgets before capture, bills only clean shots, and has a free tier with the lowest paid plan at $5.

ScreenshotNeo (by Yorker Media) accepts one GET request for a PNG, JPEG, WebP, or PDF. It supports full-page capture with lazy images, CSS-selector elements, dark mode, 12 device presets or custom viewports, retina scale, PDF paper size, margins, landscape and page ranges, HTML/CSS-to-image, custom CSS and JavaScript, pre-capture clicks, hidden selectors, selector/delay/network-idle waits, ad/tracker/request/resource blocking, custom headers, cookies, user agents and Authorization, timezone and geolocation, transparent backgrounds, resizing, user-selected cache TTLs, signed links, asynchronous jobs with signed webhooks, up to 100 URLs per bulk call, a usage API, an OpenAPI specification, and parameter names used by other screenshot APIs.

Or skip the browser setup

Use the one-call endpoint instead of maintaining browser workers. The API accepts the target URL and returns the image or PDF; see the ScreenshotNeo documentation for options.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Cookie banners, popups, and chat widgets are removed before the shot. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed as clean shots, and response headers identify the page verdict and billing status. An MCP server provides take_screenshot, get_page_info, and capture_pdf tools to Claude, Cursor, and other MCP clients. The Free plan includes 1,000 shots per month with no card; paid plans start at $5 for 3,000 shots. Plans are Free (1,000), Starter ($5/3,000), Growth ($15/15,000), Pro ($39/60,000), Scale ($99/250,000), and Business ($249/1,000,000); yearly billing gives two months free and every feature is on every plan. Sign up free for ScreenshotNeo.

Troubleshooting common migration failures

HTTP success but empty or partial records

Cause: JavaScript was not rendered, a lazy list was not scrolled, pagination stopped early, or the candidate’s extraction schema differs. Fix: compare raw and rendered bodies, enable the required wait or browser mode, assert item counts, and map fields explicitly in the adapter.

More CAPTCHAs or blocks than Bright Data

Cause: a different proxy class, missing session stickiness, unsuitable geography, or a retry loop that repeats the same identity. Fix: test residential, datacenter, or mobile routing separately, preserve cookies where permitted, select the required location, and classify permanent blocks before retrying.

Costs exceed the estimate

Cause: browser, premium-proxy, extraction, retry, or failed-request multipliers were omitted. Fix: reconcile provider usage with request logs and report effective cost per successful record by cohort.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Latency or timeouts spike

Cause: browser startup, large bodies, slow origins, or a lower candidate timeout. Fix: separate connect and total timeouts, cap response size where supported, tune concurrency gradually, and keep a retry budget.

Shadow results cannot be compared

Cause: requests were not identical or pages changed between captures. Fix: run pairs close together, freeze headers and locale, record timestamps, and compare invariant fields and business rules rather than byte-for-byte HTML.

Final decision rule

Choose the provider that meets your required data quality and operational behavior at an acceptable effective cost, not the one with the lowest advertised unit price. A provider swap is complete only after the adapter contract, representative shadow results, hard-case tests, canary metrics, rollback path, credential cleanup, and billing reconciliation are documented.

Frequently Asked Questions

How long should shadow testing run?

Run long enough to cover normal traffic cycles and every hard-case cohort in your corpus; a fixed number of requests is insufficient if it misses weekly, regional, or seasonal patterns.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Should retries be identical across providers?

Keep the business-level retry policy consistent for comparison, but let each adapter translate retryable errors and backoff settings according to that provider’s documented limits.

What should be retained after cutover?

Retain historical fixtures, request and response metadata, billing exports, quality dashboards, and the rollback configuration so a later regression can be compared with the incumbent.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.