Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

A universal web scraper API is not a magic endpoint that extracts every website perfectly. It is a configurable service that accepts a URL and extraction contract, chooses a direct HTTP crawler or an isolated browser worker, applies per-domain policy and rate limits, validates the returned records, and exposes stable job and result APIs. Build the HTTP path first, add browser rendering only for pages that demonstrably need it, and treat extraction rules, access policy, and monitoring as part of every request.

Define “universal” correctly

“Universal” should describe the execution system, not promise that every target is scrapeable. Sites differ in markup, JavaScript behavior, authentication, robots.txt instructions, bot controls, pagination, and data quality. Your API can provide one contract while selecting different workers and extraction rules underneath.

A useful request names the target, the fields to return, and bounded crawl options. A useful response makes success, empty extraction, policy refusal, timeout, and upstream failure distinct outcomes. Never hide a partial or empty result behind a successful HTTP status.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use a layered architecture

1. API boundary

Expose a small public surface such as POST /v1/scrapes for asynchronous work and GET /v1/scrapes/{job_id} for status and results. For very small, predictable pages, you can offer a synchronous endpoint, but keep the same internal job model.

#1 Best Overall
CanaKit Raspberry Pi 5 Starter Kit PRO - Turbine Black (128GB Edition) (8GB RAM)
  • Includes Raspberry Pi 5 with 2.4Ghz 64-bit quad-core CPU (8GB RAM)
  • Includes 128GB Micro SD Card pre-loaded with 64-bit Raspberry Pi OS, USB MicroSD Card Reader
  • CanaKit Turbine Black Case for the Raspberry Pi 5
  • CanaKit Low Noise Bearing System Fan
  • Mega Heat Sink - Black Anodized
  • Request: URL, field schema, optional selector map, render mode, maximum pages, timeout, and an idempotency key.
  • Response: job identifier, status, records, source metadata, and structured errors. Do not return internal credentials, queue names, or worker addresses.
  • Limits: cap URL length, page count, response bytes, redirects, browser time, and total job duration.

2. Policy and validation

Validate schemes before scheduling. Accept only http and https; reject embedded credentials and malformed hostnames. Apply your own destination policy for private networks, loopback addresses, cloud metadata endpoints, and disallowed ports, and re-check the resolved destination after redirects. The exact SSRF defense depends on your deployment, so make the checks explicit and test them.

Check the target’s published access policy before fetching. Keep an audit record of the decision, the URL, and the policy version. An API customer should be able to tell whether a job was refused by policy or failed during retrieval.

3. Scheduler and queue

Partition queued work by registrable domain (or another target key you control). This lets you enforce concurrency and delay for the site being fetched instead of accidentally sending a burst from many workers. Use bounded retries with backoff, a maximum attempt count, and cancellation support. Retain status transitions such as queued, running, succeeded, empty, blocked, timed_out, and failed.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

4. Fetch tiers

Target behavior Preferred worker Reason Controls to expose
Static HTML, JSON, XML, or a documented endpoint Direct HTTP Lower latency and resource use; easy to scale Headers, cookies, timeout, redirects, byte limit
Content appears after JavaScript execution Playwright browser Runs page scripts and waits for rendered content Viewport, wait condition, browser timeout, proxy
Clicks, form submission, infinite scroll, or interaction required Playwright browser with a declared action plan Models the required interaction instead of guessing at HTML Allowed actions, step timeout, page and resource limits
Official API, search endpoint, or bulk export exists That endpoint Usually faster for the caller and cheaper for the target than crawling pages Provider quotas, pagination, authentication, response validation

Scrapy supplies a conventional crawl lifecycle—spiders, requests and responses, selectors, items, pipelines, middleware, scheduling, statistics, and exporters. Use it for the direct HTTP tier when you need mature crawl orchestration. Playwright is an optional execution path; its Browser API supports HTTP and SOCKS proxies. Keeping these tiers separate prevents a browser-heavy workload from consuming the capacity intended for ordinary HTTP jobs.

Design a stable extraction contract

Selectors are configuration, not universal intelligence

Accept a declared field map such as {"title":"h1", "price":".price"}, using CSS or XPath selectors. Store extraction rules by site and version them. When a selector matches zero elements, return an explicit extraction error or an empty outcome according to your contract; do not silently emit a record with missing required fields.

Normalize and validate records

Convert text to a predictable representation, trim whitespace, normalize repeated spaces, and parse dates or numbers only when the schema says how. Validate required fields and types before publishing a result. Keep the source URL, fetch timestamp, HTTP status, and worker type as metadata so downstream users can diagnose changes.

Rank #2
CanaKit Raspberry Pi 4 4GB Starter PRO Kit - 4GB RAM
  • Includes Raspberry Pi 4 4GB Model B with 1.5GHz 64-bit quad-core CPU (4GB RAM)
  • Includes Pre-Loaded 32GB EVO+ Micro SD Card (Class 10), USB MicroSD Card Reader
  • CanaKit Premium High-Gloss Raspberry Pi 4 Case with Integrated Fan Mount, CanaKit Low Noise Bearing System Fan
  • CanaKit 3.5A USB-C Raspberry Pi 4 Power Supply (US Plug) with Noise Filter, Set of Heat Sinks, Display Cable - 6 foot (Supports up to 4K60p)
  • CanaKit USB-C PiSwitch (On/Off Power Switch for Raspberry Pi 4)

Choose an output format

Your service can return JSON while allowing internal exporters for JSON Lines, XML, CSV, or object storage. A stable envelope is easier to evolve:

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
{
  "job_id": "job_01J...",
  "status": "succeeded",
  "records": [{"title": "Example", "price": 19.99}],
  "meta": {"url": "https://example.com/item", "worker": "http", "http_status": 200},
  "errors": []
}

Build a small HTTP-first reference service

The following Python example is intentionally narrow: it demonstrates a synchronous endpoint, selector-based extraction, and schema validation. Install the dependencies with pip install fastapi uvicorn httpx parsel, save it as main.py, and run uvicorn main:app --reload. Add authentication, a queue, destination controls, and persistent storage before exposing it to untrusted callers.

from typing import Dict, List
from urllib.parse import urlparse

import httpx
from fastapi import FastAPI, HTTPException
from pydantic import BaseModel, Field
from parsel import Selector

app = FastAPI(title="Universal Scraper API")

class ScrapeRequest(BaseModel):
    url: str
    fields: Dict[str, str] = Field(min_length=1)
    timeout_seconds: float = Field(default=20, gt=0, le=90)

class ScrapeResponse(BaseModel):
    status: str
    records: List[dict]
    meta: dict
    errors: List[dict]

def validate_url(raw: str) -> str:
    parsed = urlparse(raw)
    if parsed.scheme not in {"http", "https"} or not parsed.hostname:
        raise HTTPException(status_code=400, detail="Only http and https URLs are accepted")
    if parsed.username or parsed.password:
        raise HTTPException(status_code=400, detail="Embedded credentials are not accepted")
    return raw

@app.post("/v1/scrape", response_model=ScrapeResponse)
async def scrape(req: ScrapeRequest):
    url = validate_url(req.url)
    try:
        async with httpx.AsyncClient(
            follow_redirects=True,
            headers={"User-Agent": "UniversalScraper/1.0"},
            timeout=req.timeout_seconds,
        ) as client:
            response = await client.get(url)
    except httpx.TimeoutException:
        return ScrapeResponse(status="timed_out", records=[], meta={"url": url},
                              errors=[{"code": "timeout"}])
    except httpx.HTTPError as exc:
        return ScrapeResponse(status="failed", records=[], meta={"url": url},
                              errors=[{"code": "fetch_error", "message": str(exc)}])

    if response.status_code >= 400:
        return ScrapeResponse(status="failed", records=[],
                              meta={"url": str(response.url), "http_status": response.status_code},
                              errors=[{"code": "upstream_http", "status": response.status_code}])

    selector = Selector(text=response.text)
    record = {}
    missing = []
    for name, css in req.fields.items():
        value = selector.css(css).xpath("string(.)").get()
        if value is None or not value.strip():
            missing.append(name)
        else:
            record[name] = " ".join(value.split())

    meta = {"url": str(response.url), "http_status": response.status_code, "worker": "http"}
    if missing:
        return ScrapeResponse(status="empty", records=[], meta=meta,
                              errors=[{"code": "required_fields_missing", "fields": missing}])
    return ScrapeResponse(status="succeeded", records=[record], meta=meta, errors=[])

This example fetches one document and one record. A production crawler should move the work to a queue, add pagination as an explicit option, enforce byte and redirect limits, and use a domain scheduler. Scrapy can provide the request scheduling, middleware, selectors, pipelines, and statistics once the HTTP path needs crawl breadth.

Add browser workers only when needed

Detect browser dependence with a documented rule: the required selector is absent from the HTTP response, or the request explicitly asks for rendering or interaction. Send only that job to a Playwright worker. Reuse browser processes carefully, isolate contexts per job, and close pages even on errors.

from playwright.async_api import async_playwright

async def rendered_html(url: str, wait_for: str | None = None) -> str:
    async with async_playwright() as pw:
        browser = await pw.chromium.launch(headless=True)
        context = await browser.new_context()
        page = await context.new_page()
        try:
            await page.goto(url, wait_until="domcontentloaded", timeout=60000)
            if wait_for:
                await page.wait_for_selector(wait_for, timeout=30000)
            return await page.content()
        finally:
            await context.close()
            await browser.close()

Make interaction steps declarative and bounded—for example, a permitted click selector followed by a wait selector. Do not allow arbitrary customer JavaScript in the same security context as your service. Browser jobs consume more CPU and memory than direct HTTP jobs; measure your own workload rather than assuming a universal cost or speed ratio.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Respect robots.txt and target pacing

Robots.txt handling must be an explicit setting. Scrapy’s robots middleware does not automatically apply Crawl-delay or Request-rate directives. Parse those values (where applicable) and translate them into your scheduler’s delay and concurrency settings. Record the effective policy with each job so operators can explain why a request was delayed or refused.

Rank #3
ELECROW CrowPi Case Kit for Raspberry Pi 5, 9-Inch Display
  • Not including the Raspberry Pi 5 (8GB), the Crowpi advanced version comes with the Raspberry Pi 5
  • ELECROW Black Case for the Raspberry Pi 5, CrowPi is equipped with a 9-inch HD touchscreen along with a camera; All the regular components used in DIY electronics are packed into the CrowPi development board, such as LCD, LED matrix, buzzer, light sensor, PIR sensor, ultrasonic sensor, IR sensor, etc
  • Raspberry Pi Sensors: The Crowpi raspberry pi 5 programming kit is jam-packed with lots of buttons such as 19 different sensors in a tidy easy to use package; You don't have to wait and wire things
  • Build Quality: Solid ABS shell and well made components in one place make it strong and convenient to travel
  • Programming Lessons: This raspberry pi 5 learning kit ships with step by step instructions and provides 21 lessons to take you through identifying components reading code and running it in the terminal

Set separate limits for each domain: maximum in-flight requests, minimum delay, maximum retry count, and a circuit breaker for repeated failures. If a target starts returning throttling responses, reduce concurrency and back off instead of retrying immediately. Prefer a documented API, search endpoint, or bulk export whenever one exists.

Make jobs observable and recoverable

  • Metrics: queue wait, fetch latency, browser startup time, status-code distribution, bytes, retries, empty results, and extraction failures, all grouped by target.
  • Logs: job ID, target host, worker type, attempt number, policy decision, and a redacted error category. Never log authorization headers or cookies.
  • Cancellation: mark a job cancelled in durable state and have workers check that state between pages and interaction steps.
  • Retention: set limits for raw HTML, screenshots, cookies, and extracted records. Keep only what your product contract requires.
  • Capacity: autoscale HTTP workers separately from browser workers, and cap per-tenant concurrency so one customer cannot monopolize a domain or your fleet.

Use bounded retries. Retry connection resets and selected transient gateway responses; do not retry malformed requests, policy refusals, authentication failures, or deterministic selector errors. Preserve the final error and the last successful response metadata.

Performance, reliability, and cost decisions

Keep the fast path fast

Reuse HTTP connections, stream or cap large responses, avoid downloading unnecessary resource types, and cache only when the caller accepts potentially stale data. For browser jobs, block unneeded resources, reuse a controlled browser pool, and set navigation and overall job deadlines.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Separate freshness from duplication

Accept an idempotency key for submissions and define whether a repeated request returns an existing job, a cached result, or a new fetch. Expose cache age in metadata. Never claim a cache hit is a fresh crawl.

Price from your workload

Compare actual HTTP transfer, browser CPU and memory, queue infrastructure, storage, and third-party endpoint costs. The available framework guidance supports avoiding unnecessary page crawling when an API or export exists, but it does not establish a universal price or performance winner.

Troubleshooting common failures

HTTP returns a page shell with no fields

Cause: data is rendered by JavaScript. Fix: inspect the page’s documented endpoint first; otherwise route the job to Playwright and wait for a specific selector rather than an arbitrary long sleep.

Rank #4
CanaKit Raspberry Pi 5 Desktop PC with SSD (Fully Assembled) (256 GB SSD)
  • Fully assembled for plug-and-play operation
  • Includes Raspberry Pi 5 with 8GB RAM
  • 256 GB PCIe Pi NVMe SSD (Pre-loaded with Pi 64-Bit OS)
  • M.2 HAT+
  • CanaKit Turbine Black Case for the Pi 5

Results suddenly become empty

Cause: markup changed, a selector is too broad or narrow, or the response is a consent, login, or bot-check page. Fix: store a redacted response sample, classify the page, version the selector rule, and return an explicit extraction failure.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Many 429 or 403 responses

Cause: pacing, access policy, or credentials are wrong. Fix: verify authorization, honor robots and published limits, reduce per-domain concurrency, increase delay, and stop retrying aggressively.

Browser jobs time out

Cause: waiting for a selector that never appears, blocked resources, or an overloaded worker. Fix: capture timing telemetry, use a precise wait condition, set an overall deadline, and move heavy jobs to isolated capacity.

Private or internal addresses are reachable

Cause: destination checks were applied only to the original hostname, not redirects or resolved addresses. Fix: enforce an allow/deny policy before fetch and after every redirect and DNS resolution, with network-level egress controls as a second boundary.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup

When your task is taking a clean screenshot rather than extracting arbitrary records, ScreenshotNeo provides a single HTTP call. It accepts cookie or consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each cleanup step can be disabled. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and whether it was billed. It also offers an MCP server with take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

See the parameter reference in the ScreenshotNeo documentation. cURL:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Every plan includes the same feature set: full-page and element capture, device presets, custom viewport and retina scale, PDF controls, HTML/CSS rendering, custom CSS and JavaScript, clicks, waits, request blocking, headers, cookies, user agents, authorization, timezone and geolocation, transparent backgrounds, resizing, configurable-TTL caching, signed links, asynchronous webhooks, bulk capture of up to 100 URLs per call, a usage API, and an OpenAPI specification. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 screenshots. Create a free ScreenshotNeo account.

Best Value
RasTech Raspberry Pi 5 8GB Kit with Active Cooler and Pi5 Case
  • 【What you Get】You will get 1*Pi 5 8GB Single Board,1*RasTech Case,1*Active Cooler,1*Screwdriver,1*Installation instructions,12-month free warranty, lifetime service, 24-hour prompt and friendly response.
  • 【More Connectors】There are two USB 3.0 ports(5Gbps simultaneously) and two USB 2.0 ports, which triple total bandwidth ,support any combination of up to two cameras or displays. Peak SD card performance is doubled through support for the SDR104 high-speed mode. It provides a smooth desktop experience for you. Offer Gigabit Ethernet and a PCIe interface, along with dual-band Wi-Fi and Bluetooth 5.0/BLE wireless capability. The RasTech Pi 5 Kit use the new 27W 5.1V 5A USB-C power connector.
  • 【 Support Dual 4Kp60 Display 】Each of the two microHDMI sockets can control a 4K display at 60 Hertz, now support HDR, offering super HD video for media streaming projects. RPi 5 is the first RPi model that comes with a PCI Express port (PCIe 2.0 x1 with 500 MB/s) to attach SSDs (requires separate M.2 HAT).
  • 【 Excellent Chips And Applications】Pi 5 is a full-size Pi computer using silicon built in-house at Pi. The RP1 “southbridge” provides the bulk of the I/O capabilities for Pi 5. Pi 5 is more friendly and convenient in the development of Internet of Things, Web development, machine identification, automatic control and other electronic equipment applications and network.
  • 【 Faster CPU, Better GPU 】 Pi 5 features a Broadcom BCM2712 64-bit quad-core Arm Cortex-A76 processor running at 2.4GHz, it delivers a 2–3× increase in CPU performance relative to RaspberryPi 4. The 800MHz VideoCore VII GPU is compatible to OpenGL ES 3.1 and Vulkan 1.2, substantial uplift in graphics performance. Pi 5 Offers lightning-fast CPU speed, a PCI Express interface, a Real Time Clock (RTC) and a power button and runs significantly cooler than Pi 4.

FAQ

Should every request launch a browser?

No. Use direct HTTP for static documents and reserve browser workers for demonstrated rendering or interaction requirements.

Can one selector rule work forever?

No. Treat selectors as versioned configuration, validate required fields, and alert on empty or malformed results.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Is robots.txt the same as a legal permission?

No single crawler setting resolves every legal or contractual question. Implement an explicit policy process and obtain authorization for the targets you operate on.

When should a scrape be asynchronous?

Use a job queue when work can involve multiple pages, browser rendering, retries, or unpredictable latency. Return a job ID and let clients poll or receive a callback.

Frequently Asked Questions

Should every request launch a browser?

No. Use direct HTTP for static documents and reserve browser workers for demonstrated rendering or interaction requirements.

Can one selector rule work forever?

No. Treat selectors as versioned configuration, validate required fields, and alert on empty or malformed results.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Is robots.txt the same as a legal permission?

No single crawler setting resolves every legal or contractual question. Implement an explicit policy process and obtain authorization for the targets you operate on.

When should a scrape be asynchronous?

Use a job queue when work can involve multiple pages, browser rendering, retries, or unpredictable latency. Return a job ID and let clients poll or receive a callback.

Quick Recap

Bestseller No. 1
CanaKit Raspberry Pi 5 Starter Kit PRO - Turbine Black (128GB Edition) (8GB RAM)
CanaKit Raspberry Pi 5 Starter Kit PRO - Turbine Black (128GB Edition) (8GB RAM)
Includes Raspberry Pi 5 with 2.4Ghz 64-bit quad-core CPU (8GB RAM); CanaKit Turbine Black Case for the Raspberry Pi 5
$259.95
Bestseller No. 2
CanaKit Raspberry Pi 4 4GB Starter PRO Kit - 4GB RAM
CanaKit Raspberry Pi 4 4GB Starter PRO Kit - 4GB RAM
Includes Raspberry Pi 4 4GB Model B with 1.5GHz 64-bit quad-core CPU (4GB RAM); Includes Pre-Loaded 32GB EVO+ Micro SD Card (Class 10), USB MicroSD Card Reader
$159.99
Bestseller No. 4
CanaKit Raspberry Pi 5 Desktop PC with SSD (Fully Assembled) (256 GB SSD)
CanaKit Raspberry Pi 5 Desktop PC with SSD (Fully Assembled) (256 GB SSD)
Fully assembled for plug-and-play operation; Includes Raspberry Pi 5 with 8GB RAM; 256 GB PCIe Pi NVMe SSD (Pre-loaded with Pi 64-Bit OS)
$339.97

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.