Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
A universal web scraper API is not a magic endpoint that extracts every website perfectly. It is a configurable service that accepts a URL and extraction contract, chooses a direct HTTP crawler or an isolated browser worker, applies per-domain policy and rate limits, validates the returned records, and exposes stable job and result APIs. Build the HTTP path first, add browser rendering only for pages that demonstrably need it, and treat extraction rules, access policy, and monitoring as part of every request.
Define “universal” correctly
“Universal” should describe the execution system, not promise that every target is scrapeable. Sites differ in markup, JavaScript behavior, authentication, robots.txt instructions, bot controls, pagination, and data quality. Your API can provide one contract while selecting different workers and extraction rules underneath.
A useful request names the target, the fields to return, and bounded crawl options. A useful response makes success, empty extraction, policy refusal, timeout, and upstream failure distinct outcomes. Never hide a partial or empty result behind a successful HTTP status.
Free tools Windows power users keep installed
One-click scans. No signup required.
Use a layered architecture
1. API boundary
Expose a small public surface such as POST /v1/scrapes for asynchronous work and GET /v1/scrapes/{job_id} for status and results. For very small, predictable pages, you can offer a synchronous endpoint, but keep the same internal job model.
#1 Best Overall
- Includes Raspberry Pi 5 with 2.4Ghz 64-bit quad-core CPU (8GB RAM)
- Includes 128GB Micro SD Card pre-loaded with 64-bit Raspberry Pi OS, USB MicroSD Card Reader
- CanaKit Turbine Black Case for the Raspberry Pi 5
- CanaKit Low Noise Bearing System Fan
- Mega Heat Sink - Black Anodized
- Request: URL, field schema, optional selector map, render mode, maximum pages, timeout, and an idempotency key.
- Response: job identifier, status, records, source metadata, and structured errors. Do not return internal credentials, queue names, or worker addresses.
- Limits: cap URL length, page count, response bytes, redirects, browser time, and total job duration.
2. Policy and validation
Validate schemes before scheduling. Accept only http and https; reject embedded credentials and malformed hostnames. Apply your own destination policy for private networks, loopback addresses, cloud metadata endpoints, and disallowed ports, and re-check the resolved destination after redirects. The exact SSRF defense depends on your deployment, so make the checks explicit and test them.
Check the target’s published access policy before fetching. Keep an audit record of the decision, the URL, and the policy version. An API customer should be able to tell whether a job was refused by policy or failed during retrieval.
3. Scheduler and queue
Partition queued work by registrable domain (or another target key you control). This lets you enforce concurrency and delay for the site being fetched instead of accidentally sending a burst from many workers. Use bounded retries with backoff, a maximum attempt count, and cancellation support. Retain status transitions such as queued, running, succeeded, empty, blocked, timed_out, and failed.
4. Fetch tiers
| Target behavior | Preferred worker | Reason | Controls to expose |
|---|---|---|---|
| Static HTML, JSON, XML, or a documented endpoint | Direct HTTP | Lower latency and resource use; easy to scale | Headers, cookies, timeout, redirects, byte limit |
| Content appears after JavaScript execution | Playwright browser | Runs page scripts and waits for rendered content | Viewport, wait condition, browser timeout, proxy |
| Clicks, form submission, infinite scroll, or interaction required | Playwright browser with a declared action plan | Models the required interaction instead of guessing at HTML | Allowed actions, step timeout, page and resource limits |
| Official API, search endpoint, or bulk export exists | That endpoint | Usually faster for the caller and cheaper for the target than crawling pages | Provider quotas, pagination, authentication, response validation |
Scrapy supplies a conventional crawl lifecycle—spiders, requests and responses, selectors, items, pipelines, middleware, scheduling, statistics, and exporters. Use it for the direct HTTP tier when you need mature crawl orchestration. Playwright is an optional execution path; its Browser API supports HTTP and SOCKS proxies. Keeping these tiers separate prevents a browser-heavy workload from consuming the capacity intended for ordinary HTTP jobs.
Design a stable extraction contract
Selectors are configuration, not universal intelligence
Accept a declared field map such as {"title":"h1", "price":".price"}, using CSS or XPath selectors. Store extraction rules by site and version them. When a selector matches zero elements, return an explicit extraction error or an empty outcome according to your contract; do not silently emit a record with missing required fields.
Normalize and validate records
Convert text to a predictable representation, trim whitespace, normalize repeated spaces, and parse dates or numbers only when the schema says how. Validate required fields and types before publishing a result. Keep the source URL, fetch timestamp, HTTP status, and worker type as metadata so downstream users can diagnose changes.
Rank #2
- Includes Raspberry Pi 4 4GB Model B with 1.5GHz 64-bit quad-core CPU (4GB RAM)
- Includes Pre-Loaded 32GB EVO+ Micro SD Card (Class 10), USB MicroSD Card Reader
- CanaKit Premium High-Gloss Raspberry Pi 4 Case with Integrated Fan Mount, CanaKit Low Noise Bearing System Fan
- CanaKit 3.5A USB-C Raspberry Pi 4 Power Supply (US Plug) with Noise Filter, Set of Heat Sinks, Display Cable - 6 foot (Supports up to 4K60p)
- CanaKit USB-C PiSwitch (On/Off Power Switch for Raspberry Pi 4)
Choose an output format
Your service can return JSON while allowing internal exporters for JSON Lines, XML, CSV, or object storage. A stable envelope is easier to evolve:
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
{
"job_id": "job_01J...",
"status": "succeeded",
"records": [{"title": "Example", "price": 19.99}],
"meta": {"url": "https://example.com/item", "worker": "http", "http_status": 200},
"errors": []
}
Build a small HTTP-first reference service
The following Python example is intentionally narrow: it demonstrates a synchronous endpoint, selector-based extraction, and schema validation. Install the dependencies with pip install fastapi uvicorn httpx parsel, save it as main.py, and run uvicorn main:app --reload. Add authentication, a queue, destination controls, and persistent storage before exposing it to untrusted callers.
from typing import Dict, List
from urllib.parse import urlparse
import httpx
from fastapi import FastAPI, HTTPException
from pydantic import BaseModel, Field
from parsel import Selector
app = FastAPI(title="Universal Scraper API")
class ScrapeRequest(BaseModel):
url: str
fields: Dict[str, str] = Field(min_length=1)
timeout_seconds: float = Field(default=20, gt=0, le=90)
class ScrapeResponse(BaseModel):
status: str
records: List[dict]
meta: dict
errors: List[dict]
def validate_url(raw: str) -> str:
parsed = urlparse(raw)
if parsed.scheme not in {"http", "https"} or not parsed.hostname:
raise HTTPException(status_code=400, detail="Only http and https URLs are accepted")
if parsed.username or parsed.password:
raise HTTPException(status_code=400, detail="Embedded credentials are not accepted")
return raw
@app.post("/v1/scrape", response_model=ScrapeResponse)
async def scrape(req: ScrapeRequest):
url = validate_url(req.url)
try:
async with httpx.AsyncClient(
follow_redirects=True,
headers={"User-Agent": "UniversalScraper/1.0"},
timeout=req.timeout_seconds,
) as client:
response = await client.get(url)
except httpx.TimeoutException:
return ScrapeResponse(status="timed_out", records=[], meta={"url": url},
errors=[{"code": "timeout"}])
except httpx.HTTPError as exc:
return ScrapeResponse(status="failed", records=[], meta={"url": url},
errors=[{"code": "fetch_error", "message": str(exc)}])
if response.status_code >= 400:
return ScrapeResponse(status="failed", records=[],
meta={"url": str(response.url), "http_status": response.status_code},
errors=[{"code": "upstream_http", "status": response.status_code}])
selector = Selector(text=response.text)
record = {}
missing = []
for name, css in req.fields.items():
value = selector.css(css).xpath("string(.)").get()
if value is None or not value.strip():
missing.append(name)
else:
record[name] = " ".join(value.split())
meta = {"url": str(response.url), "http_status": response.status_code, "worker": "http"}
if missing:
return ScrapeResponse(status="empty", records=[], meta=meta,
errors=[{"code": "required_fields_missing", "fields": missing}])
return ScrapeResponse(status="succeeded", records=[record], meta=meta, errors=[])
This example fetches one document and one record. A production crawler should move the work to a queue, add pagination as an explicit option, enforce byte and redirect limits, and use a domain scheduler. Scrapy can provide the request scheduling, middleware, selectors, pipelines, and statistics once the HTTP path needs crawl breadth.
Add browser workers only when needed
Detect browser dependence with a documented rule: the required selector is absent from the HTTP response, or the request explicitly asks for rendering or interaction. Send only that job to a Playwright worker. Reuse browser processes carefully, isolate contexts per job, and close pages even on errors.
from playwright.async_api import async_playwright
async def rendered_html(url: str, wait_for: str | None = None) -> str:
async with async_playwright() as pw:
browser = await pw.chromium.launch(headless=True)
context = await browser.new_context()
page = await context.new_page()
try:
await page.goto(url, wait_until="domcontentloaded", timeout=60000)
if wait_for:
await page.wait_for_selector(wait_for, timeout=30000)
return await page.content()
finally:
await context.close()
await browser.close()
Make interaction steps declarative and bounded—for example, a permitted click selector followed by a wait selector. Do not allow arbitrary customer JavaScript in the same security context as your service. Browser jobs consume more CPU and memory than direct HTTP jobs; measure your own workload rather than assuming a universal cost or speed ratio.
Respect robots.txt and target pacing
Robots.txt handling must be an explicit setting. Scrapy’s robots middleware does not automatically apply Crawl-delay or Request-rate directives. Parse those values (where applicable) and translate them into your scheduler’s delay and concurrency settings. Record the effective policy with each job so operators can explain why a request was delayed or refused.
Rank #3
- Not including the Raspberry Pi 5 (8GB), the Crowpi advanced version comes with the Raspberry Pi 5
- ELECROW Black Case for the Raspberry Pi 5, CrowPi is equipped with a 9-inch HD touchscreen along with a camera; All the regular components used in DIY electronics are packed into the CrowPi development board, such as LCD, LED matrix, buzzer, light sensor, PIR sensor, ultrasonic sensor, IR sensor, etc
- Raspberry Pi Sensors: The Crowpi raspberry pi 5 programming kit is jam-packed with lots of buttons such as 19 different sensors in a tidy easy to use package; You don't have to wait and wire things
- Build Quality: Solid ABS shell and well made components in one place make it strong and convenient to travel
- Programming Lessons: This raspberry pi 5 learning kit ships with step by step instructions and provides 21 lessons to take you through identifying components reading code and running it in the terminal
Set separate limits for each domain: maximum in-flight requests, minimum delay, maximum retry count, and a circuit breaker for repeated failures. If a target starts returning throttling responses, reduce concurrency and back off instead of retrying immediately. Prefer a documented API, search endpoint, or bulk export whenever one exists.
Make jobs observable and recoverable
- Metrics: queue wait, fetch latency, browser startup time, status-code distribution, bytes, retries, empty results, and extraction failures, all grouped by target.
- Logs: job ID, target host, worker type, attempt number, policy decision, and a redacted error category. Never log authorization headers or cookies.
- Cancellation: mark a job cancelled in durable state and have workers check that state between pages and interaction steps.
- Retention: set limits for raw HTML, screenshots, cookies, and extracted records. Keep only what your product contract requires.
- Capacity: autoscale HTTP workers separately from browser workers, and cap per-tenant concurrency so one customer cannot monopolize a domain or your fleet.
Use bounded retries. Retry connection resets and selected transient gateway responses; do not retry malformed requests, policy refusals, authentication failures, or deterministic selector errors. Preserve the final error and the last successful response metadata.
Performance, reliability, and cost decisions
Keep the fast path fast
Reuse HTTP connections, stream or cap large responses, avoid downloading unnecessary resource types, and cache only when the caller accepts potentially stale data. For browser jobs, block unneeded resources, reuse a controlled browser pool, and set navigation and overall job deadlines.
Quick wins for a faster PC:
Clear out junk files and repair common Windows errorsFree Scan →Scan for outdated or missing drivers - takes under a minuteDriver Scan →Repair Windows errors before they cause bigger problemsFix Now →Separate freshness from duplication
Accept an idempotency key for submissions and define whether a repeated request returns an existing job, a cached result, or a new fetch. Expose cache age in metadata. Never claim a cache hit is a fresh crawl.
Price from your workload
Compare actual HTTP transfer, browser CPU and memory, queue infrastructure, storage, and third-party endpoint costs. The available framework guidance supports avoiding unnecessary page crawling when an API or export exists, but it does not establish a universal price or performance winner.
Troubleshooting common failures
HTTP returns a page shell with no fields
Cause: data is rendered by JavaScript. Fix: inspect the page’s documented endpoint first; otherwise route the job to Playwright and wait for a specific selector rather than an arbitrary long sleep.
Rank #4
- Fully assembled for plug-and-play operation
- Includes Raspberry Pi 5 with 8GB RAM
- 256 GB PCIe Pi NVMe SSD (Pre-loaded with Pi 64-Bit OS)
- M.2 HAT+
- CanaKit Turbine Black Case for the Pi 5
Results suddenly become empty
Cause: markup changed, a selector is too broad or narrow, or the response is a consent, login, or bot-check page. Fix: store a redacted response sample, classify the page, version the selector rule, and return an explicit extraction failure.
Many 429 or 403 responses
Cause: pacing, access policy, or credentials are wrong. Fix: verify authorization, honor robots and published limits, reduce per-domain concurrency, increase delay, and stop retrying aggressively.
Browser jobs time out
Cause: waiting for a selector that never appears, blocked resources, or an overloaded worker. Fix: capture timing telemetry, use a precise wait condition, set an overall deadline, and move heavy jobs to isolated capacity.
Private or internal addresses are reachable
Cause: destination checks were applied only to the original hostname, not redirects or resolved addresses. Fix: enforce an allow/deny policy before fetch and after every redirect and DNS resolution, with network-level egress controls as a second boundary.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Or skip the browser setup
When your task is taking a clean screenshot rather than extracting arbitrary records, ScreenshotNeo provides a single HTTP call. It accepts cookie or consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each cleanup step can be disabled. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and whether it was billed. It also offers an MCP server with take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients.
PC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minuteSee the parameter reference in the ScreenshotNeo documentation. cURL:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python:
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
Every plan includes the same feature set: full-page and element capture, device presets, custom viewport and retina scale, PDF controls, HTML/CSS rendering, custom CSS and JavaScript, clicks, waits, request blocking, headers, cookies, user agents, authorization, timezone and geolocation, transparent backgrounds, resizing, configurable-TTL caching, signed links, asynchronous webhooks, bulk capture of up to 100 URLs per call, a usage API, and an OpenAPI specification. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 screenshots. Create a free ScreenshotNeo account.
Best Value
- 【What you Get】You will get 1*Pi 5 8GB Single Board,1*RasTech Case,1*Active Cooler,1*Screwdriver,1*Installation instructions,12-month free warranty, lifetime service, 24-hour prompt and friendly response.
- 【More Connectors】There are two USB 3.0 ports(5Gbps simultaneously) and two USB 2.0 ports, which triple total bandwidth ,support any combination of up to two cameras or displays. Peak SD card performance is doubled through support for the SDR104 high-speed mode. It provides a smooth desktop experience for you. Offer Gigabit Ethernet and a PCIe interface, along with dual-band Wi-Fi and Bluetooth 5.0/BLE wireless capability. The RasTech Pi 5 Kit use the new 27W 5.1V 5A USB-C power connector.
- 【 Support Dual 4Kp60 Display 】Each of the two microHDMI sockets can control a 4K display at 60 Hertz, now support HDR, offering super HD video for media streaming projects. RPi 5 is the first RPi model that comes with a PCI Express port (PCIe 2.0 x1 with 500 MB/s) to attach SSDs (requires separate M.2 HAT).
- 【 Excellent Chips And Applications】Pi 5 is a full-size Pi computer using silicon built in-house at Pi. The RP1 “southbridge” provides the bulk of the I/O capabilities for Pi 5. Pi 5 is more friendly and convenient in the development of Internet of Things, Web development, machine identification, automatic control and other electronic equipment applications and network.
- 【 Faster CPU, Better GPU 】 Pi 5 features a Broadcom BCM2712 64-bit quad-core Arm Cortex-A76 processor running at 2.4GHz, it delivers a 2–3× increase in CPU performance relative to RaspberryPi 4. The 800MHz VideoCore VII GPU is compatible to OpenGL ES 3.1 and Vulkan 1.2, substantial uplift in graphics performance. Pi 5 Offers lightning-fast CPU speed, a PCI Express interface, a Real Time Clock (RTC) and a power button and runs significantly cooler than Pi 4.
FAQ
Should every request launch a browser?
No. Use direct HTTP for static documents and reserve browser workers for demonstrated rendering or interaction requirements.
Can one selector rule work forever?
No. Treat selectors as versioned configuration, validate required fields, and alert on empty or malformed results.
Is robots.txt the same as a legal permission?
No single crawler setting resolves every legal or contractual question. Implement an explicit policy process and obtain authorization for the targets you operate on.
When should a scrape be asynchronous?
Use a job queue when work can involve multiple pages, browser rendering, retries, or unpredictable latency. Return a job ID and let clients poll or receive a callback.
Frequently Asked Questions
Should every request launch a browser?
No. Use direct HTTP for static documents and reserve browser workers for demonstrated rendering or interaction requirements.
Can one selector rule work forever?
No. Treat selectors as versioned configuration, validate required fields, and alert on empty or malformed results.
Do these 3 things before closing this tab:
1Scan for outdated or missing drivers - takes under a minute2Repair Windows errors before they cause bigger problems3Fix the driver behind crashes, sound loss and screen glitchesIs robots.txt the same as a legal permission?
No single crawler setting resolves every legal or contractual question. Implement an explicit policy process and obtain authorization for the targets you operate on.
When should a scrape be asynchronous?
Use a job queue when work can involve multiple pages, browser rendering, retries, or unpredictable latency. Return a job ID and let clients poll or receive a callback.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

