Hardware FixRecommendedDevice not working? Your driver may be the problemCheck updates for common hardware issues.Fix DriversOctober DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsClean PCRecommendedOne scan can reveal what keeps slowing WindowsLook for cleanup and repair opportunities.Run Scan×
Skip to content
FastAPI

How to Build a Scraper REST API with Pyppeteer or Selenium

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Build a small Python API that accepts a validated page URL and a constrained set of CSS selectors, waits for the rendered content, and returns extracted fields as JSON. Use Pyppeteer when an asyncio-first Chromium workflow fits your deployment; use Selenium when WebDriver, cross-browser work, or remote browser execution matters. In either case, put limits, timeouts, cleanup, and destination checks around browser work before exposing the endpoint to other callers.

What the API should accept and return

Keep the HTTP interface narrower than the browser. A POST request with a JSON body avoids putting long selector specifications in a query string and gives you room to validate both the destination and extraction instructions.

A starter contract can accept a URL, a named map of CSS selectors, and an optional readiness selector. For each named field, allow only the selector and an extraction mode such as text or an attribute. Return the normalized target and a map of field names to values. Decide deliberately whether a missing field is an empty string, null, or an error; do not let callers guess what an absent value means.

{
  "url": "https://example.com/products/42",
  "wait_for": ".product-title",
  "fields": {
    "title": {"selector": ".product-title", "mode": "text"},
    "price": {"selector": ".price", "mode": "text"},
    "image": {"selector": ".product-image", "mode": "attribute", "attribute": "src"}
  }
}

For predictable failures, return a stable error object, for example {"error":{"code":"selector_timeout","message":"The page did not show the required content in time."}}. Use HTTP status codes consistently: malformed input is a client error, disallowed destinations are forbidden, timeouts can be reported as gateway timeouts, and unexpected browser failures should not reveal stack traces. Include a request or correlation ID so operators can connect the response to safe server-side logs.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Choose Pyppeteer or Selenium

Consideration Pyppeteer Selenium
Execution model Asyncio-oriented Python coroutines for controlling Chromium. WebDriver interface; typical Python calls are synchronous.
Browser and deployment fit Useful when Chromium and an async application fit. The surfaced API reference is version 0.0.25 and says compatibility is best with its bundled Chromium revision; arbitrary browser executables are not guaranteed. Useful when you need WebDriver workflows, supported browser choices, or execution locally and remotely through Selenium Server.
Operational choice Keep browser contexts and pages isolated per job; verify that your package and browser revisions work together before pinning deployment images. Blocking WebDriver work needs a threadpool/worker boundary or a dedicated job system in an async web application.

Neither API model establishes a universal speed advantage. Measure the pages, browser builds, and deployment environment you actually plan to run. Selenium’s documentation describes WebDriver as driving a browser “natively, as a user would,” locally or remotely using Selenium Server. It also documents WebDriver BiDi, a WebSocket-enabled standard protocol for browser events. Pyppeteer’s surfaced documentation is old, so check current package support and browser compatibility before choosing it for a new long-lived service.

Build a bounded FastAPI service with Pyppeteer

This example illustrates the request contract, selector wait, extraction, and cleanup. It launches a browser for each operation, making it simple to understand but not a high-throughput architecture. Install FastAPI, Uvicorn, and a Pyppeteer build compatible with the Chromium you intend to run; the Pyppeteer reference identified here is version 0.0.25, not a current-release recommendation. Confirm the package’s present installation and browser requirements for your platform before deployment.

from typing import Literal
from urllib.parse import urlsplit

from fastapi import FastAPI, HTTPException
from pydantic import BaseModel, Field
from pyppeteer import launch
from pyppeteer.errors import TimeoutError as PyppeteerTimeout

app = FastAPI()

class FieldSpec(BaseModel):
    selector: str = Field(min_length=1, max_length=500)
    mode: Literal["text", "attribute"] = "text"
    attribute: str | None = Field(default=None, max_length=100)

class ScrapeRequest(BaseModel):
    url: str = Field(min_length=1, max_length=2048)
    wait_for: str | None = Field(default=None, max_length=500)
    fields: dict[str, FieldSpec] = Field(min_length=1, max_length=30)

def validate_url(url: str) -> None:
    parts = urlsplit(url)
    if parts.scheme not in ("http", "https") or not parts.hostname:
        raise HTTPException(status_code=422, detail="url must be an absolute http or https URL")
    if parts.username or parts.password:
        raise HTTPException(status_code=422, detail="credentials in URLs are not accepted")
    # This syntactic check is not sufficient protection from SSRF.
    # Enforce destination policy and network egress controls at deployment level.

async def extract(req: ScrapeRequest) -> dict:
    validate_url(req.url)
    browser = None
    page = None
    try:
        browser = await launch(headless=True, args=["--no-sandbox"])
        page = await browser.newPage()
        page.setDefaultNavigationTimeout(20000)
        page.setDefaultTimeout(10000)
        await page.goto(req.url, {"waitUntil": "domcontentloaded", "timeout": 20000})
        if req.wait_for:
            await page.waitForSelector(req.wait_for, {"timeout": 10000})

        values = {}
        for name, spec in req.fields.items():
            values[name] = await page.evaluate(
                """(selector, mode, attribute) => {
                    const el = document.querySelector(selector);
                    if (!el) return null;
                    if (mode === 'attribute') return el.getAttribute(attribute);
                    return (el.innerText || el.textContent || '').trim();
                }""",
                spec.selector, spec.mode, spec.attribute
            )
        return {"url": req.url, "fields": values}
    except PyppeteerTimeout:
        raise HTTPException(status_code=504, detail="navigation or selector wait timed out")
    except HTTPException:
        raise
    except Exception:
        # Log a correlation ID and safe diagnostic details server-side.
        raise HTTPException(status_code=502, detail="browser could not complete the scrape")
    finally:
        if page is not None:
            try:
                await page.close()
            except Exception:
                pass
        if browser is not None:
            try:
                await browser.close()
            except Exception:
                pass

@app.post("/scrape")
async def scrape(req: ScrapeRequest):
    return await extract(req)

Run locally with uvicorn app:app --host 127.0.0.1 --port 8000, assuming the code is saved as app.py and dependencies are installed. Test with a permitted page and a request such as:

curl -X POST http://127.0.0.1:8000/scrape 
  -H 'Content-Type: application/json' 
  -d '{"url":"https://example.com","wait_for":"h1","fields":{"heading":{"selector":"h1","mode":"text"}}}'

The response is JSON with the target URL and a fields object. Missing selectors in the extraction map produce null in this example; only the optional wait_for selector causes a timeout when absent.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Validate destinations before production

The code checks URL syntax and rejects credentials, but that is not an SSRF defense. A caller-controlled browser URL can target loopback, private or link-local networks, internal services, or cloud metadata endpoints. Resolve and evaluate destinations under an explicit allow/deny policy, account for redirects and DNS rebinding, and enforce network egress restrictions outside the application process. Do not rely on one hostname check before navigation as a complete security boundary. Also authenticate callers, rate-limit by client, cap request size and selector count, and limit response size. Have the policy reviewed against your deployment and applicable security requirements.

Do not expose an unrestricted fetch proxy. Use only authorized sources, check site terms and applicable rules, and do not bypass access controls or disguise automation. If a site denies access, seek permission or use an authorized API.

Replace per-request launch when the workload grows

The sample’s launch-and-close lifecycle makes cleanup visible, but browser startup is repeated. A developer asking about a REST scraping server reported in a 2021 Stack Overflow question that opening and closing Chrome for every request delayed responses and used many resources. That is one person’s report, not a general benchmark. For higher volume, measure browser startup, navigation, extraction, memory, and failures, then move toward bounded worker processes, a queue, or remote browser sessions.

Do not share a mutable page between unrelated callers. If you retain a browser process, create isolated contexts or otherwise isolate cookies and page state, close each page/context in a finally path, and restart unhealthy workers. Set a maximum number of active jobs and a queue limit; when full, reject or defer work rather than accepting unlimited browser sessions. There is no universal safe worker count: it depends on page behavior, memory, CPU, browser build, and host limits.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use readiness conditions, not arbitrary sleeps

domcontentloaded means the initial document was parsed; it does not guarantee that a client-rendered product listing or other target data has appeared. Wait for a selector or application-specific condition that indicates the content you need, then extract it. Pyppeteer documents navigation waits and waitForSelector; its click/navigation guidance also warns that a navigation wait can race with the click that triggers navigation, so coordinate those operations when a scrape requires interaction.

A fixed delay can be useful as an additional bounded pause for a known page behavior, but it should not be the sole readiness test. Distinguish “the page loaded and the field is absent” from “the required selector never appeared.” For interactive sites, content may require scrolling, clicking, or authentication; there is no generic recipe guaranteed to work across sites. Keep any interaction specific, authorized, and covered by a timeout.

Use Selenium without blocking FastAPI’s event loop

FastAPI recommends async def for path operations that await an async library and ordinary def when the library does not support await. Pyppeteer fits the first case. Selenium’s common Python WebDriver calls are synchronous, so do not call them directly inside an async route and block the event loop.

For a low-volume service, a synchronous FastAPI path operation can run in the framework’s sync execution path; for a larger service, put Selenium work behind a controlled threadpool or dedicated worker process. A worker boundary makes it easier to cap simultaneous sessions, apply queue limits, isolate browser crashes, and set job deadlines. Remote WebDriver can put browser execution on another machine or Selenium Server, but it does not itself provide queueing, resource limits, retries, or cleanup.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Keep the service contract and extraction rules independent of the browser implementation. That lets you switch between a Pyppeteer coroutine and a Selenium worker without changing callers’ JSON. With either driver, explicitly wait for the content condition, set navigation and selector deadlines, close the page/session even after errors, and return controlled error codes rather than raw traces.

Timeouts, limits, and failure handling

A production endpoint needs limits at more than one layer. Choose values from measured behavior and service objectives rather than treating the example timeouts as universal recommendations.

  • Input: require HTTP or HTTPS, cap URL length, field count, selector length, and request-body size; reject unsupported extraction modes.
  • Destination: allow only destinations your service is permitted to fetch, and block access to internal address ranges through both application checks and network egress policy.
  • Execution: set navigation and selector timeouts, an overall job deadline, a maximum number of active browsers, and a bounded queue.
  • Output: cap extracted text or response size and avoid returning whole page HTML unless the product contract explicitly requires it.
  • Operations: authenticate callers, apply per-client request limits, record correlation IDs and safe diagnostics, and monitor queue wait, browser startup, navigation, extraction, memory, and failure rates.

Handle malformed URLs, blocked destinations, DNS and connection failures, navigation errors, HTTP error responses, missing selectors, timeouts, browser crashes, and oversized output separately where the distinction helps clients recover. Never expose credentials, cookies, internal addresses, or raw browser stack traces in public error responses.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Common problems and fixes

  • The route hangs or other requests stall: synchronous Selenium may be blocking an async route. Move it to a synchronous path, controlled threadpool, or worker queue; cap concurrency.
  • Fields are empty although the page looks populated: the page may render them after initial document parsing. Wait for the field’s selector or a meaningful page state before extraction.
  • Every job is slow or memory use climbs: inspect whether each request launches a full browser and whether pages, contexts, or processes are left open. Measure the workload, close resources reliably, and test a bounded worker design rather than increasing parallelism blindly.
  • Navigation wait times out: the site may be slow, unavailable, redirecting, or waiting on resources. Set a finite navigation deadline, inspect safe server-side diagnostics, and do not treat an arbitrary longer timeout as a fix for every failure.
  • The browser cannot launch in deployment: verify that the browser binary, operating-system dependencies, permissions, and automation package are compatible. Pyppeteer 0.0.25’s reference cautions that its bundled Chromium version is the compatibility target; do not assume an arbitrary system Chromium will work.
  • A caller can reach an internal service: syntactic URL validation is not enough. Add destination policy, redirect-aware checks, and network egress restrictions before allowing untrusted callers.
  • Clicking a control races with navigation: arrange the navigation wait and click so the wait is active for the navigation event, and bound both operations.

When a managed browser API is a better fit

If you do not want to maintain browser binaries and browser workers, a managed browser API is another deployment model. Browserless documents HTTP endpoints for rendered HTML, selector extraction, screenshots, PDFs, and related tasks; its REST overview describes stateless, single-action operations that launch a browser, perform a task, and close the session. Its selector scrape flow runs client-side JavaScript and waits for selectors before returning selected text, HTML, or attributes as JSON. The documentation gives a default selector wait of up to 30 seconds for that endpoint.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Compare self-hosting and managed execution on the controls your application needs: data handling, isolation, latency, request limits, pricing, operational responsibility, and vendor dependency. No price or head-to-head performance result is established here, so assess the provider’s current terms and documentation directly. A managed browser task can simplify infrastructure, but your API still needs validation, caller authentication, output limits, and an intentional response contract.

Or skip the browser setup

If your goal is a screenshot or PDF rather than structured field extraction, ScreenshotNeo is a website screenshot API and MCP server from Yorker Media. It does not replace the selector-based JSON scraper above; it is useful when the output you need is a visual capture. One GET request can return a PNG, JPEG, WebP, or PDF. Example using the documented API pattern:

curl -G "https://api.screenshotneo.com/v1/shot" 
  -d access_key=YOUR_API_KEY 
  --data-urlencode url=https://example.com 
  -o shot.webp

See the ScreenshotNeo API documentation for request options. It accepts cookie/consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be disabled. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing status. Its MCP server provides take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients. The Free plan includes 1,000 shots per month with no card; paid plans start at $5 for 3,000 shots.

Sign up for ScreenshotNeo’s free plan to try 1,000 screenshots a month with no card.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Frequently Asked Questions

Can I use this endpoint to scrape any website?

No. Fetch only pages you are authorized to access, and follow applicable site terms and rules. The API’s destination policy should also restrict where callers can direct the browser.

Should I return extracted data or the whole rendered page?

For a stable client-facing contract, return the named fields the caller requested. Return full HTML only when your use case specifically needs it and you have bounded its size and exposure.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Read next

Recommended PC Tool
Recommended PC Tool
PC Slower Than It Used to Be?Free scan - under a minute
Crashes, No Sound, or Screen Glitches?Free driver scan

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.