Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

An asynchronous crawler API lets you submit a crawl or extraction request, receive a run ID, and retrieve the results after the provider finishes. The practical pattern is to save that ID, poll for completion or use a documented callback, then validate and store the returned data. Use browser rendering only when the information you need appears after JavaScript runs; a plain HTTP fetch cannot see content that exists only in the rendered page.

What asynchronous extraction changes

A synchronous API keeps the request open while it fetches and processes a page, then returns the result. An asynchronous API separates those steps: your application submits work and gets an acknowledgement, while the crawl continues independently. Your application later checks its status and obtains the output.

This model is useful when a crawl may take longer than a normal web request, when a batch contains many URLs, or when you want workers to process results without holding a user-facing connection open. It does not make extraction itself more accurate or guarantee that a page can be accessed. It changes how work is scheduled and results are collected.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The lifecycle

  1. Submit: send the target URL or batch and the provider-specific extraction settings.
  2. Record: persist the returned run ID along with the input, options, creation time, and an application-generated idempotency key.
  3. Wait: poll a status endpoint with bounded backoff, or receive a callback if the provider documents one.
  4. Fetch: once complete, retrieve the structured response or dataset items.
  5. Validate and save: check the output contract and write accepted records to your database or warehouse.

Scrapy.io documents this managed lifecycle: an asynchronous scraper run, status polling through GET /v1/runs/{runId}, dataset retrieval through GET /v1/runs/{runId}/dataset/items, and recurring schedules. Its documentation also describes a batch POST endpoint. The exact host, authentication, submission body, and response fields are provider-specific; use the current API documentation for those values rather than assuming that every service uses the same contract.

Choose HTTP or browser rendering based on the data

Start with the simplest method that can actually see the information. If the server’s HTTP response already contains the required HTML or JSON, direct HTTP extraction avoids the work of running a browser. If JavaScript creates or changes the content after the initial response, use a browser-capable extraction mode. Zyte puts the distinction plainly: “HTTP responses do not reflect HTML content rendered by a web browser that executes JavaScript code.” Its API reference describes both HTTP and browser extraction modes, as well as automatic extraction types such as article, product, job-posting, and SERP data. The documented extraction endpoint is https://api.zyte.com/v1/extract.

Use direct HTTP when

  • The needed fields are present in the response HTML or a public JSON response.
  • You can parse the source without waiting for client-side scripts or user interactions.
  • Lower operational complexity matters more than browser-specific behavior.

Use a browser when

  • Required text or elements appear only after JavaScript executes.
  • The page needs browser state such as cookies, a session, or a particular user-agent context.
  • Your extraction depends on behavior the provider’s browser mode supports; verify the exact controls in that provider’s documentation.

Rendering does not mean every page is accessible. A site may still deny requests, require authorization, display a challenge, or change its markup. Treat a missing field as a result to investigate, not proof that the page had no data.

Submit, poll, and retrieve without losing track of work

Keep the provider-specific HTTP details behind a small adapter in your application. That makes it easier to switch providers or update an endpoint without mixing API plumbing into parsing and business logic. The following Python example shows the orchestration pattern for an API that implements the illustrated contract. It is a generic adapter example, not a claim about the precise submission schema or field names of Scrapy.io or Zyte. Configure the URLs and payload to match the provider’s current documentation.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
import os
import time
import requests

API_TOKEN = os.environ["CRAWLER_API_TOKEN"]
SUBMIT_URL = os.environ["CRAWLER_SUBMIT_URL"]
# Set these to provider-documented URL patterns, including {run_id}.
STATUS_URL = os.environ["CRAWLER_STATUS_URL"]
RESULTS_URL = os.environ["CRAWLER_RESULTS_URL"]

session = requests.Session()
session.headers.update({
    "Authorization": f"Bearer {API_TOKEN}",
    "Content-Type": "application/json",
    "Accept": "application/json",
})

def run_crawl(target_url, payload, timeout_seconds=900):
    # payload must use the exact schema required by your provider.
    response = session.post(SUBMIT_URL, json=payload, timeout=30)
    response.raise_for_status()
    submitted = response.json()
    run_id = submitted["runId"]  # Map this to the provider's actual ID field.

    deadline = time.monotonic() + timeout_seconds
    delay = 1.0
    while time.monotonic() < deadline:
        status_response = session.get(
            STATUS_URL.format(run_id=run_id), timeout=30
        )
        status_response.raise_for_status()
        status_data = status_response.json()
        state = status_data["status"]  # Map to the provider's status field.

        if state == "succeeded":
            result_response = session.get(
                RESULTS_URL.format(run_id=run_id), timeout=60
            )
            result_response.raise_for_status()
            records = result_response.json()
            return {"run_id": run_id, "url": target_url, "records": records}
        if state in {"failed", "cancelled"}:
            raise RuntimeError(f"Run {run_id} ended with status {state}: {status_data}")

        time.sleep(delay)
        delay = min(delay * 2, 30.0)

    raise TimeoutError(f"Run {run_id} did not finish before the client deadline")

if __name__ == "__main__":
    target = "https://example.com"
    # Replace with a documented provider payload and validate returned records.
    result = run_crawl(target, {"url": target})
    print(result)

This script deliberately uses an application deadline and caps polling intervals; tune them to the provider’s published limits and expected job duration. In production, save the run ID immediately after submission, before beginning to poll. If the process stops, a separate worker can resume by loading the saved run ID instead of creating a second crawl.

Rank #2
Sale
HTML and CSS: Design and Build Websites
  • HTML CSS Design and Build Web Sites
  • Comes with secure packaging
  • It can be a gift option

Use bounded polling

Polling every few milliseconds wastes requests and may run into provider rate limits. A short initial delay followed by exponential backoff reduces needless checks while keeping a ceiling on how long each interval grows. Add jitter when many workers may poll together, and stop at a defined deadline. If the provider supports callbacks or webhooks, use them when their delivery and retry behavior fit your system; still provide a recovery path for missed callbacks.

Make submission recoverable

A network timeout during submission is ambiguous: the provider may have accepted the request even if your client never received the run ID. Do not blindly repeat a non-idempotent submission. Use a provider-supported idempotency mechanism if documented. Otherwise, persist an application request key and the exact submission details, then reconcile uncertain submissions using whatever lookup facility the provider offers. Keep retries separate from creation so a transient polling error does not accidentally start another crawl.

Validate results before they reach downstream systems

Successful job status only means the provider says the run completed; it does not prove that the output is complete or suitable for your application. Validate the returned records at the boundary, before writing them into a warehouse or serving them to users.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Confirm the output is valid JSON or the expected format and conforms to your schema.
  • Check required fields, data types, and sensible value ranges.
  • Retain the source URL and capture or extraction timestamp for provenance.
  • Detect duplicate records using stable business keys, not just array position.
  • Record the run ID, input options, status, and provider error payload for diagnosis.

For changing websites, version your own data contract and monitor field coverage. A page redesign can produce a technically successful run with missing or renamed fields. Send malformed or partial data to a quarantine path rather than silently replacing a good prior record.

Hosted crawler API or a self-managed Scrapy project?

A hosted service trades some control for less infrastructure to operate. Zyte’s usage documentation lists hosted capabilities including proxies and IP controls, geolocation, cookies, sessions, browser automation, screenshots, and automatic structured extraction. Its developer page describes a scriptable headless browser and automatic extraction for articles, products, and job listings. Check the provider’s current terms for exact availability and limits.

Self-managed Scrapy is a better fit when your team needs to own spider code, parsing rules, schedules, deployment, and data contracts. Scrapy’s current API includes crawl_async() and asyncio-compatible runner classes whose tasks finish when crawling finishes. That control also leaves your team responsible for the scheduler, storage, observability, browser and proxy layer if needed, and failure handling.

Scrapy.io occupies a middle ground: its official documentation describes managed synchronous or asynchronous scraper runs, polling by run ID, dataset item export, and recurring schedules. Consider the following decision points before choosing:

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Execution: Is an immediate response sufficient, or do you need a background job and later retrieval?
  • Rendering: Is the required data in the HTTP response, or does it require JavaScript-capable browser execution?
  • Control: Do you need custom spider logic and parsing, or provider-managed extraction fields?
  • Operations: Who will own proxies, sessions, retries, concurrency, storage, and monitoring?
  • Output: Do you want raw HTML or JSON under your own parser, or a provider-defined structured response?
  • Scale and cost: Compare current concurrency limits, request pricing, and dataset retention directly in the applicable plan terms; no price or benchmark is assumed here.

Performance, reliability, and cost considerations

Asynchronous execution improves how your system handles waiting; it does not automatically make the crawl faster. End-to-end time can include queueing, fetching, browser rendering, extraction, and result transfer. For throughput, measure each stage in your own workload and respect the provider’s concurrency and rate limits. Keep the number of in-flight runs bounded so a large batch does not overwhelm your account, your own result processor, or the target sites.

Rank #4
Sale
Web Design with HTML, CSS, JavaScript and jQuery Set
  • Brand: Wiley
  • Set of 2 Volumes
  • A handy two-book set that uniquely combines related technologies Highly visual format and accessible language makes these books highly effective learning tools Perfect for beginning web designers and front-end developers

For reliability, separate transient failures from permanent ones. A brief network error or rate limit may justify a delayed retry; an authorization failure, invalid request, or persistent access denial usually requires changing configuration or obtaining access, not repeating the same call. Retry only operations that are safe to repeat, preserve the original run ID and error detail, and set a maximum attempt count. Avoid logging credentials or sensitive page data.

Costs depend on the provider’s current billing unit and plan terms. Before production, establish whether billing counts submitted URLs, completed pages, browser time, retries, or another unit, and how failed or duplicate work is treated. Dataset retention and concurrency can also affect architecture even when they are not billed as separate line items. The available technical sources do not establish comparable prices or statistical success rates, so verify current terms rather than relying on a generic estimate.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Compliance and access are part of the design

Before crawling, confirm that you are authorized to access the pages and process the resulting data. Review the target site’s terms and robots guidance, apply reasonable rate limits, and take extra care with personal or sensitive information. A hosted API can supply technical access mechanisms; that does not transfer your responsibility to decide whether a particular collection is permitted. Store only what your use case needs, set retention rules, and restrict access to raw results.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Or skip the browser setup

If the task is to capture a page visually rather than extract structured records, ScreenshotNeo is a website screenshot API and MCP server, not a general-purpose crawler or structured-data extractor. It can return a PNG, JPEG, WebP, or PDF from a URL in one GET request. Its cleanup options accept cookie or consent banners like a visitor and remove 60+ known consent platforms, newsletter popups, and chat widgets; each cleanup step can be turned off. The response identifies whether a page was a bot check, blank, timeout, failed load, or cache hit, and only clean shots are billed.

For a one-call screenshot, see the ScreenshotNeo API documentation:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

ScreenshotNeo also has an MCP server with take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000. Cookie banners, popups, and chat widgets are removed before the shot; bot checks, blank pages, and failed loads are never billed. Sign up free for 1,000 screenshots a month with no card.

Frequently Asked Questions

Does asynchronous mean the crawl runs in parallel?

Not necessarily. It means submission and result retrieval are separate; parallelism depends on your provider, account limits, and how you schedule runs.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Can I use a screenshot API to get structured product or article fields?

A screenshot is visual output, not a structured extraction result. Use a crawler or extraction API for records such as titles, prices, or article fields; use a screenshot API when the page image or PDF is the desired output.

Should I keep raw crawl results after parsing?

Keep them only if your audit, debugging, or retention requirements justify it. Apply access controls and a retention period, especially when page content may contain personal information.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.