Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

Use a webhook to turn a scrape run into an event your application can consume: start the run, save its provider run ID, configure success and failure callbacks, authenticate and validate the request, record it idempotently, return 2xx immediately, and let a durable worker fetch results and update your systems. Event names, payloads, timeout limits and retry schedules belong to the scraping provider, not to webhooks in general.

What a scraping webhook does

A webhook is a provider-initiated HTTP request sent when a configured event occurs. Instead of polling a scrape endpoint every few seconds, your application exposes a callback URL and receives a POST, commonly with JSON describing the run and event.

Apify’s API is a concrete example: its webhook creation endpoint accepts a requestUrl, selected eventTypes and a condition, then posts JSON to the target when an Actor, task or run changes state. See the Apify create-webhook reference for the current contract. Available run events documented there include success, failure, abort, timeout and resurrection.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Do not copy Apify’s names or delivery behavior blindly to another service. Treat them as provider-specific and verify the current documentation for your scraper.

The production flow

  1. Create a correlation record. Generate your own job or request ID. Start the scrape and persist that ID beside the provider’s run ID, target URL, tenant and expected result location.
  2. Configure the callback. Subscribe to the states your product needs. Success and failure are typical; include timeout or abort when users must see those outcomes. If your provider supports idempotency for webhook creation, use it. Apify documents an idempotencyKey to prevent repeated creation calls from producing duplicate webhook definitions.
  3. Protect the endpoint. Use HTTPS and a secret credential, such as a high-entropy token in a header or URL parameter supported by the provider. Apify recommends a secret token. Never log the token. A token is not the same as a cryptographic signature; claim signature verification only when your selected provider documents signed requests.
  4. Validate before enqueueing. Check the method, content type, credential, JSON shape, provider name, event type and run ID. Reject malformed or unauthenticated requests without creating work.
  5. Deduplicate atomically. Persist a stable provider event ID when one exists. Otherwise derive a key from provider, run ID, event type and a provider delivery identifier or payload hash. Enforce a unique database constraint and make every downstream action safe to repeat.
  6. Acknowledge quickly. Once the event is authenticated and durably recorded or placed on a durable queue, return a 2xx response. Do not download large result sets, run transformations or call several downstream APIs inside the webhook request.
  7. Process asynchronously. A worker consumes the queue, fetches or reads the scrape output, transforms it, writes application data and records completion. Keep retries and dead-letter handling in this worker layer.
  8. Reconcile independently. Keep a status lookup or durable run ledger so delayed or exhausted notifications cannot leave a job permanently unknown. Use the provider’s status and result API, or another authoritative record, to repair gaps.

Provider delivery rules are not universal

Apify documents that a webhook endpoint must return a 2xx response. A failed request is retried with exponential backoff: approximately one minute, two minutes, four minutes and so on, through an eleventh retry at about 32 hours, after which retries stop. Its HTTP request timeout is two minutes. These are Apify values, not a general webhook standard. Apify advises immediate responses and an internal queue for lengthy work.

Delivery can occur more than once. Apify’s webhook action guidance says: “In rare cases, the webhook might be invoked more than once. Design your code to be idempotent to handle duplicate calls.” Assume at-least-once delivery unless your provider explicitly guarantees something stronger.

What to compare before choosing a provider

Question Why it matters
Which events exist? Success, failure, timeout and abort determine what your state machine can represent.
What identifies an event? A stable ID makes deduplication reliable; otherwise define a deterministic key.
How do you obtain results? Payloads may contain data, a result URL, or only a run reference.
What are timeout and retries? They define how fast you must acknowledge and how long recovery may take.
How is the endpoint authenticated? Check token, header, signature and rotation support in the current contract.
How are missed events repaired? A status/result API or export ledger is essential for reconciliation.

A minimal receiver with a durable queue

The following Python example uses Flask and a database abstraction named events. Replace the storage calls with your database and queue client. The important properties are validation, an atomic uniqueness check, fast acknowledgement and worker-side processing.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
import os, json
from flask import Flask, request, jsonify

app = Flask(__name__)
WEBHOOK_TOKEN = os.environ["SCRAPER_WEBHOOK_TOKEN"]

@app.post("/webhooks/scraper")
def scraper_webhook():
    if request.headers.get("Authorization") != f"Bearer {WEBHOOK_TOKEN}":
        return jsonify(error="unauthorized"), 401
    if not request.is_json:
        return jsonify(error="expected JSON"), 415

    body = request.get_json(silent=True) or {}
    run_id = body.get("runId") or body.get("run_id")
    event = body.get("eventType") or body.get("event")
    event_id = body.get("id") or f"{run_id}:{event}"
    if not run_id or not event:
        return jsonify(error="missing run or event"), 400

    # INSERT ... ON CONFLICT DO NOTHING must be atomic.
    inserted = events.insert_once(
        event_id=event_id, run_id=run_id, event_type=event,
        payload=json.dumps(body)
    )
    if inserted:
        queue.publish("scrape-events", {"event_id": event_id})
    return jsonify(accepted=True), 202

@app.errorhandler(Exception)
def handle_error(error):
    app.logger.exception("webhook failure")
    return jsonify(error="temporary failure"), 500

Return a 2xx only after the event is safely stored or queued. If persistence is unavailable, return a non-2xx so a provider with retries can try again. Do not acknowledge first and then hope an in-memory task survives a process crash.

Worker logic

def handle_event(message):
    event = events.get(message["event_id"])
    if event.processed_at:
        return                         # safe duplicate delivery
    run = provider.get_run(event.run_id)
    if event.event_type == "success":
        result = provider.get_result(run)
        records = transform(result)
        application_store.upsert(records, key="source_record_id")
    elif event.event_type in {"failure", "abort", "timeout"}:
        application_store.mark_failed(event.run_id, event.event_type)
    events.mark_processed(event.event_id)

Use bounded worker retries, exponential backoff and a dead-letter queue. Preserve the original payload and provider run ID for diagnosis. If a result can change after a notification, fetch it only after checking the run state and use versioning or upserts to avoid stale writes.

Configuring an Apify-style webhook

Apify’s API uses a POST to its webhook resource with a request URL, event types and a condition that selects the Actor, task or run. Consult the live API reference for required fields and authentication. Its Python SDK documentation also covers webhook concepts at the webhook guide.

At job creation time, save both identifiers:

job_id = create_local_job(target_url)
run = apify.start_actor(input={"url": target_url})
link_provider_run(job_id, run["id"])
# Create a webhook conditioned on this run, with success/failure events

Keep the callback URL stable while deployments change behind it. During secret rotation, accept the old and new credential for a short overlap, then remove the old one. Rate-limit unauthenticated requests and cap body size before parsing JSON.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Testing, observability and recovery

  • Test each terminal state: success, failure, timeout and abort, plus malformed JSON and invalid credentials.
  • Test duplicates: deliver the same payload twice and verify one queue message and one business effect.
  • Test slow dependencies: make the worker’s result API call hang; the callback should still return within its provider window.
  • Measure the path: log event ID, run ID, receive time, acknowledgement latency, queue time, processing attempts and final state. Exclude secrets and sensitive scraped content.
  • Alert on gaps: alert on rising non-2xx responses, queue age, dead-letter count and runs with no terminal event. A scheduled reconciliation job should compare active local jobs with provider status.

Apify’s finite retry behavior makes reconciliation important: once its documented retry limit is exhausted, a transient outage will not automatically deliver the event again. The exact recovery API depends on your provider.

Scraping APIs that do not document webhooks

ScrapingBee’s official HTML API documentation describes request-response scraping and an Spb-request-id returned on responses, including errors; it recommends retrying a 500 response. That documentation does not establish callback webhooks. Do not design around webhook events until the provider confirms event names, authentication, retries and result lookup. You can still wrap a request-response call in your own job queue and emit an internal webhook when your worker finishes.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup

If your workflow needs a rendered page image or PDF as one step, ScreenshotNeo provides a GET-based screenshot API and an MCP server for AI agents. It accepts consent banners before capture and removes more than 60 known consent platforms, newsletter popups and chat widgets; each step can be disabled. Only clean shots are billed: bot checks or CAPTCHAs, blank pages, timeouts, failed loads and cache hits cost nothing, and response headers report the page verdict and billing status.

One call is enough (see the ScreenshotNeo API documentation):

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

It supports full-page captures with lazy images loaded, CSS-selector elements, dark mode, 12 device presets or custom viewports, retina scale, PDF paper settings and page ranges, HTML/CSS rendering, custom JavaScript, clicks, hidden selectors, waits, request blocking, headers, cookies, user agents, authorization, timezone, geolocation, transparent backgrounds, resizing, configurable caching, signed image links, asynchronous jobs with signed webhooks, bulk capture of up to 100 URLs per call, usage reporting and an OpenAPI specification. Its parameter names are compatible with those used by other screenshot APIs, easing migration. An MCP server exposes take_screenshot, get_page_info and capture_pdf to Claude, Cursor and other MCP clients.

The Free plan includes 1,000 shots each month with no card. Paid plans start at $5 for 3,000 shots; yearly billing gives two months free, and every feature is on every plan. Sign up free for ScreenshotNeo.

Troubleshooting

No webhook arrives

Confirm the event condition matches the actual run, the URL is publicly reachable over HTTPS, DNS and firewall rules allow the provider, and the provider accepted the webhook definition. Inspect delivery logs and use the status API to distinguish a run that never finished from a notification that failed.

The provider reports delivery failure

Check that your handler returns 2xx before its timeout, does not require a browser session or CSRF token, accepts the provider’s content type and can parse its actual field names. Move all slow work to the queue.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The same event is processed twice

Add a database uniqueness constraint on the event key, perform insert-and-enqueue atomically, and make result writes idempotent with upserts. Never rely on an in-memory “seen” set.

A job is stuck after retries stop

Run reconciliation against the provider’s status and result endpoints, then replay a locally stored event or synthesize an internal recovery event. Record the repair so operators can distinguish it from an original delivery.

Authentication suddenly fails

Check secret rotation, proxy header forwarding and URL encoding. Keep credentials in a secret manager, redact them from logs and accept overlapping keys only for the planned rotation window.

FAQ

Should the webhook download the scraped data?

No. Acknowledge after durable recording, then let a worker fetch or read the result. This avoids provider timeout limits and makes retries controllable.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Can I assume exactly-once delivery?

No. Design for duplicate and out-of-order notifications unless your provider explicitly documents stronger guarantees.

What if my scraper has no webhook feature?

Use a queue and polling worker, then emit your own internal event when the request-response operation reaches a terminal state.

Frequently Asked Questions

How do I get notified when a scrape finishes?

Configure the provider’s success webhook for the run, store the provider run ID, validate and persist the POST, return 2xx quickly, and let a worker process the result.

How do I handle duplicate webhook events?

Use a unique event key with an atomic insert, make downstream writes idempotent, and treat every delivery as at least once.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.