Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A web-scraping API webhook lets the provider notify a URL you control when a configured event occurs—such as a job succeeding or failing. Build the receiver to validate the request, record or enqueue the event, and return a successful HTTP response quickly. Then fetch the scrape result through the provider’s documented result mechanism. The notification and the result are separate steps, and delivery, retry, and payload behavior vary by provider.

What a scraping API webhook does

In a webhook workflow, your application starts a scrape or configures an event subscription, and the provider sends an HTTP request to your receiver when that event occurs. Your server acknowledges the delivery; your application can then do slower work, such as fetching data, storing it, or notifying a user.

This is different from repeatedly polling a job-status endpoint. A callback can reduce unnecessary status checks, but it does not necessarily contain the scraped data. For example, Bright Data documents an asynchronous flow in which a trigger returns a snapshot ID; the client can check progress and download results after the snapshot is ready. A notify URL can provide a completion notification. The documented result-retrieval flow still matters.

There is no universal scraping-webhook contract. Check each provider’s documentation for event names, payload fields, authentication, response requirements, retries, timeout, and result retrieval before you build against it.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Plan the workflow before configuring a callback

  1. Choose the event. Decide whether the receiver needs successful completion, failure, or another lifecycle event. Scope it to the relevant job, task, or scraper where the provider allows that.
  2. Decide what the callback should trigger. Usually the callback should mark a job’s state and enqueue follow-up work, rather than download or process a large result inside the request handler.
  3. Choose the result source. Determine whether the provider puts results in a storage location, exposes a result endpoint, or requires a separate download step. Record any job or snapshot identifier needed to retrieve them.
  4. Design for repeat delivery. Assume a callback might arrive more than once unless the provider explicitly guarantees otherwise. Use a stable event or job identifier and make downstream updates safe to repeat.
  5. Protect the receiver. Use HTTPS, keep secrets out of source code, validate that a request is expected, and avoid accepting arbitrary callback data as trusted instructions.

Configure a provider-specific webhook

Apify: create a webhook for an event

Apify documents a webhook-creation API that accepts a request URL, event types, and a condition. Actor-run and build event types are among the available options. The create request uses JSON with Content-Type: application/json. Its payload template can include defined variables such as the event type, event data, and the resource that triggered the event; the resulting template must be valid JSON.

Use the current Apify documentation for the exact endpoint, authentication, accepted event names, condition syntax, and template variables. The essential configuration has this shape, but it is illustrative rather than a complete API request:

{
  "requestUrl": "https://your.example.com/webhooks/scraper",
  "eventTypes": ["<documented event type>"],
  "condition": { "<provider-specific condition>": "<value>" }
}

Do not send the illustrative placeholders as a real request. Replace them with values supported by your chosen event and the current API schema. Apify also supports an idempotency key when creating a webhook, which can prevent duplicate webhook records if your client repeats the create request. That creation feature does not deduplicate callback deliveries at your receiver.

Bright Data: connect notification to snapshot retrieval

Bright Data documents an asynchronous Web Scraper API flow: trigger a job, retain its snapshot ID, monitor progress, and retrieve the result when it is ready. Its documentation also describes a notify URL for completion notification. The progress flow exposes starting, running, ready, and failed states, and the documented API uses bearer-token authorization.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use Bright Data’s current endpoint reference for the exact request and notification payload. Treat the notification as a signal to advance the workflow, not as proof that result bytes are included. Associate it with the relevant snapshot, check or otherwise confirm the documented state, and fetch the result using the documented retrieval method.

Build a receiver that acknowledges fast

The receiver should do only the synchronous work needed to accept a delivery: validate it, persist enough information to recover, enqueue the work, and return a 2xx response. A durable queue or equivalent persistent job record helps prevent a server restart from losing work after acknowledgment.

Here is a small Node.js HTTP receiver skeleton using only built-in modules. It shows the control flow, not a provider-specific signature check or payload schema. Replace the validation and enqueue functions with implementations appropriate to your provider and infrastructure. In production, use a durable queue or database-backed job table rather than the in-memory set shown here.

const http = require('node:http');

const seen = new Set(); // Demonstration only: not durable or multi-process safe.

function isExpected(req) {
  // Replace with provider-specific authentication and request validation.
  return req.headers['x-webhook-secret'] === process.env.WEBHOOK_SECRET;
}

async function enqueue(event) {
  // Persist the event or enqueue it durably before acknowledging it.
  console.log('Queued event:', event.id);
}

const server = http.createServer(async (req, res) => {
  if (req.method !== 'POST' || req.url !== '/webhooks/scraper') {
    res.writeHead(404).end();
    return;
  }

  if (!isExpected(req)) {
    res.writeHead(401).end();
    return;
  }

  let raw = '';
  for await (const chunk of req) raw += chunk;

  let event;
  try {
    event = JSON.parse(raw);
  } catch {
    res.writeHead(400).end('Invalid JSON');
    return;
  }

  // Map this to a stable identifier actually present in your provider's payload.
  const eventId = event.id;
  if (!eventId) {
    res.writeHead(400).end('Missing event identifier');
    return;
  }

  if (seen.has(eventId)) {
    res.writeHead(200).end('Already accepted');
    return;
  }

  try {
    await enqueue(event);
    seen.add(eventId);
    res.writeHead(200).end('Accepted');
  } catch (err) {
    // A non-2xx response can prompt provider retries; do not acknowledge lost work.
    res.writeHead(500).end('Could not queue event');
  }
});

server.listen(process.env.PORT || 3000);

Do not assume a provider sends an x-webhook-secret header or an id field: those names in the skeleton are examples. Apify documents a secret token in the webhook URL as a security measure and also permits a headers template, subject to provider-controlled headers. Use the authentication method and fields your provider actually documents. If a secret is in a URL, protect it as a credential: do not expose it in logs, analytics, or public links.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

For reliable deduplication, save the event identifier in durable storage with a uniqueness constraint or equivalent atomic check, and enqueue or update work transactionally where possible. A process-local Set disappears on restart and is not shared between server instances. If the provider supplies no stable delivery ID, deduplicate using a documented job or resource identifier plus event type and state, or make the resulting state update idempotent. Avoid a scheme that drops distinct valid events merely because they refer to the same job.

Handle acknowledgment, retries, and duplicates

For Apify, a non-2xx response is treated as a delivery error and can trigger retries with exponential backoff. Its documentation specifies up to eleven retries, with the eleventh approximately 32 hours after the initial attempt, and a two-minute webhook request timeout. Those figures describe Apify’s documented behavior, not a general scraping API standard. Apify recommends responding immediately and moving time-consuming work to an internal message queue.

A quick acknowledgment should mean you have safely accepted responsibility for the event—not that every downstream task has finished. Return 2xx only after the event is validated and durably recorded or queued. If persistence fails, returning a failure response allows a provider retry where supported. If you acknowledge first and then lose the event, the provider may consider delivery complete even though your application never processed it.

Retries can produce duplicate notifications. Apify explicitly warns: “In rare cases, the webhook might be invoked more than once. Design your code to be idempotent to handle duplicate calls.” Make repeated handling harmless—for example, by upserting a job’s state rather than creating a new record each time, and by guarding one-time side effects with a durable idempotency record.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Retrieve results and represent failures

Keep notification handling and result processing as separate workflow stages. Store the provider event, the associated job or snapshot identifier, and the reported status. A worker can then obtain the result through the provider’s documented endpoint or storage mechanism. For Bright Data’s documented flow, track the snapshot state and retrieve results after it becomes ready; a failed state should follow an error path rather than the normal result-download path.

Define application states that make operational sense, such as queued, running, succeeded, and failed, and map provider states into them. Keep the provider’s raw status and event metadata where useful for diagnosis. Do not treat delivery success as scrape success: a successful HTTP acknowledgment means your receiver accepted the callback, not that the remote scrape succeeded or that the result has been downloaded.

Security and operations checklist

  • HTTPS: expose a TLS-protected receiver URL and do not send secrets over plain HTTP.
  • Authentication: use the provider’s documented token, secret, signature, or header mechanism. Reject requests that cannot be validated.
  • Minimal payload: include only the fields your receiver needs, using the provider’s supported payload template where available.
  • Fast, durable acceptance: validate and persist or enqueue, then respond; move downloads and expensive processing to workers.
  • Idempotent effects: deduplicate deliveries in durable storage and make state updates safe to repeat.
  • Observability: log a correlation identifier, job/snapshot ID, event type, receipt time, and processing outcome, while redacting secrets and sensitive scraped data.
  • Recovery: provide a way to inspect failed jobs and retry your own processing without relying solely on the provider to redeliver.
  • Provider review: verify current event names, retry schedule, timeout, response requirements, authentication options, and result endpoints before deployment. These details can change.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Troubleshoot common webhook problems

The provider reports delivery failures

Check that the URL is publicly reachable over HTTPS, routes to the expected handler, and does not require an interactive login. Inspect server and proxy logs for status code and response time. Confirm that the endpoint returns the status code the provider requires and that network rules allow inbound requests.

The provider keeps retrying an event

Inspect the actual response code and whether the handler exceeded the provider’s timeout. Validate JSON parsing and authentication, then make queue insertion fast. For Apify, non-2xx responses are errors and lead to retries; other providers may use different rules. Do not blindly return success to suppress retries if the event was not safely recorded.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The same job is processed twice

Assume duplicate delivery is possible and check whether deduplication is durable across restarts and server instances. Use a stable provider identifier if documented, or make the job-state update idempotent. Remember that an idempotency key used to create an Apify webhook prevents duplicate webhook creation, not duplicate callback execution.

The callback arrives but results are missing

Confirm that your code treats the notification as a lifecycle signal rather than the scrape output. Save the job or snapshot identifier and use the provider’s result endpoint or storage flow. For Bright Data’s documented flow, check progress and download when the snapshot is ready; handle failed separately.

The callback is rejected as unauthorized

Compare your validation logic with the provider’s documented authentication format and inspect headers without logging secret values. Apify recommends a secret token in the webhook URL and supports a headers template, but not every provider uses the same mechanism. Rotate any credential that has appeared in logs or a public repository.

Or skip the browser setup

If your scraping workflow also needs website screenshots, ScreenshotNeo is a separate screenshot API and MCP server—not a scraping-job webhook provider. One GET request returns an image or PDF, so you can avoid setting up and managing a browser capture worker for that part of the workflow. It is not a replacement for the webhook pattern above.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

For example, save a screenshot as WebP with cURL:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

See the ScreenshotNeo API documentation for request options. Cookie banners, popups, and chat widgets are removed before the shot; bot checks, blank pages, and failed loads are never billed. Its MCP server lets AI agents take screenshots. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. Sign up free for ScreenshotNeo.

Frequently Asked Questions

Can a webhook contain the complete scraped dataset?

It depends on the provider and its configured payload. The documented Bright Data flow uses a snapshot ID and a separate result download; do not assume the callback itself contains the dataset.

Does an HTTP 200 response mean the scrape succeeded?

No. It acknowledges receipt by your receiver. Check the provider’s job or snapshot status to learn whether the scrape itself succeeded.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.