To scrape a known list of URLs, send the list to a provider’s batch endpoint. For a small job, a synchronous request may return the results in one response; for longer or larger jobs, submit an asynchronous batch, save its job or task identifiers, then collect results by polling or webhook. Treat each URL as its own outcome: a batch can contain both successes and failures.
Choose batch scraping, not crawling
A batch request is for URLs you already know. You provide an explicit list and ask the API to process each address. Crawling is different: it discovers or follows links from one or more starting pages. If the target set is fixed, a batch endpoint is usually the direct fit. Firecrawl’s documentation distinguishes its explicit-list batch operation from crawling (Firecrawl batch-scrape documentation).
First choose whether the client should wait for results. A synchronous batch is convenient when processing is short and the provider supports returning all results in the original request. An asynchronous batch is more suitable when the work may outlast an ordinary HTTP request or when you want to submit work and collect it later. Firecrawl documents both modes for an explicit URL list; ScraperAPI’s batch endpoint and Oxylabs’ Push-Pull workflow are asynchronous (Firecrawl; ScraperAPI; Oxylabs).
Plan the batch before sending it
Prepare inputs and shared options
Put the target addresses in an array and decide what you need back: rendered HTML, extracted fields, or another documented output format. Batch APIs often apply shared settings to every URL, but request formats and supported options differ. ScraperAPI, for example, documents a JSON request with an apiKey and a urls array sent to https://async.scraperapi.com/batchjobs. Do not transfer endpoint paths, authentication fields, or request bodies from one provider to another; use the chosen provider’s current documentation (ScraperAPI batch requests).
#1 Best Overall
Keep credentials outside source code, such as in environment variables or a secrets manager. Validate the list before submitting: remove accidental duplicates if they are not useful, reject malformed URLs, and retain the original input alongside each request. These steps make it easier to match returned tasks to the pages that generated them.
Set a safe batch size and concurrency
“Batch” does not mean unlimited parallel work. Providers set different maximums, submission rates, and concurrency controls, sometimes according to plan. In documentation accessed in 2026, ScraperAPI states a maximum of 50,000 URLs per batch job, while Oxylabs states up to 5,000 URL or query values per Push-Pull batch POST. Those are product-specific limits, not general API standards, and the documentation is undated, so confirm current account limits before relying on them (ScraperAPI; Oxylabs).
Firecrawl documents a per-job maxConcurrency setting; its example of 50 means 50 simultaneous scrapes and is an example, not a universal recommendation. Scrape.do lists asynchronous concurrency by plan, and Oxylabs says submission rates depend on subscription plan. Check your current plan limits and begin with a size and concurrency level that your application can track and recover safely (Firecrawl; Scrape.do; Oxylabs).
Submit an asynchronous batch and collect its results
- Submit the URL list. Send the provider-specific request body and authentication fields. Record the submission time, batch identifier if provided, and the original URL list.
- Save each returned identifier. Some APIs return one task or job record for every input URL, not only a single batch ID. Store the input URL with its returned task ID or status URL so you can reconcile results later.
- Wait for completion. Poll the documented status endpoint at a controlled interval, or configure a webhook or callback if the provider supports one. Avoid repeatedly requesting status at high frequency.
- Fetch and persist output. Retrieve completed results and store any content or extracted fields the application needs. Do not assume the provider’s API is a permanent archive.
- Reconcile every input. Mark each URL as succeeded, failed, pending, or expired, and retain useful error details and attempt counts.
For Scrape.do, the documented flow is create a job, check job status, then fetch task results by job and task ID. The documentation recommends exponential backoff for status checks, identifies 429 as a rate-limit response, and recommends webhooks for production use to reduce polling. Firecrawl supports tracking an asynchronous batch by ID with status polling or webhooks; its webhook documentation describes per-page notifications and started, completed, and failed events. It documents HMAC-SHA256 signature verification using the X-Firecrawl-Signature header (Scrape.do async API; Firecrawl batch-scrape documentation).
Use backoff for polling
If you poll, wait longer between successive checks rather than issuing a tight loop. Exponential backoff increases the delay after each unsuccessful check; cap the delay at a limit suitable for your application and stop when the job completes, fails, or reaches your own deadline. Honor any provider-supplied retry guidance. If the API returns 429, reduce request frequency and follow the provider’s documented rate-limit instructions instead of immediately retrying at the same pace.
Make webhooks safe to process
A webhook can avoid repeated status requests, but your receiver still needs to handle delayed or repeated deliveries. Verify signatures when available, acknowledge notifications promptly, and make processing idempotent so that a duplicate event does not create duplicate records. Store the event and update the task identified by the provider; do not treat receipt of one event as proof that every URL in the batch succeeded.
Handle partial failures and retries per URL
Do not assume a batch is atomic. A provider may report progress for the overall job while individual pages fail, and task-level status or error details are important. Firecrawl documents an operation for inspecting failed URLs; Scrape.do advises checking each task’s status. Keep a record for each item with at least the submitted URL, provider task ID, status, error or response information, attempt count, and final result (Firecrawl; Scrape.do).
When retrying, select failed or otherwise retryable items rather than resubmitting successful ones by default. This selective-retry approach follows from APIs exposing per-URL statuses and errors; retry conditions and any idempotency support depend on the provider. Avoid endless retry loops: cap attempts, record the final failure, and distinguish transient problems from inputs or pages that need a different treatment.
The Tool Desk
Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Retrieve results before they expire
Retention is provider-specific. Scrape.do warns that task results are temporary and should be retrieved before the task’s ExpiresAt time. Firecrawl says batch results remain available through its API for 24 hours after completion, after which activity logs remain available. These are different retention policies, not a general guarantee for scraping APIs (Scrape.do; Firecrawl).
Fetch and persist the data your application needs while it is available. If your system requires a durable record, save the result, relevant metadata, and task outcome in storage you control; do not rely on an API’s temporary result endpoint as your archive.
Compare providers by workflow, not just batch size
The documented products expose different workflows. Limits below are vendor statements from undated documentation accessed in 2026 where noted; they can change. They do not establish a performance or reliability ranking.
| Provider | Documented batch workflow | Useful distinction | Limit or retention stated in the cited documentation |
|---|---|---|---|
| Firecrawl | Explicit URL list; synchronous or asynchronous batch | Same structured-extraction schema can be applied across URLs; configurable per-job concurrency; status polling or webhooks | Results available by API for 24 hours after completion, per Firecrawl documentation accessed in 2026. Documentation |
| ScraperAPI | Asynchronous POST to https://async.scraperapi.com/batchjobs; response includes an individual job record per URL |
Documented JSON body uses apiKey and urls |
Up to 50,000 URLs per batch job, per undated documentation accessed in 2026. Documentation |
| Oxylabs Web Scraper API | Push-Pull asynchronous method for larger workloads | Results can be delivered by callback or written to cloud storage | Up to 5,000 URL or query values per batch POST; Push-Pull results available for at least 24 hours; submission limits depend on plan, per undated documentation accessed in 2026. Documentation |
| Scrape.do | Create job, check job, fetch task results | Task-level status and result retrieval; documentation recommends backoff and production webhooks | Its accessed documentation lists separate async concurrency limits: Free 2, Hobby 3, Pro 15, Business 30, Advanced 60, and Custom/Enterprise 30% of plan limit. Vendor-reported and volatile. Documentation |
Before committing to a provider, check explicit URL-list support, synchronous versus asynchronous behavior, per-URL errors, concurrency controls, webhook or callback support, output format, retention, and account-specific submission rates. A high maximum batch count alone does not tell you whether the API returns the format or task-level visibility your application needs.
Free tools Windows power users keep installed
One-click scans. No signup required.
Reliability, performance, and cost considerations
Keep the system observable
Track submitted, pending, successful, failed, retried, and expired counts, along with job duration and request errors. Preserve enough identifiers to find a result without reconstructing a batch from logs. For asynchronous jobs, keep submission and collection as separate operations so a worker restart does not lose the ability to resume fetching results.
Control workload rather than assuming a batch is faster
Batching can simplify submission and result management, but the cited documentation does not provide independent speed benchmarks or a basis for claiming that one provider is faster. Throughput depends on provider limits, target response behavior, selected options, and available concurrency. Avoid sending more simultaneous jobs than your account permits; a large submission can still wait in a queue or encounter per-task failures.
Budget around actual provider terms
Pricing and billing units are not established by the cited batch documentation, so check the current plan and how it counts successful, failed, and retried requests before estimating cost. Do not assume that a batch is priced as one request or that every URL is billed identically. Also account for the engineering cost of result storage and retries, especially when retention windows are short.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Troubleshooting common batch problems
- The request is rejected immediately: Check the provider’s required HTTP method, endpoint, authentication field, JSON shape, URL encoding, and maximum batch size. Request formats are not interchangeable.
- The job was accepted but results are missing: Confirm you are checking the correct batch and task identifiers and using the provider’s documented result endpoint. Some APIs return a separate record per URL.
- Status checks return HTTP 429: You are exceeding a rate limit or polling too frequently. Back off, reduce status-request frequency, and use a supported webhook or callback where appropriate.
- Some URLs failed while others succeeded: Inspect task-level statuses and error details, then retry only eligible failures. Do not interpret overall batch progress as proof of individual success.
- A result can no longer be fetched: Check whether the task expired or the provider’s retention period ended. Retrieve results before the stated expiry and persist them in your own storage.
- Submissions are queued or slower than expected: Verify plan-specific concurrency and submission limits, then reduce simultaneous jobs or adjust batch size. The documentation cited here does not establish comparative provider performance.
Or skip the browser setup
If your goal is to capture website pages as screenshots or PDFs rather than extract structured page data, ScreenshotNeo offers a one-request screenshot API. It accepts a URL and can return PNG, JPEG, WebP, or PDF; use a screenshot API for visual capture, not as a substitute for a scraper that needs page content or structured fields.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
For a single capture, the cURL request is:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
See the ScreenshotNeo API documentation for request options. ScreenshotNeo accepts cookie and consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; those steps can each be turned off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and responses identify the page verdict and billing status in headers. Its MCP server provides take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients. The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 screenshots, and every feature is available on every plan.
Sign up for ScreenshotNeo’s free plan: 1,000 screenshots a month, no card required.
Frequently Asked Questions
Should I use a batch endpoint or a crawl endpoint?
Use a batch endpoint when you already have the URLs. Use a crawl workflow when the task is to discover or traverse pages from starting points.
Does a batch request guarantee every URL succeeds?
No. Track individual task outcomes and errors; one batch may contain both successful and failed URLs.
How often should I poll an asynchronous job?
Follow the provider’s guidance and use increasing delays rather than a rapid loop. Prefer a supported webhook or callback when it fits your application.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




