Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
When a web-scraping API request fails, separate the problem into layers: request construction, authentication, HTTP response, transport, pagination, and parsing. Record the exact request and response first; then use the status code and structured error body to choose a fix. A timeout is not proof that the remote service returned no data, and a 200 response is not proof that your extraction is complete.
What to capture before changing code
Make one reproducible diagnostic record for the failing call. It should contain the method, endpoint, query parameters, body, relevant headers, authentication method, timeout, start time, elapsed time, status code, response headers and body, and redirect history. Also record the retry count and any request ID returned by the service.
Redact credentials and sensitive values before saving or sharing logs. Never record API keys, bearer tokens, passwords, or private cookies. If the payload may contain personal or customer data, store only a short redacted sample or a hash rather than the full response.
Do these 3 things before closing this tab:
1Repair Windows errors before they cause bigger problems2Fix the driver behind crashes, sound loss and screen glitches3Clear out junk files and repair common Windows errorsKeep the raw HTTP result distinct from your scraper’s interpretation of it. A useful record lets you tell whether the server rejected the request, the connection failed, the request timed out, or your parser mishandled a successful response.
#1 Best Overall
Check authentication and request construction
401: missing or invalid credentials
A 401 usually points to absent or invalid authentication. Confirm that the key belongs to the intended account or project, is active, and is being sent using the authentication scheme the API expects. Scrapy.io’s Platform API documentation requires an API key and recommends Bearer authentication; it specifically says not to put the key in a query parameter or browser-delivered code: Scrapy.io authentication documentation.
Check the final request as actually sent, not only the configuration you intended to use. Verify that the authorization header name and value are correct, that whitespace or quoting has not been introduced, and that a proxy or redirect is not stripping the header. Do not paste a live key into an issue, chat, or shared log.
400: validation or malformed input
A 400 can indicate an invalid body, missing required field, wrong field type, malformed JSON, or invalid pagination parameter. Read the error body before rewriting the scraper. Scrapy.io documents structured error types including validation_error; its API reference also describes invalid limits as validation errors: Scrapy.io error reference.
Quick wins for a faster PC:
Scan for outdated or missing drivers - takes under a minuteDriver Scan →Clear out junk files and repair common Windows errorsFree Scan →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Compare the actual serialized request with the endpoint’s expected schema. Common mistakes include sending JSON without the matching content type, encoding a nested value incorrectly, or passing a cursor or limit in the wrong location. Reduce the request to the smallest valid case, then add optional fields back one at a time.
Check the complete URL and headers
Log the endpoint and a redacted representation of the parameters. Confirm that reserved URL characters are encoded, that required headers are present, and that the method matches the endpoint. Avoid logging secrets simply to make the request easier to reproduce.
Read the HTTP status and structured error together
A status code narrows the cause, but the API’s error type and message often identify the specific fix. Scrapy.io documents these mappings for its Platform API: Scrapy.io status and error types.
| Status | Meaning to investigate | First useful check |
|---|---|---|
| 400 | validation_error |
Request body, required fields, parameter types, and pagination values. |
| 401 | unauthorized |
Credential presence, validity, scope, and authorization header. |
| 402 | insufficient_credits |
Account balance or plan allowance for the requested operation. |
| 403 | forbidden |
Whether the authenticated account is allowed to access this endpoint or resource. |
| 404 | not_found |
Endpoint path, resource identifier, and whether the resource exists. |
| 409 | conflict |
Whether the request conflicts with the current state or a prior operation. |
| 429 | rate_limit_exceeded |
Request frequency, concurrency, and any retry guidance in the response. |
| 500 | internal_error |
Whether the issue is transient; retain the request ID and report a minimal reproduction if it persists. |
These mappings describe Scrapy.io’s documented API, not a universal promise about every scraping provider. Other APIs may use different status codes or response formats, so consult the specific service’s error documentation.
Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minutePC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Separate transport failures from extraction failures
A scraper’s HTTP client can fail before it receives an HTTP response. In Python Requests, a Timeout is distinct from ConnectionError and HTTPError. A timeout means the client stopped waiting; it does not establish that the remote service completed no work or produced no data. Requests recommends setting explicit timeouts because calls without one may wait indefinitely: Requests timeout guidance.
Inspect the redirect history and final URL when your client follows redirects. A redirected request may reach a login page, a different endpoint, or a destination where authorization headers are treated differently. Check the content type and a small, redacted response sample before handing the body to a JSON or HTML parser.
Call raise_for_status() or use the equivalent in your HTTP library before interpreting the response as successful data. This keeps an error document from being misdiagnosed as malformed scraped content. Then validate the expected content type and schema explicitly.
Rank #3
Python diagnostic example
This example uses Requests to capture the essential diagnostic facts without printing the API key. Replace the endpoint and authentication header with the service’s documented values.
The Tool Desk
Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →import time
import requests
endpoint = "https://api.example.com/v1/scrape"
headers = {
"Authorization": "Bearer " + "YOUR_API_KEY",
"Accept": "application/json",
}
params = {"url": "https://example.org"}
started = time.monotonic()
try:
response = requests.get(
endpoint,
headers=headers,
params=params,
timeout=(5, 60), # connect timeout, read timeout in seconds
)
elapsed = time.monotonic() - started
print("status:", response.status_code)
print("elapsed_seconds:", round(elapsed, 3))
print("final_url:", response.url)
print("redirects:", [r.status_code for r in response.history])
print("request_id:", response.headers.get("X-Request-ID"))
print("content_type:", response.headers.get("Content-Type"))
print("body_sample:", response.text[:500]) # redact before retaining
response.raise_for_status()
data = response.json()
except requests.exceptions.Timeout as exc:
print("client timed out:", exc)
except requests.exceptions.ConnectionError as exc:
print("connection failed:", exc)
except requests.exceptions.HTTPError as exc:
print("HTTP error:", exc)
except requests.exceptions.RequestException as exc:
print("other request failure:", exc)
The example’s endpoint and request shape are illustrative, not a working Scrapy.io call. Use the actual endpoint, required method, schema, and authentication documented by your provider. Choose connect and read timeouts appropriate to the service’s expected work; do not treat one arbitrary timeout as suitable for every API.
Retry only requests that are safe to repeat
Do not retry every failure in a tight loop. Repeating a request can duplicate work, consume credits, or create duplicate jobs if the endpoint is not idempotent.
- Retry idempotent GET or HEAD requests when the failure is plausibly transient.
- Retry a POST only when the API documents idempotency support and the request includes an
Idempotency-Key. - For transient 429 or 5xx responses, use bounded exponential backoff, respect any retry timing supplied by the service, cap the number of attempts, and log each attempt.
- Do not repeatedly retry 400, 401, or 403 responses without changing the invalid request or authorization; those usually require a correction, not waiting.
For example, a backoff can increase the pause after each failed attempt while imposing both a maximum delay and a maximum attempt count. Add jitter if many workers may retry together. Preserve the original failure details so that a later successful retry does not hide a recurring rate-limit or server issue.
Validate pagination and completeness
A successful first page is not necessarily a complete scrape. Check the API’s pagination contract: whether it uses a cursor, page number, or offset; the permitted limit; the field indicating the next page; and the documented signal that the list has ended.
For each page, record the requested cursor or page, the echoed pagination values if present, and the number of items returned. Ensure the next cursor changes and stop when the API’s documented end condition is met. A repeated cursor can cause an infinite loop; a prematurely empty page can indicate a bad parameter, not necessarily that there are no more records.
Scrapy.io documents consistent pagination on list endpoints and validation errors for invalid limits: Scrapy.io pagination documentation. A 200 response should still be checked for expected data fields, plausible item counts, and a terminal pagination state.
Diagnose empty or incorrectly parsed results
When the response is 200 but your output is empty, inspect the raw response before changing selectors or parsing logic. Confirm that it is the expected format, that the response contains the target data, and that your parser reads the right field and type. APIs sometimes return a valid envelope with records nested under a key rather than at the top level.
- Check that the response body is not an error object wrapped in a successful HTTP response.
- Compare the returned schema with the parser’s assumptions, including capitalization, nesting, null values, and array types.
- Check whether the target page itself rendered content only after client-side loading, consent interaction, or another browser step; an HTTP fetch and a rendered-browser capture are different methods.
- For paginated results, verify that the parser processes every page rather than only the first response.
Keep transport and parsing tests separate: save a sanitized response fixture, then test your parser against it without making another network request. This makes it easier to distinguish an upstream change from a local code regression.
Recommended Free Tools
Use this troubleshooting sequence
- Reproduce one request with the same method, endpoint, parameters, body, headers, and authentication as the failing run.
- Set explicit connect and read timeouts, then capture elapsed time, status, response headers, a redacted body sample, final URL, and redirects.
- If an HTTP response exists, read the status and structured error body before parsing or retrying.
- Fix authentication or validation errors directly; do not mask them with repeated retries.
- For timeouts and connection failures, determine whether the client could connect, whether the server may still be working, and whether repeating the operation is safe.
- For 429 and transient server errors, retry only within a bounded backoff policy and retain a record of each attempt.
- On a 200, validate the schema, item count, and pagination end condition before declaring the scrape complete.
Choose an API with useful debugging visibility
For any scraping service, assess whether it exposes raw response headers and structured errors, supports safe authentication and secret handling, offers timeout and retry controls, preserves redirect details, documents pagination, and lets you log or export results safely. Also establish whether work is synchronous or runs as an asynchronous job, how job polling and dataset export work, and what the total request cost is.
Best Value
- Used Book in Good Condition
Scrapy.io documents synchronous calls, asynchronous runs, run polling, dataset export, and schedules for its Platform API: Scrapy.io Platform API documentation. Those capabilities may suit jobs that outlast a client-side timeout, but the right mode depends on the endpoint and workload; confirm the service’s current behavior and pricing before building around it.
Or skip the browser setup
If the problem is specifically capturing a rendered webpage rather than debugging a general scraping API, ScreenshotNeo returns a screenshot or PDF with one GET request. Its API accepts URL and capture options; parameter names used by other screenshot APIs also work, which can make migration easier. See the ScreenshotNeo API documentation.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
ScreenshotNeo accepts cookie and consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each of those steps can be turned off. Bot checks and CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and the response includes X-Page-Verdict and X-Billed headers. An MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients. The Free plan includes 1,000 screenshots per month without a card; paid plans start at $5 for 3,000 shots. Sign up for ScreenshotNeo free: 1,000 screenshots a month, no card required.
Frequently Asked Questions
Does a 200 status code guarantee that every result was scraped?
No. Validate the response schema, item count, and pagination end condition; HTTP success alone does not establish extraction completeness.
Should I put my scraping API key in the URL?
Follow the provider’s documented authentication method. Scrapy.io recommends Bearer authentication and says not to pass its key as a query parameter.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

