What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Short answer: an HTTP status code tells you what happened at the layer that returned it—not necessarily whether the target page was scraped correctly. A scraping API may expose its own result, a proxy response, or the target website’s response. Treat 200 as “the responding layer completed an HTTP request,” then validate the body, content type and expected fields before accepting the data.
The practical workflow is to record the code and response, identify the responding layer, inspect the payload, and apply the provider’s documented retry, concurrency and billing rules. The standards define status semantics; each scraping service decides which layer it surfaces and how it handles CAPTCHAs, retries and charges.
Why one scraping request can have several status layers
A normal browser talks directly to a website. A scraping API adds at least one intermediary: your client calls the API, the API calls the target (often through a proxy), and the target returns a response. The service can then return the target body, its own error, or a normalized result.
Consequently, a 403 might mean the target rejected the proxy, while another provider could use 403 for an API-level permission problem. A 407 specifically concerns proxy authentication, whereas 401 concerns credentials for the target resource under RFC 9110. Read the provider’s error format, headers and documentation before changing code. The normative definitions are in RFC 9110 and the MDN status reference.
#1 Best Overall
- Funny design. funny HTTP status code featuring a green thumbs up and the words "200 OK". A fun tee for any web developer or web programmer with a sense of humor
- Hardcover journal with 240 line-ruled pages (120 sheets)
- Built-in elastic closure and ribbon bookmark
- Includes an expandable inner storage pocket and a pen holder
Questions to answer for every response
- Which component generated the status: your API endpoint, its proxy, or the target site?
- Did the client follow redirects, and what was the final response?
- Does the body contain the expected page and fields, or a login form, CAPTCHA, block page or empty document?
- Does this provider retry, bill, or classify this outcome in a special way?
The five HTTP status classes
| Class | Meaning | Scraping interpretation |
|---|---|---|
| 1xx | Informational | Interim protocol messages; most scraping clients do not expose them as the final result. |
| 2xx | Successful | The responding layer completed the request. Validate the body; success does not guarantee useful data. |
| 3xx | Redirection | Inspect redirect-following settings and the final URL and status. |
| 4xx | Client error | Request, authentication, permission, rate or resource problems; the exact layer matters. |
| 5xx | Server error | Failure at the API, proxy or target server; identify which one before retrying. |
Common codes and the action that makes sense
200 OK: transport success, not data success
200 means the responding server successfully handled the HTTP request. It does not prove that the intended article, product record or JSON object is present. A target can return a CAPTCHA, login page, “access denied” document or empty shell with status 200. ScraperAPI documents CAPTCHA detection in successful responses as a provider-specific workflow, not a universal HTTP rule (ScraperAPI status documentation).
Check the final URL, Content-Type, minimum body length and required selectors or JSON keys. Reject a response that lacks those invariants, even when its status is 200.
301, 302 and other 3xx redirects
Redirects indicate that the requested location points elsewhere. Whether the client follows them, preserves authorization, changes methods or exposes the intermediate response depends on the code and client. Enable the provider’s documented redirect behavior, then validate the final URL and body. A redirect to a sign-in or consent page is not a successful scrape of the original resource.
400 Bad Request
This commonly indicates a malformed or unsupported request: an invalid URL, missing required parameter or unacceptable option. ScraperAPI labels its 400 responses as malformed requests and advises checking the URL. Do not blindly retry an unchanged request; log the provider’s error payload, correct the input and try once.
The Tool Desk
Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →401 Unauthorized
RFC 9110 defines 401 as a request lacking valid authentication credentials for the target resource. A scraping service may instead use 401 for its own API-key failure; ScraperAPI lists an invalid API key as one cause. Determine whether the credentials belong to your scraping account, the target website, or a proxy before rotating keys.
403 Forbidden
The responding layer refuses access. Supplying credentials is not automatically a fix: 401 and 403 have different semantics. A target may block an IP, user agent, geography or automation fingerprint. ScraperAPI notes that protected domains may require a premium request option, but that advice applies to that provider, not every API. Check permission settings, proxy options and the target’s terms, and inspect the response body for a block explanation.
404 Not Found
The requested resource was not found at the responding layer. Verify URL encoding, trailing slashes, locale paths, redirects and whether the record was removed. An API can also return 404 for a missing API resource rather than a missing target page. ScraperAPI counts 404 among successful requests for billing purposes, illustrating why “successful request” in a billing policy does not mean “content found.”
407 Proxy Authentication Required
407 is proxy authentication, distinct from target-resource authentication (401). It usually points to proxy credentials or configuration between the scraping service and target. Confirm the proxy username, password, host and port in the provider’s settings; changing target-site login credentials will not solve a proxy challenge.
Free tools Windows power users keep installed
One-click scans. No signup required.
429 Too Many Requests
429 signals excessive request volume. The API may apply it to account rate limits, simultaneous-request limits or a target response passed through by the service. Reduce concurrency, add exponential backoff with jitter, honor Retry-After when supplied, and review your plan’s limits. ScraperAPI specifically documents excessive simultaneous requests and recommends checking plan concurrency. Retrying immediately at the same rate usually extends the outage.
5xx server errors
A 5xx code can originate at the target, proxy or scraping API. Capture the provider request ID, timestamp and target URL, then consult status pages or support if the provider is failing. Retry transient failures according to documented policy, preferably with bounded exponential backoff; do not retry a deterministic 501/505-style incompatibility forever. ScraperAPI says requests that fail after 70 seconds of retrying are not charged—an implementation and billing rule that must not be generalized to other services.
Rank #3
- Quick reference Statistics chart
- This 8.5" x 11" 4-page laminated Guide provides an easy to follow summary of all basic principles that are the foundation to Statistics and Probabilities
- Detailed descriptions and examples of theory
- Using a combination of charts and sample equations, the key concepts are developed and the essential Statistics theories are outlined.
- Easy-to-read to promoted memory retention. Great quick reference aid.
A diagnostic sequence you can automate
- Log the complete event. Store status, final URL, selected headers, response size, content type, provider request ID, target URL and UTC timestamp. Redact API keys, cookies and personal data.
- Identify the layer. Look for provider-specific error JSON, headers and documentation. A target status passed through by the API has different remediation from an API authentication error.
- Validate successful bodies. For 2xx responses, check content type, required selectors or JSON keys, expected title/URL and block-page signatures such as CAPTCHA or login text.
- Choose a targeted fix. Correct malformed URLs for 400; fix the relevant credentials for 401/407; investigate access policy for 403; verify resource existence for 404; reduce rate or concurrency for 429.
- Retry selectively. Use the service’s documented retry rules. Never repeat unchanged malformed or unauthorized requests, and cap attempts for 5xx responses.
- Record billing outcome. Providers differ on whether 404, CAPTCHA, timeout and exhausted retries consume credits. Treat billing headers or the provider dashboard as authoritative.
Minimal client-side validation examples
Python
import requests, time, random
url = "https://example.com/data"
for attempt in range(4):
r = requests.get("https://api.example.test/fetch",
params={"url": url, "api_key": "YOUR_KEY"},
timeout=90)
print(r.status_code, r.url, r.headers.get("content-type"))
if r.status_code == 200:
text = r.text.lower()
if "captcha" not in text and "required-field" in text:
data = r.json()
break
raise ValueError("HTTP 200 but expected content is missing")
if r.status_code in (429, 500, 502, 503, 504):
time.sleep((2 ** attempt) + random.random())
continue
r.raise_for_status()
else:
raise RuntimeError("Retries exhausted")
cURL
curl --fail-with-body --max-time 90
-G "https://api.example.test/fetch"
--data-urlencode "api_key=YOUR_KEY"
--data-urlencode "url=https://example.com/data"
-D response-headers.txt -o response-body.html
Use --fail-with-body so non-2xx responses fail the command while preserving the provider’s diagnostic body. Do not treat a successful exit code as proof that the page contains your data.
Node.js
const q = new URLSearchParams({ api_key: 'YOUR_KEY', url: 'https://example.com/data' });
const res = await fetch(`https://api.example.test/fetch?${q}`, { signal: AbortSignal.timeout(90000) });
const body = await res.text();
if (res.status === 200 && body.includes('required-field') && !body.toLowerCase().includes('captcha')) {
console.log('validated response');
} else if ([429,500,502,503,504].includes(res.status)) {
throw new Error(`transient failure: ${res.status}`);
} else {
throw new Error(`scrape failed: ${res.status} ${body.slice(0,200)}`);
}
Retries, concurrency, performance and cost
Backoff without a thundering herd
Use exponential delays with random jitter, honor Retry-After, and bound total attempts and wall-clock time. A queue that limits simultaneous requests is safer than launching thousands of promises and hoping the provider absorbs the burst.
Separate transport metrics from extraction metrics
Track HTTP status, body-validation failures, field completeness, latency and bytes independently. A rising 200 rate alongside falling field completeness often means the target changed its markup or began serving a challenge page.
Billing is provider-specific
Do not infer charges from status classes. Read current plan documentation for 200, 404, CAPTCHA, timeout, cache-hit and exhausted-retry treatment. ScraperAPI’s documented policies are examples, not industry defaults.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Do not confuse scraping APIs with search crawlers
Google says its crawlers temporarily slow after 429 and 5xx responses, and that a 2xx response does not guarantee indexing (Google’s HTTP status guidance). Those rules describe Google crawling and indexing, not a universal behavior for scraping clients or APIs.
How to compare scraping APIs
| Evaluation question | Why it matters |
|---|---|
| Which response layer is surfaced? | Determines whether you fix your request, proxy credentials or target access. |
| Is the body checked for CAPTCHA or block pages? | Prevents false “success” from 200 responses. |
| What are retry and concurrency rules? | Controls latency, 429 frequency and duplicate work. |
| How are 200, 404 and failures billed? | Changes the real cost of unreliable targets. |
| Are error bodies and request IDs clear? | Shortens diagnosis and support cycles. |
Or skip the browser setup
For screenshots rather than raw HTML extraction, ScreenshotNeo is the first service to try: it removes cookie banners, newsletter popups and chat widgets before capture, bills only clean shots, and exposes verdict headers so you can distinguish bot checks, blank pages, timeouts, failed loads and cache hits. It also provides an MCP server for AI agents and a free allowance of 1,000 screenshots per month without a card; paid plans start at $5 for 3,000 shots.
Recommended Free Tools
One GET request returns PNG, JPEG, WebP or PDF:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
In production, inspect X-Page-Verdict and X-Billed headers instead of assuming every 200 is a usable capture. The service supports full-page and element captures, device and viewport settings, retina scale, PDF options, custom CSS/JavaScript, waits, blocking rules, headers, cookies, caching, signed links, asynchronous webhooks and bulk capture. See the ScreenshotNeo documentation for parameter names and response details.
Create a free ScreenshotNeo account to get 1,000 screenshots a month with no card.
Frequently Asked Questions
Can a scraping API return 200 when the target blocked me?
Yes. A target or intermediary can return a CAPTCHA, login page or block document with HTTP 200. Validate the body and expected fields; status alone is insufficient.
Should every 429 or 5xx response be retried?
No. Retry only transient conditions under the provider’s documented policy, with bounded exponential backoff and reduced concurrency. Do not retry unchanged malformed or unauthorized requests.
Do these 3 things before closing this tab:
1Repair Windows errors before they cause bigger problems2Fix the driver behind crashes, sound loss and screen glitches3Clear out junk files and repair common Windows errorsWho is responsible for a 403 returned by my scraping API?
You must identify the responding layer from headers, error format and documentation. It may be the target site, proxy or API itself, and each requires a different fix.
Does a 404 always mean the target URL is gone?
No. It can represent a missing target resource or a missing resource at the scraping API endpoint. Check the final URL and provider error payload.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




