A 403 usually means the site or its security layer is refusing access; a 429 means you have sent requests too quickly. Diagnose which response you received before changing your scraper: honor rate limits and retry a 429 cautiously, but treat a 403 as an authorization or policy issue—not as a signal to disguise the request. If you are not authorized to access the content, stop and ask the site operator for an approved route.
What a 403 or 429 tells you
Both are HTTP client errors, but they call for different responses. RFC 6585 defines 429 as a rate-limit response: the client has sent too many requests in a period. The response may include a Retry-After header telling the client when it may try again. A 403 is an access-denied decision. Possible reasons include missing permission, an IP or country restriction, a firewall rule, or a site policy. A Cloudflare-protected site can return access-denied responses for several of these reasons; a challenge may also be produced by WAF rules, Bot Management, Bot Fight Mode, Turnstile, HTTP DDoS protection, or Under Attack Mode.
Do not infer the cause from the status code alone. A site can customize its error pages, and an intermediary such as a CDN or WAF may generate the response rather than the origin server. Inspect the headers and a limited sample of the response body, then compare with a request you are authorized to make.
| Response | Likely interpretation | Safe next action |
|---|---|---|
| 429 Too Many Requests | Rate limit or quota reached | Check Retry-After and rate-limit headers, reduce load, and retry within a bounded budget. |
| 403 Forbidden | Permission, policy, network, or security rule denied access | Verify authorization and request requirements; use an approved API or contact the operator. |
| Challenge or interstitial page | A security check is asking for an allowed browser interaction or denying automation | Do not automate a bypass. Use an authorized browser flow if the site permits it, or request access. |
Diagnose the response before changing your scraper
Capture enough evidence to distinguish a temporary quota from a policy denial. Keep credentials and personal data out of logs; redact secrets and limit stored body samples.
Free tools Windows power users keep installed
One-click scans. No signup required.
#1 Best Overall
- Record the request context. Log the URL (redact sensitive query values), method, timestamp, status, redirect chain, request identity, concurrency, and a bounded response-body sample. Record relevant response headers, but never log authorization tokens or session cookies in plaintext.
- Look for rate-limit signals. Check
Retry-After,Ratelimit, andRatelimit-Policy. Cloudflare documentsretry-afteras the seconds until more capacity is available and describes quota headers. Header presence and spelling can vary; absence does not make an unlimited retry safe. - Check for challenge evidence. Look for an interstitial HTML page, challenge-related cookies, vendor headers, or a request identifier such as a Cloudflare Ray ID. Preserve the identifier for the site operator. Do not assume the response body is the requested page just because the HTTP request completed.
- Compare against an authorized request. Confirm the endpoint, method, credentials, required headers, cookies, TLS behavior, and network origin. A manually authorized request and a scraper request can differ in authentication or endpoint usage even when they use the same URL.
- Classify before retrying. A 429 may be temporary; a 403 or challenge is generally not fixed by repeating the same request. Treat 5xx origin failures separately, and do not retry every 4xx automatically.
Fix 429 Too Many Requests without making the block worse
The first response is to slow down, not to fan out across more connections. RFC 6585 says a 429 response indicates that the user has sent too many requests in a given amount of time; its response may explain the condition and may include Retry-After. Treat that header as a minimum wait, parse either its delay-in-seconds form or its HTTP-date form, and do not retry before it expires.
Reduce demand at the source
- Lower per-host concurrency and use a per-host token bucket or other rate limiter so workers share one request budget.
- Cache responses where freshness requirements allow, deduplicate identical URLs, and avoid fetching pages already collected.
- Spread scheduled work over a longer window instead of launching a large burst. If the site publishes quotas, configure your client below them and monitor actual response headers.
- Stop when 429s continue without recovery, or when the account or IP is explicitly blocked. Ask the operator about the limit rather than trying to route around it.
Use bounded retries with backoff and jitter
If Retry-After is absent, use exponential backoff with random jitter, a maximum delay, and a small retry budget. Jitter prevents many workers from waking and retrying simultaneously. Retry only idempotent operations (for example, a documented read-only GET) unless the API owner explicitly documents safe retries for another method. A retry budget should count attempts, not run indefinitely.
Rank #2
Cloudflare’s documented API limits are 1,200 requests per five minutes per user or account token and 200 requests per second per IP, according to Cloudflare’s 2026 documentation. These figures apply to Cloudflare APIs, not to arbitrary websites protected by Cloudflare and not to a target site’s scraping allowance.
Fix 403 Forbidden through permission and configuration
Think of 403 as a decision to deny this request, not a temporary overload. First confirm you have permission to collect the content and that your method complies with the site’s terms and robots guidance. Robots rules are not a grant of access; obtain the site owner’s approval where required.
Rank #3
- Used Book in Good Condition
Check credentials and the intended access route
- Use the site’s official API, export, feed, or licensed data channel when one exists. Confirm the endpoint and method in its documentation.
- Verify that the account is authorized for the resource. Refresh expired credentials and include documented session or CSRF state only when the site’s approved flow requires it.
- Check whether access is restricted by IP, country, account role, or network. If you own the site, review the relevant firewall and WAF logs for the matched rule.
- If a browser challenge is intended for interactive visitors, use a normal browser flow only where authorized. Otherwise request an allowlist entry, API credential, or other approved access from the operator.
Changing User-Agent strings or rotating proxies does not establish permission and is not a reliable or appropriate fix for an intentional block. It can obscure the diagnostic trail or violate site rules. Do not attempt to evade CAPTCHA, bot checks, geographic controls, account restrictions, or other access controls.
If you operate the protected site
Review the firewall or WAF event and identify the rule, route, and traffic characteristic that matched. For rate rules, tune the expression, counting characteristics, period, requests-per-period, and mitigation duration to match the intended policy. Cloudflare notes that counters can take a few seconds to update, so thresholds are approximate at enforcement time. Test a narrow adjustment and monitor both legitimate traffic and abusive requests rather than broadly disabling protection.
Rank #4
- Used Book in Good Condition
Build a response classifier and bounded retry loop
Keep response handling explicit. A useful classifier has states such as success, rate_limited, access_denied, challenge, auth_required, and origin_error. Send 429 responses to a rate-limited path; stop and seek authorization or operator support for access-denied and challenge responses. Preserve request IDs such as Cloudflare Ray IDs. Cloudflare structured errors can expose fields including retryable, retry_after, owner_action_required, and error_category; use those fields when the relevant documented API returns them.
This Python example retries a read-only GET only for 429, honors a valid Retry-After value, and stops after a fixed number of attempts. It deliberately does not retry a 403 or attempt to bypass a challenge. Install the dependency with python -m pip install requests, then set TARGET_URL to a URL you are authorized to request.
Best Value
import email.utils
import random
import time
from datetime import datetime, timezone
import requests
TARGET_URL = "https://example.com/permitted-resource"
MAX_ATTEMPTS = 4 # Includes the first request
MAX_DELAY_SECONDS = 120
def retry_after_seconds(value):
if not value:
return None
try:
return max(0.0, float(value))
except ValueError:
try:
when = email.utils.parsedate_to_datetime(value)
if when.tzinfo is None:
when = when.replace(tzinfo=timezone.utc)
return max(0.0, (when - datetime.now(timezone.utc)).total_seconds())
except (TypeError, ValueError, OverflowError):
return None
with requests.Session() as session:
for attempt in range(MAX_ATTEMPTS):
response = session.get(TARGET_URL, timeout=(5, 30))
print("status:", response.status_code)
print("request id:", response.headers.get("cf-ray", "not provided"))
if response.status_code != 429:
if response.status_code == 403:
raise SystemExit("403 access denied: verify permission or contact the site operator")
if response.status_code in (401, 407):
raise SystemExit("Authentication or proxy authentication is required")
if 500 <= response.status_code <= 599:
raise SystemExit("Origin/server error: handle under the service's documented policy")
response.raise_for_status()
print("success; bytes:", len(response.content))
break
if attempt == MAX_ATTEMPTS - 1:
raise SystemExit("429 persisted; retry budget exhausted")
advertised = retry_after_seconds(response.headers.get("Retry-After"))
if advertised is not None:
delay = advertised # Never retry before the server's requested wait.
else:
base = min(MAX_DELAY_SECONDS, 2 ** attempt)
delay = random.uniform(base / 2, base)
if delay > MAX_DELAY_SECONDS:
raise SystemExit("Retry-After exceeds local wait limit; stop and reschedule later")
print("rate limited; waiting seconds:", round(delay, 2))
time.sleep(delay)
For a production client, add structured logging, a per-host limiter shared by all workers, a total job deadline, and a policy for rescheduling work when the retry budget is exhausted. Do not print or persist full response bodies by default; error pages may contain sensitive information.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Choose the access method that fits the job
Compare approaches on permission, freshness, request volume, latency, implementation effort, stability under WAF changes, observability, cost, and contractual fit. An official API or licensed feed is generally the most stable option. A slower, authorized crawl may be suitable where no API exists. An interactive browser is appropriate only if the site permits that access path. If the only goal is a visual record of a page you are allowed to view, a screenshot is not a substitute for permission to scrape or extract its data.
Or skip the browser setup
If your task is to capture a page you are authorized to view—not to bypass a block or harvest its contents—ScreenshotNeo provides a website screenshot API and MCP server. One GET request returns a PNG, JPEG, WebP, or PDF. See the ScreenshotNeo API documentation for parameters and formats.
Quick Recap
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
ScreenshotNeo accepts cookie or consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Only clean shots are billed: bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and responses report the page verdict and billing status in headers. Its MCP server gives AI agents tools named take_screenshot, get_page_info, and capture_pdf. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 shots. It does not grant access to a site that has denied your request. See ScreenshotNeo, or sign up free for 1,000 screenshots a month with no card.
Recommended Free Tools
Troubleshoot common failure patterns
- 429 repeats immediately: The client may be ignoring
Retry-After, sharing no limit across workers, or retrying too many URLs at once. Confirm the parsed delay, lower concurrency, and use one host-level budget. - 429 has no retry header: The server has not supplied a wait value. Use capped exponential backoff with jitter and a small retry budget; stop if the service does not recover.
- 403 persists after credentials are refreshed: The issue may be authorization scope, IP/country policy, a WAF rule, or an endpoint/method mismatch. Check the response ID and ask the operator which approved route applies.
- The response is HTML instead of expected data: Inspect a bounded sample and headers for a challenge or interstitial. Do not parse the challenge as content or automate a workaround; request an allowed API or browser path.
- Requests work manually but fail in the scraper: Compare endpoint, method, authentication, required cookies or headers, TLS configuration, and source network. Avoid copying browser session secrets into an unattended job unless the site’s documented terms permit that use.
- Cloudflare API limits are mistaken for site limits: The 1,200-per-five-minutes-per-token and 200-per-second-per-IP figures cited above describe Cloudflare API limits, not the limit for a third-party website behind Cloudflare. Obtain the target’s own policy.
Decide whether to retry, pause, or stop
| Condition | Decision |
|---|---|
429 with a usable Retry-After |
Pause at least as long as instructed, then retry only within the job’s bounded budget. |
| 429 without recovery, or retries exhausted | Stop the job and reschedule later or contact the service owner. |
| 403, explicit block, or challenge | Stop automatic retries; verify access rights and use an approved access path. |
| Successful response from an official API or licensed feed | Follow its documented quotas, authentication rules, and permitted-use terms. |
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

