To tell whether a website is blocking your scraper, look at several signals together: the HTTP status and headers, the response body, how the result compares with an authorized control request, whether the behavior repeats, and—if you operate the site—what the server or security logs show. A single failed request or status code does not prove a deliberate block.
This guide is for diagnosing permitted scraping and site-owner security events, not bypassing a challenge or access restriction. If a site presents a challenge or explicitly denies access, follow its published rules.
What evidence points to a block?
A likely block is a response that differs from the expected page in a way that is repeatable and consistent with an access-control action. A challenge page, interstitial, or substitute response is more informative than a status alone, especially when an authorized control request to the same URL returns the expected content.
Keep the diagnosis provisional until you have enough context. Temporary site problems, client configuration, proxies, gateways, and security rules can all affect what arrives. The same outward symptom can have different causes.
#1 Best Overall
- Status: Record it as part of the response, but do not infer a specific cause from it alone.
- Headers: Save them as context. The sources reviewed do not establish that any particular header always identifies a block.
- Body: Check whether the content is the intended page, a challenge, an interstitial, an error, or an empty substitute.
- Repeatability: A consistent difference across permitted attempts strengthens the diagnosis, but does not by itself prove intent.
- Corroboration: For site operators, a matching WAF event, rule action, or bot-analytics signal is stronger evidence than client output alone.
A repeatable diagnostic workflow
- Record the full response. Save the requested URL, final URL after redirects, method, status, response headers, time, and response body or a safe fingerprint of it. Keep the relevant request configuration, such as the User-Agent you intended to send. Avoid storing secrets or personal data in diagnostic logs.
- Inspect the returned content. Do not stop at an apparently successful HTTP response. Examine the actual HTML or rendered result for challenge language, interstitial markup, or content unrelated to the expected page. A crawler-measurement study’s abstract describes checking both status and HTML when identifying block or challenge pages: study abstract.
- Compare with an authorized control. Where the site permits it, request the same URL and method through an ordinary, permitted control client. Compare the content and relevant response details with the scraper result. A difference confined to one client pattern is useful evidence; it is not permission to evade a restriction.
- Check whether the pattern repeats. Note isolated versus repeated failures, timing, and changes in request behavior. Do not increase request volume to “test” a limit. Site-specific rate limits and detection signals are not universal thresholds.
- Verify client and intermediary details. Confirm the request metadata your client actually sent and consider whether a proxy or gateway changed it. Cloudflare notes that a missing or empty User-Agent can receive its lowest bot score, and that a corporate proxy stripping the header may explain unexpected scoring. This illustrates one possible confounder, not a rule for every website: Cloudflare bot score documentation.
- Correlate with site telemetry if you own the site. Match the request time and endpoint against server logs, WAF events, bot analytics, and the rule or challenge action. Cloudflare recommends using Bot Analytics before applying bot rules; score availability varies by plan: Cloudflare Bot Analytics.
- Stop or change course within the site’s rules. A challenge or explicit restriction is a signal to respect the site’s access controls and terms. Contact the site owner or use an authorized API if you need access.
How to compare the responses
Use the same dimensions for each attempt so that differences are meaningful. A compact incident record can prevent a generic network failure from being mistaken for a security decision.
| What to compare | What to record or ask | How to interpret it |
|---|---|---|
| Status and headers | Status, response headers, final URL, and timestamp | Useful context about what returned; neither status nor a particular header proves a block by itself. |
| Expected versus received body | Expected page markers versus received HTML, text, or a safe content fingerprint | A challenge or unrelated substitute body is strong practical evidence when checked against a valid control. |
| Control versus scraper | Same URL and method under authorized, comparable conditions | A consistent client-specific difference narrows the diagnosis, but may still have more than one cause. |
| One-time versus repeated | Whether the symptom recurs, and whether request behavior changed | Repeatability strengthens a pattern; temporary errors can still recur. |
| Client evidence versus site telemetry | Client logs alongside server, WAF, or bot-analytics events | Operator-side records can identify a rule or challenge action and correlate it with the request. |
What site operators should check
If you operate the website, start with the endpoint and event rather than applying a broad rule based on one suspicious request. Cloudflare documents several bot-detection approaches, including heuristics, JavaScript detections, machine learning, and behavioral methods; availability depends on plan. Its current bot-detection-engine documentation says the legacy Anomaly Detection engine is being deprecated and new customers are not being onboarded to it, so do not treat that legacy engine as a generally available option for new deployments: Cloudflare bot detection engines.
Review scraping detections and challenge scope
Cloudflare’s scraping-detection guidance describes zone-level detection of anomalous behavior and managed challenges. It says detections are dynamically recalculated rather than permanently flagging a fingerprint from one observation, and advises excluding API calls that should not receive challenges. Check that legitimate API and application routes are not caught by a challenge intended for browser traffic: Cloudflare scraping detection.
Validate a rate-limit rule against the endpoint
Use analytics to verify which endpoint is generating the pattern and whether the configured rule action matches the intended protection. Cloudflare’s examples include counting failed operations and limiting price-lookup operations that could otherwise enable catalog scraping. These are examples of site-owner controls, not a safe or universal request rate for scrapers: Cloudflare rate limiting rules.
Quick wins for a faster PC:
Repair Windows errors before they cause bigger problemsFix Now →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Clear out junk files and repair common Windows errorsFree Scan →Rank #3
Interpret bot scores carefully
Cloudflare documents bot scores from 1 to 99, with lower scores indicating more automated traffic. Its documentation says granular scores require Enterprise Bot Management. A score of 0 means the request was not evaluated; it does not mean the request is human or safe. A low score can also reflect missing metadata, such as a User-Agent stripped by a proxy, so correlate the score with logs and request details before changing a rule: Cloudflare bot score documentation and Cloudflare Bot Analytics.
Common symptoms and next steps
- The request fails once, with no challenge content. Treat it as an unresolved request failure, not a confirmed block. Record the response and check whether the site or connection recovered without changing access controls.
- The response body contains a challenge or interstitial. Compare it with an authorized control and, if you own the site, look for the corresponding security event. Do not attempt to defeat the challenge.
- The status looks successful but the page is wrong. Inspect the body. HTTP-level success does not establish that the expected page was delivered; a challenge or substitute page may still be present.
- The scraper differs from a permitted control. Verify that both requests used the same URL and method, and inspect intended and actual request metadata. Consider whether a proxy or gateway altered the request.
- Only some endpoints are challenged. For an operator, review the matching rule, endpoint analytics, and whether API routes should be excluded from challenge behavior.
- Failures appear after a change in request pattern. Compare timing and site telemetry, but do not infer a general rate threshold. Limits and behavioral signals depend on the site’s configuration.
Or skip the browser setup
If your goal is to capture a visual record of a permitted page while diagnosing what it returns, ScreenshotNeo is a website screenshot API and MCP server. It can capture a page as PNG, JPEG, WebP, or PDF; for a block diagnosis, compare the resulting capture with an authorized expected view, while remembering that a screenshot is not a substitute for HTTP response headers or server logs. Its clean-shot steps accept cookie or consent banners and remove 60+ known consent platforms, newsletter popups, and chat widgets before capture, and each step can be turned off. Bot checks/CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed; responses identify the page verdict and billing status in headers. AI agents can use its MCP server tools: take_screenshot, get_page_info, and capture_pdf.
One GET request returns a capture. Replace the target URL as needed; see the ScreenshotNeo API documentation for request options.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python and Node.js examples are also available if those fit your tooling better:
PC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchimport requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 shots. Sign up for ScreenshotNeo’s free plan.
Best Value
Limits of a client-side diagnosis
Without access to site-side logs or security analytics, a scraper can usually establish that it received unexpected or challenge content, not which rule produced it or whether the site intentionally targeted that client. The reviewed sources do not support a definitive status-code map for blocks, a universal rate limit, or a header that always proves blocking. Treat the conclusion as “likely block” when several signals agree, and respect the site’s rules regardless of whether the cause can be confirmed.
Frequently Asked Questions
Does an HTTP 200 response mean the page was not blocked?
No. Check the response body for a challenge or substitute page; status alone does not establish that the expected page arrived.
Can a missing User-Agent explain an unexpected bot signal?
It can be a contributing factor in Cloudflare’s scoring, including when a proxy strips the header, but it does not prove a block or explain how every site behaves.
Is there a request rate that is safe for every website?
No universal threshold is established here. Follow the site’s published rules and any limits specific to its service.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

