Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

To reduce the risk of a website blocking your scraper, use an authorized access route, check the site’s current terms and robots.txt, request only what you need at a conservative pace, and stop or slow down when the server signals a limit or refusal. No universal delay or technique guarantees access: each site sets its own policies and technical thresholds.

Start with permission and an approved access route

Before writing a crawler, check whether the site offers an official API, data export, licensed feed, or another documented way to obtain the information. An API or feed is generally the first route to investigate because the provider defines it as an intended access method; not every site offers one.

Review the target’s current terms and any restrictions relevant to your purpose and jurisdiction. Whether a particular scraping project is lawful depends on its facts and location; general crawler guidance cannot settle that question. If you need access that the site has not made available, ask the owner for permission rather than trying to defeat a refusal.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Check robots.txt correctly

For a site at https://example.com, its robots file is normally at https://example.com/robots.txt. Check the file for rules that apply to your crawler’s product token and the paths you intend to request. RFC 9309 describes robots.txt as crawler instructions, not authorization: “These rules are not a form of access authorization.” RFC 9309 also makes clear that it is not a security boundary or permission to access restricted content.

If the file loads

Follow the parseable rules relevant to your crawler. Keep in mind that robots.txt expresses crawler preferences; it does not replace the site’s terms, grant access to private data, or make a prohibited use permissible.

If the file is unreachable

Under RFC 9309, when a crawler cannot fetch robots.txt because of network or server errors, it must assume complete disallow. Do not treat a temporary inability to read the file as permission to crawl. The RFC also says crawlers should not use a cached robots.txt copy for more than 24 hours unless the file is unreachable. That recommendation applies to caching this file, not to the interval between ordinary page requests. Read the protocol.

Build a conservative crawler

There is no source-backed request interval that guarantees a site will accept a scraper. Begin with a small, low-impact workload, watch the site’s responses, and adjust downward or stop if it signals strain or refusal. Fetch only the pages and fields your task requires.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Limit concurrency. Avoid sending a large burst of simultaneous requests. Start with one worker unless the site documents a higher allowed level.
  • Avoid redundant fetches. Cache results responsibly and use conditional requests when the server supports them, so unchanged pages need not be downloaded repeatedly.
  • Identify your crawler honestly. RFC 9309 recommends that a crawler’s identification string describe its purpose and include its product token. Do not pretend to be a browser or rotate identities to hide the crawler.
  • Use a bounded workload. Scope the URLs, frequency, and duration in advance. Stop when you have the needed data rather than continuing to recrawl indefinitely.
  • Monitor responses. Record status codes and relevant response headers so your crawler can distinguish rate limiting, temporary outages, and refusals.

These are cautious operating practices, not a formula for bypassing a site’s controls. If an owner asks you to stop, stop.

Respond to HTTP status codes instead of retrying blindly

A retry policy should branch on the response. Repeating the same request at the same pace can worsen rate limiting, and an unchanged request after a refusal is expected to fail again.

Response What it signals Appropriate action
429 Too Many Requests The client sent too many requests in a period. The server may include Retry-After. Pause, reduce the request rate, and honor the indicated wait if present. Do not resume at the same pace. MDN: 429
503 Service Unavailable The server is temporarily unable to handle the request; a Retry-After header may give an estimated recovery time. Wait for the stated recovery period if supplied. If the service remains unavailable, stop and reassess rather than repeatedly probing it. MDN: 503
403 Forbidden The server understood the request and refused to process it. Treat it as a refusal. Do not repeat an unchanged request or disguise, reroute, or otherwise evade the restriction. Seek permission or use an approved source. MDN: 403

Understand Retry-After

The Retry-After value can be an HTTP date or a non-negative number of seconds. When the server supplies it, use that value to determine how long to wait before a follow-up request; it is not a signal to keep retrying before the indicated time. MDN explains the header’s formats and behavior.

Choose a data-collection route that fits the job

Compare access routes by permission and terms compliance, whether the provider offers an API or export, whether the route respects robots.txt and server limits, the freshness and completeness you need, and the maintenance burden. A page scraper may be a poor fit when an API or licensed data source provides the same information more reliably. If no authorized route is clear, clarify access with the site owner before collecting data.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

When your task is a screenshot, use a screenshot route

If you need a visual record of a page rather than its underlying structured data, scraping and parsing HTML may be unnecessary. ScreenshotNeo is a website screenshot API and MCP server for developers. A single GET request can return a PNG, JPEG, WebP, or PDF. Its clean-shot flow accepts cookie and consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; individual steps can be turned off. A screenshot is billed only when it is a clean shot: bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, with the outcome identified in response headers.

That does not grant permission to access a site or override its terms, robots rules, or technical restrictions. Use it only for pages and purposes you are authorized to access.

Or skip the browser setup

For a one-off capture, make a GET request to the API. This cURL example saves a WebP screenshot of Stripe:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Replace YOUR_API_KEY with your key and change the target URL to a page you are permitted to capture. See the ScreenshotNeo API documentation for request parameters and response details.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Cookie banners, popups, and chat widgets are removed before the shot; bot checks, blank pages, and failed loads are never billed. An MCP server lets AI agents use screenshot tools, and 1,000 screenshots a month are free with no card; paid plans start at $5 for 3,000. Sign up for free.

Common problems and what to do

The crawler receives 429 repeatedly

Reduce concurrency and request frequency, then wait for Retry-After if provided. Repeated 429 responses mean the current pace is still too high or the site is not accepting the workload. If the site does not recover, stop and seek an approved route instead of increasing retries.

The crawler receives 403

Stop the request sequence. A 403 is a refusal, not a puzzle to solve by changing identities or routing through proxies. Check whether you have permission and whether an official API or export is available.

The crawler receives 503

Check for a Retry-After value and wait until the stated recovery time. A 503 can indicate a temporary service problem; avoid adding load while the server reports it cannot serve requests.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

robots.txt cannot be fetched

Do not assume the site permits crawling. RFC 9309 directs crawlers to assume complete disallow when robots.txt is unreachable because of network or server errors. Retry the file fetch later if appropriate, but do not proceed with page crawling in the meantime.

The site changes its rules or page structure

Recheck the current robots.txt and terms before a new crawl, and validate that your parser still extracts the intended fields. If the site’s rules or access signals are unclear, pause and ask the owner rather than broadening the crawl.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Keep the operating boundary clear

Responsible scraping is not about finding a delay or disguise that defeats blocking. It is about using a permitted route, honoring crawler instructions and server signals, minimizing load, and stopping when access is denied. The target’s current policies and applicable legal rules depend on the site, purpose, and jurisdiction; verify those specifics for your project.

Frequently Asked Questions

Does robots.txt mean I have permission to scrape a site?

No. RFC 9309 says robots.txt rules are not a form of access authorization. They communicate crawler preferences and do not grant permission.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Is there a safe number of seconds to wait between requests?

No universal interval is established that guarantees acceptance. Use a conservative rate, follow the site’s own guidance, and respond to its status codes and headers.

Should I use rotating proxies or CAPTCHAs to avoid a block?

No. Do not evade a site’s refusal or conceal crawler identity. Stop and seek permission or an approved API or data source.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.