Quick wins for a faster PC:
Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Clear out junk files and repair common Windows errorsFree Scan →Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
To reduce the risk of a website blocking your scraper, use an authorized access route, check the site’s current terms and robots.txt, request only what you need at a conservative pace, and stop or slow down when the server signals a limit or refusal. No universal delay or technique guarantees access: each site sets its own policies and technical thresholds.
Start with permission and an approved access route
Before writing a crawler, check whether the site offers an official API, data export, licensed feed, or another documented way to obtain the information. An API or feed is generally the first route to investigate because the provider defines it as an intended access method; not every site offers one.
Review the target’s current terms and any restrictions relevant to your purpose and jurisdiction. Whether a particular scraping project is lawful depends on its facts and location; general crawler guidance cannot settle that question. If you need access that the site has not made available, ask the owner for permission rather than trying to defeat a refusal.
Check robots.txt correctly
For a site at https://example.com, its robots file is normally at https://example.com/robots.txt. Check the file for rules that apply to your crawler’s product token and the paths you intend to request. RFC 9309 describes robots.txt as crawler instructions, not authorization: “These rules are not a form of access authorization.” RFC 9309 also makes clear that it is not a security boundary or permission to access restricted content.
#1 Best Overall
If the file loads
Follow the parseable rules relevant to your crawler. Keep in mind that robots.txt expresses crawler preferences; it does not replace the site’s terms, grant access to private data, or make a prohibited use permissible.
If the file is unreachable
Under RFC 9309, when a crawler cannot fetch robots.txt because of network or server errors, it must assume complete disallow. Do not treat a temporary inability to read the file as permission to crawl. The RFC also says crawlers should not use a cached robots.txt copy for more than 24 hours unless the file is unreachable. That recommendation applies to caching this file, not to the interval between ordinary page requests. Read the protocol.
Build a conservative crawler
There is no source-backed request interval that guarantees a site will accept a scraper. Begin with a small, low-impact workload, watch the site’s responses, and adjust downward or stop if it signals strain or refusal. Fetch only the pages and fields your task requires.
PC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware match- Limit concurrency. Avoid sending a large burst of simultaneous requests. Start with one worker unless the site documents a higher allowed level.
- Avoid redundant fetches. Cache results responsibly and use conditional requests when the server supports them, so unchanged pages need not be downloaded repeatedly.
- Identify your crawler honestly. RFC 9309 recommends that a crawler’s identification string describe its purpose and include its product token. Do not pretend to be a browser or rotate identities to hide the crawler.
- Use a bounded workload. Scope the URLs, frequency, and duration in advance. Stop when you have the needed data rather than continuing to recrawl indefinitely.
- Monitor responses. Record status codes and relevant response headers so your crawler can distinguish rate limiting, temporary outages, and refusals.
These are cautious operating practices, not a formula for bypassing a site’s controls. If an owner asks you to stop, stop.
Respond to HTTP status codes instead of retrying blindly
A retry policy should branch on the response. Repeating the same request at the same pace can worsen rate limiting, and an unchanged request after a refusal is expected to fail again.
| Response | What it signals | Appropriate action |
|---|---|---|
429 Too Many Requests |
The client sent too many requests in a period. The server may include Retry-After. |
Pause, reduce the request rate, and honor the indicated wait if present. Do not resume at the same pace. MDN: 429 |
503 Service Unavailable |
The server is temporarily unable to handle the request; a Retry-After header may give an estimated recovery time. |
Wait for the stated recovery period if supplied. If the service remains unavailable, stop and reassess rather than repeatedly probing it. MDN: 503 |
403 Forbidden |
The server understood the request and refused to process it. | Treat it as a refusal. Do not repeat an unchanged request or disguise, reroute, or otherwise evade the restriction. Seek permission or use an approved source. MDN: 403 |
Understand Retry-After
The Retry-After value can be an HTTP date or a non-negative number of seconds. When the server supplies it, use that value to determine how long to wait before a follow-up request; it is not a signal to keep retrying before the indicated time. MDN explains the header’s formats and behavior.
Choose a data-collection route that fits the job
Compare access routes by permission and terms compliance, whether the provider offers an API or export, whether the route respects robots.txt and server limits, the freshness and completeness you need, and the maintenance burden. A page scraper may be a poor fit when an API or licensed data source provides the same information more reliably. If no authorized route is clear, clarify access with the site owner before collecting data.
When your task is a screenshot, use a screenshot route
If you need a visual record of a page rather than its underlying structured data, scraping and parsing HTML may be unnecessary. ScreenshotNeo is a website screenshot API and MCP server for developers. A single GET request can return a PNG, JPEG, WebP, or PDF. Its clean-shot flow accepts cookie and consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; individual steps can be turned off. A screenshot is billed only when it is a clean shot: bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, with the outcome identified in response headers.
That does not grant permission to access a site or override its terms, robots rules, or technical restrictions. Use it only for pages and purposes you are authorized to access.
Rank #3
Or skip the browser setup
For a one-off capture, make a GET request to the API. This cURL example saves a WebP screenshot of Stripe:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Replace YOUR_API_KEY with your key and change the target URL to a page you are permitted to capture. See the ScreenshotNeo API documentation for request parameters and response details.
Cookie banners, popups, and chat widgets are removed before the shot; bot checks, blank pages, and failed loads are never billed. An MCP server lets AI agents use screenshot tools, and 1,000 screenshots a month are free with no card; paid plans start at $5 for 3,000. Sign up for free.
Common problems and what to do
The crawler receives 429 repeatedly
Reduce concurrency and request frequency, then wait for Retry-After if provided. Repeated 429 responses mean the current pace is still too high or the site is not accepting the workload. If the site does not recover, stop and seek an approved route instead of increasing retries.
The crawler receives 403
Stop the request sequence. A 403 is a refusal, not a puzzle to solve by changing identities or routing through proxies. Check whether you have permission and whether an official API or export is available.
The crawler receives 503
Check for a Retry-After value and wait until the stated recovery time. A 503 can indicate a temporary service problem; avoid adding load while the server reports it cannot serve requests.
robots.txt cannot be fetched
Do not assume the site permits crawling. RFC 9309 directs crawlers to assume complete disallow when robots.txt is unreachable because of network or server errors. Retry the file fetch later if appropriate, but do not proceed with page crawling in the meantime.
The site changes its rules or page structure
Recheck the current robots.txt and terms before a new crawl, and validate that your parser still extracts the intended fields. If the site’s rules or access signals are unclear, pause and ask the owner rather than broadening the crawl.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Keep the operating boundary clear
Responsible scraping is not about finding a delay or disguise that defeats blocking. It is about using a permitted route, honoring crawler instructions and server signals, minimizing load, and stopping when access is denied. The target’s current policies and applicable legal rules depend on the site, purpose, and jurisdiction; verify those specifics for your project.
Best Value
Frequently Asked Questions
Does robots.txt mean I have permission to scrape a site?
No. RFC 9309 says robots.txt rules are not a form of access authorization. They communicate crawler preferences and do not grant permission.
Is there a safe number of seconds to wait between requests?
No universal interval is established that guarantees acceptance. Use a conservative rate, follow the site’s own guidance, and respond to its status codes and headers.
Should I use rotating proxies or CAPTCHAs to avoid a block?
No. Do not evade a site’s refusal or conceal crawler identity. Stop and seek permission or an approved API or data source.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

