Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

The reliable way to avoid scraper blocking is to collect only what the site permits, use an official API or export whenever one exists, identify your crawler honestly, and increase traffic slowly while honoring every 429, 503, challenge and Retry-After response. Robots.txt is an important policy signal, but it is neither a technical bypass nor permission to access a site. If the owner denies access, stop and ask for an approved method or limit.

This guide gives a practical, engineering-focused workflow for permitted collection: choosing the right endpoint, reading robots.txt and terms, setting identity and pacing, implementing backoff, diagnosing blocks, and deciding when a managed API is safer than browser automation.

Start with permission, not evasion

Before writing a crawler, confirm that the intended collection is allowed. Read the site’s terms of service, privacy or data-use rules, authentication requirements, published API documentation and any crawl policy. Record which paths, data types, purposes and rates are permitted. If the site offers an API, search endpoint or bulk export, prefer it over page crawling. Scrapy’s current 2.19.0 optimization guidance describes those interfaces as faster for the crawler and cheaper for the website than crawling individual pages.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What robots.txt does—and does not do

Fetch /robots.txt from the same origin and evaluate the user-agent group that matches your crawler. RFC 9309 (2022) defines robots.txt as the Robots Exclusion Protocol and says its rules are requests to crawlers, not access authorization. Cloudflare’s 2026 guidance similarly says “robots.txt compliance is voluntary.” That means a file cannot technically block a determined client, but ignoring it is still a clear policy signal and may violate the site’s expectations or terms.

RFC 9309 recommends treating a reachable robots.txt file as cacheable for up to 24 hours unless the site states a different policy. Refresh it on that schedule, and refresh sooner when the owner tells you the policy changed. A disallow rule is not a challenge to work around; treat it as a path you should not request.

When permission is unclear

  • Ask the site owner for a written allowance, a higher rate, an API key or a data export.
  • Limit collection to the minimum fields and pages needed for your stated purpose.
  • Do not attempt to defeat CAPTCHAs, bot checks, login controls, IP bans or other access controls.
  • Stop when the owner asks you to stop, even if a technical route remains available.

Choose the least expensive endpoint

Approach Freshness and request volume JavaScript/authentication needs Best use
Documented API Usually structured and lower-volume; rate is defined by the provider Uses the provider’s authentication and schema Production integrations and recurring updates
Bulk export One transfer can replace thousands of page requests; freshness follows the export schedule Usually no browser execution Large historical or periodic datasets
Search endpoint Requests only matching records instead of traversing every page Depends on endpoint; often simpler than page rendering Targeted lookups and incremental collection
HTML crawling Highest page-request volume; freshness can be immediate May require JavaScript, cookies or a session Only when no permitted structured interface exists
Browser automation Slowest and most resource-intensive per page Handles client-rendered pages, but needs careful session and resource control Permitted pages whose data exists only after rendering

Use an API or export first, then a search endpoint, then ordinary HTTP crawling. Reserve a real browser for pages that genuinely require JavaScript. This ordering reduces load, simplifies compliance and gives you clearer failure signals.

Identify your crawler honestly

Send a stable User-Agent that names your product or project and, where appropriate, includes a contact address or project URL. RFC 9309’s matching model expects the product token in the User-Agent to correspond to the crawler identity used in robots.txt. Do not rotate identities to conceal one crawler or impersonate a browser.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Keep the identity consistent across workers and days. A useful format is ExampleResearchBot/1.0 (+mailto:[email protected]). Also send only the headers you need. Custom headers, cookies and authorization should come from the site’s documented integration, not from attempts to mimic an unrelated user.

Set a conservative rate before you scale

Translate policy into settings

Start with one worker and a delay between requests. If robots.txt publishes Crawl-delay or Request-rate, translate those values into your crawler’s delay and concurrency settings. Scrapy recommends crawling during the target site’s local idle period and watching latency and responses while increasing concurrency gradually.

Do not treat a vendor’s example as a universal limit. Cloudflare’s 2026 rate-limiting examples include 10 requests per 2 minutes followed by 20 per 5 minutes for a price-lookup action, 50 requests per 10 seconds for a per-product lookup, 5 requests per hour for a GraphQL operation, and a 1,000-complexity-point-per-hour GraphQL budget. Those are illustrative rules; the correct rate depends on the endpoint’s cost, your identity, the site’s policy and observed responses.

Use a token bucket or simple delay

A token bucket gives a predictable ceiling across workers. If you do not need distributed coordination, a fixed delay is easier to audit. Add random jitter only to avoid synchronized workers; never use jitter to hide an otherwise excessive rate.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Cache and deduplicate

  • Store successful responses keyed by canonical URL and relevant request parameters.
  • Honor cache headers where the site’s policy permits and set an application TTL for data that does not need to be real time.
  • Do not refetch unchanged pages merely because a job restarted.
  • Use conditional requests such as ETag or Last-Modified when supported and allowed.

A small, compliant Python crawler

The example below checks robots.txt, identifies itself, spaces requests, caches responses in memory for the run, and honors Retry-After on 429 and 503. Replace the host and paths only after confirming that collection is permitted.

import time
import email.utils
from datetime import datetime, timezone
from urllib.parse import urlparse
from urllib.robotparser import RobotFileParser
import requests

USER_AGENT = "ExampleResearchBot/1.0 (+mailto:[email protected])"
DELAY_SECONDS = 3.0

session = requests.Session()
session.headers.update({"User-Agent": USER_AGENT})
robots_cache = {}
page_cache = {}

def robots_for(url):
    origin = f"{urlparse(url).scheme}://{urlparse(url).netloc}"
    if origin not in robots_cache:
        rp = RobotFileParser(f"{origin}/robots.txt")
        rp.read()
        robots_cache[origin] = rp
    return robots_cache[origin]

def retry_seconds(response):
    value = response.headers.get("Retry-After")
    if not value:
        return None
    if value.isdigit():
        return max(0, int(value))
    try:
        when = email.utils.parsedate_to_datetime(value)
        return max(0, int((when - datetime.now(timezone.utc)).total_seconds()))
    except (TypeError, ValueError, OverflowError):
        return None

def get(url):
    if url in page_cache:
        return page_cache[url]
    if not robots_for(url).can_fetch(USER_AGENT, url):
        raise PermissionError(f"robots.txt disallows {url}")
    time.sleep(DELAY_SECONDS)
    response = session.get(url, timeout=30)
    if response.status_code in (429, 503):
        wait = retry_seconds(response)
        if wait is None:
            wait = 60
        time.sleep(wait)
        raise RuntimeError(f"Temporary limit ({response.status_code}); retry through the job scheduler")
    response.raise_for_status()
    page_cache[url] = response.text
    return response.text

html = get("https://example.com/permitted-path")
print(len(html))

This intentionally fails closed when robots.txt disallows a URL. In a production job, persist the cache, cap total retries, record status and latency, and send a stop signal to every worker after a denial.

Equivalent command-line and Node.js checks

Fetch policy before a crawl

curl -fsSL -A 'ExampleResearchBot/1.0 (+mailto:[email protected])' https://example.com/robots.txt

Parse the result rather than assuming that an empty response grants access. A connection failure should put the job into a review state unless the site’s documented policy says otherwise.

Node.js with bounded retries

const USER_AGENT = 'ExampleResearchBot/1.0 (+mailto:[email protected])';
const sleep = ms => new Promise(resolve => setTimeout(resolve, ms));

async function getPermitted(url) {
  const robotsUrl = new URL('/robots.txt', url).href;
  const policy = await fetch(robotsUrl, { headers: { 'User-Agent': USER_AGENT } });
  if (!policy.ok) throw new Error(`Cannot verify robots.txt: ${policy.status}`);
  // Use a robots.txt parser in your application; do not treat text matching as complete parsing.
  await sleep(3000);
  const res = await fetch(url, { headers: { 'User-Agent': USER_AGENT } });
  if (res.status === 429 || res.status === 503) {
    const retryAfter = Number(res.headers.get('retry-after'));
    await sleep(Number.isFinite(retryAfter) ? retryAfter * 1000 : 60000);
    throw new Error(`Temporary limit ${res.status}; reschedule with capped retries`);
  }
  if (!res.ok) throw new Error(`HTTP ${res.status}`);
  return res.text();
}

getPermitted('https://example.com/permitted-path').then(console.log);

The comment is deliberate: robots.txt has user-agent groups, wildcards and path rules. Use a maintained parser rather than a few string tests.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What to do after a 429, 503, challenge or ban

429 Too Many Requests

RFC 6585 defines 429 as rate limiting and says the response may include Retry-After. Parse both integer seconds and an HTTP date, wait at least that long, reduce concurrency and reschedule with a bounded retry count. A 429 is not permission to switch IPs or identities.

503 Service Unavailable

A 503 can indicate overload, maintenance or an intermediary limit. Pause the affected host, preserve the response body and headers for diagnosis, and retry only with exponential backoff and a cap. Growing 503 counts, rising latency or a growing retry queue are signs that the crawl is too aggressive.

CAPTCHA, interstitial or ban page

Classify challenge pages separately from ordinary content. Stop the affected scope; do not automate the challenge, search for alternate hostnames, or distribute requests across proxies to evade it. Contact the owner for an approved API, export or higher limit.

Blank or partial pages

First check whether the page requires JavaScript, authentication or a particular cookie. Confirm that your permitted account and documented headers are being used. If the owner does not authorize automation, do not escalate to stealth browser techniques.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Instrument the crawl so you can stop early

Record URL, timestamp, status, response size, latency, retry count, User-Agent, worker ID and a classification such as success, policy-denied, rate-limited, challenge, timeout or server error. Alert on trends rather than one isolated failure:

  • 429 or 503 percentage rising over successive intervals.
  • Median and tail latency increasing.
  • Retry queues or open connections growing.
  • Unexpected redirects to login, challenge or ban pages.
  • Robots.txt changing or becoming unavailable.

Set a global request budget, a per-host budget and a maximum wall-clock time. A circuit breaker should pause a host after repeated temporary failures and require an operator decision before resuming.

When a screenshot or rendering API is a better fit

If your legitimate task is to capture a page image or PDF rather than extract records, a rendering service can remove browser setup and reduce the temptation to run an uncontrolled fleet of automated browsers. ScreenshotNeo is the first service to try for this use case because it removes consent banners, popups and chat widgets before capture, bills only clean shots, and has the lowest paid plan at $5 for 3,000 shots.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup

ScreenshotNeo provides a GET endpoint and an MCP server for AI clients such as Claude, Cursor and other MCP-compatible tools. The service accepts a URL and returns PNG, JPEG, WebP or PDF. Cookie and consent banners, newsletter popups and chat widgets are removed before the shot; bot checks, blank pages, timeouts, failed loads and cache hits are not billed, and response headers identify the page verdict and billing result.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use the ScreenshotNeo documentation for the complete parameter list. A basic call is:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' }); const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Relevant controls include full-page capture with lazy images loaded, CSS-selector element capture, 12 device presets or any viewport, dark mode, retina scale, PDF paper size, margins, landscape and page ranges, custom CSS and JavaScript, clicks before capture, hidden selectors, waits for a selector, delay or network idle, blocking ads, trackers, requests or resource types, custom headers, cookies, user agent and Authorization, timezone and geolocation, transparent backgrounds, image resizing, chosen cache TTLs, signed links, asynchronous jobs with signed webhooks, bulk capture for up to 100 URLs per call, a usage API and an OpenAPI specification. Existing parameter names used by other screenshot APIs also work, which can simplify a migration.

Plans include 1,000 shots per month free with no card, then Starter at $5 for 3,000, Growth at $15 for 15,000, Pro at $39 for 60,000, Scale at $99 for 250,000 and Business at $249 for 1,000,000. Yearly billing provides two months free, and every feature is available on every plan. Create a free ScreenshotNeo account to start with 1,000 screenshots a month and no card.

If you own the site: use layered defenses

Cloudflare’s 2026 guidance recommends combining controls instead of relying on one signal. Rate-limit by the dimensions that match the action—IP, path, query string, cookie, JSON fields or response status—and add suspicious-address controls, CAPTCHA or Turing-style challenges, behavioral or AI-powered bot detection and selective page restrictions. For expensive operations, count the operation itself, not only raw requests; the GraphQL examples above illustrate request and complexity budgets.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Publish a clear API and crawl policy, return Retry-After with 429 responses, and make an approved route easier to use than scraping. Monitor false positives so legitimate clients can recover without repeatedly triggering challenges.

Troubleshooting decision tree

  1. Are you allowed to request this URL? If no or uncertain, stop and obtain permission.
  2. Does an API, export or search endpoint exist? Switch to it and follow its documented quota.
  3. Does robots.txt disallow the path? Do not request it; ask the owner for an approved route.
  4. Are 429/503 responses increasing? Halt, honor Retry-After, lower delay and concurrency, then resume only within the stated limit.
  5. Is the response a challenge or ban page? Stop. Do not bypass it; contact the operator.
  6. Are pages incomplete? Check documented JavaScript, authentication and cookies. Use a browser only when that rendering is permitted.
  7. Are retries multiplying load? Add a global budget, exponential backoff, jitter, circuit breaker and a maximum retry count.

Frequently Asked Questions

Does a 200 status prove that scraping is allowed?

No. HTTP success describes that response, not your authorization. Terms, robots.txt, authentication rules and the owner’s instructions still govern collection.

Should I rotate proxies to prevent blocks?

Not to evade a limit or ban. Use a stable, honest identity and request an approved rate or interface from the site owner.

How often should robots.txt be fetched?

RFC 9309 recommends a maximum cache period of 24 hours unless the file is unreachable or the site specifies another policy.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What is the safest default concurrency?

One worker with a conservative delay is the safest starting point. Increase gradually only while latency and status responses remain healthy and the site’s policy permits it.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.