What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
Web scraping is automated HTTP use: a scraper sends requests, interprets responses and headers, and extracts information from the response body when it is appropriate to do so. Reliable scraping depends on more than getting a 200 response. You also need to identify your crawler, handle redirects and errors, respect robots.txt guidance, and slow down when a site signals that it is overloaded.
How does web scraping use HTTP?
HTTP is the protocol layer a scraper uses to communicate with a website. A scraper sends a request for a resource; the server replies with a status code, response headers and usually a body. The body might contain HTML, JSON, an image or another representation. The scraper then checks what it received and extracts the fields it needs.
RFC 9110 defines HTTP semantics, including methods, status codes, headers and resource metadata. In a basic page scraper, the usual sequence is:
- Choose a URL and an appropriate request method, normally GET when retrieving a page.
- Send the request with an honest, stable User-Agent.
- Inspect the response status, final URL and headers such as Content-Type.
- Handle redirects, errors and rate-limit instructions deliberately.
- Parse only a response that is of the expected type, and record the result.
- Request another permitted resource only when your crawl rules and request schedule allow it.
A response body is not automatically usable just because it arrived. A server may return an error page, a login page, a bot check, or a different format than expected. Check the status and content type before parsing. For pages that depend on browser-side JavaScript, a plain HTTP request may not contain the same rendered content a visitor sees; browser automation may be needed, but it does not remove the need to follow site rules or manage request rates.
#1 Best Overall
What the status classes tell you
MDN groups HTTP response status codes into five classes: 1xx informational, 2xx successful, 3xx redirection, 4xx client error and 5xx server error. Treat these as operational signals, not as a substitute for checking the response body and headers.
- 1xx: informational responses during request handling.
- 2xx: the request was handled successfully at the HTTP level; verify that the body is the representation your parser expects.
- 3xx: the resource is being redirected or further action is indicated. Decide whether the destination and method are acceptable, and retain the final URL.
- 4xx: the request encountered a client-side issue, such as a missing or inaccessible resource. Do not blindly repeat an unchanged request.
- 5xx: the server or an upstream service could not complete the request. A retry may be appropriate, but only with limits and any requested delay.
Do you need to follow robots.txt?
For a crawler, robots.txt is guidance about which paths the site asks automated agents to access. RFC 9309 defines the Robots Exclusion Protocol and says its rules are requested crawler behavior, not access authorization. A robots.txt file is publicly visible; it does not protect private data or replace authentication.
Before crawling a site, fetch its top-level /robots.txt, identify the group matching your crawler’s product token or the wildcard group, and apply the most-specific matching allow or disallow rule. After a successful fetch, RFC 9309 requires crawlers to follow parseable rules. Do not interpret an allowed path as a blanket legal permission to copy or republish its contents: site terms, copyright, privacy and jurisdiction can matter separately.
The Tool Desk
Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →If robots.txt is missing or unavailable
RFC 9309 distinguishes an unavailable file from an unreachable one. A 4xx response means the robots.txt file is unavailable; under the protocol, a crawler may access resources. A 5xx response or network failure means it is unreachable, and the crawler must assume complete disallow while that condition persists. The specification permits caching and generally recommends not using a cached copy for more than 24 hours, except when the file is unreachable.
Implement these cases explicitly rather than treating every failure as permission to crawl. Log the fetch result and the policy decision so you can explain why a URL was or was not requested. Robots rules are a crawler protocol, not a way for a site to secure an endpoint.
What should a scraper put in its User-Agent?
Use a stable identifier that accurately names your crawler. Where practical, include a URL or contact route that explains the crawler’s purpose. RFC 9309 says the crawler product token should be a substring of the HTTP User-Agent identification string; its example uses the same product token in the HTTP header and the robots.txt user-agent line.
This consistency helps you select the matching robots.txt group and makes automated traffic identifiable. Do not disguise a crawler as an unrelated browser or another service. Frameworks may have their own robots-specific user-agent setting and fallback behavior; Scrapy documents such settings, so check the framework’s current configuration rather than assuming its default header matches your policy.
Free tools Windows power users keep installed
One-click scans. No signup required.
A simple HTTP request you can inspect
Here is a small Python example using the widely used Requests package. It makes one GET request and prints status, final URL, content type and a short response preview; it does not crawl links or decide whether a path is permitted. Install the dependency with python -m pip install requests, then save and run the script. Replace the example URL and contact information with your own.
import requests
url = "https://example.com/"
headers = {
"User-Agent": "ExampleResearchCrawler/1.0 (+https://example.org/crawler-info)"
}
try:
response = requests.get(url, headers=headers, timeout=20, allow_redirects=True)
print("status:", response.status_code)
print("final URL:", response.url)
print("content type:", response.headers.get("Content-Type", "not stated"))
print("retry after:", response.headers.get("Retry-After", "not stated"))
if response.ok and "text/html" in response.headers.get("Content-Type", "").lower():
print(response.text[:1000])
else:
print("Response was not a successful HTML page; inspect before parsing.")
except requests.RequestException as exc:
print("request failed:", exc)
That example is intentionally a single-request inspection, not a production crawler. A crawler should apply robots rules before fetching each target, use a per-site schedule, cap redirect chains and retries, and parse only responses with the expected content type. For an HTTP-level check without Python, cURL can show response headers and status:
curl -i -A 'ExampleResearchCrawler/1.0 (+https://example.org/crawler-info)' --max-time 20 https://example.com/
In Node.js, the built-in Fetch API can expose the same basic signals. Run this with a recent Node.js version that provides global fetch:
Rank #3
const url = 'https://example.com/';
const response = await fetch(url, {
headers: { 'User-Agent': 'ExampleResearchCrawler/1.0 (+https://example.org/crawler-info)' },
signal: AbortSignal.timeout(20000),
redirect: 'follow'
});
console.log('status:', response.status);
console.log('final URL:', response.url);
console.log('content type:', response.headers.get('content-type'));
console.log('retry after:', response.headers.get('retry-after'));
if (response.ok && (response.headers.get('content-type') || '').includes('text/html')) {
console.log((await response.text()).slice(0, 1000));
} else {
console.log('Response was not a successful HTML page; inspect before parsing.');
}
These snippets illustrate request inspection, not a complete policy engine. In particular, they do not implement robots.txt group matching. Add that policy before expanding a one-page fetcher into a crawler.
What do 429 and 503 mean for a scraper?
HTTP 429 means the client sent too many requests within a period. It is a signal to stop or reduce requests, not an invitation to rotate identities and continue at the same pace. The response may include Retry-After to indicate when to try again.
A 503 response indicates that the service is unavailable; RFC 9110 also defines Retry-After semantics for 503 responses and redirects. If the header is present, it can specify an HTTP date or a delay in seconds. Parse the value as a time instruction, wait at least as long as it indicates, and avoid retrying earlier just because the first retry failed.
Use bounded backoff, not a retry loop
For a 429 or a retryable 503, honor Retry-After when supplied. If it is absent, use a bounded backoff strategy: increase the wait after repeated failures and add jitter so multiple workers do not all retry simultaneously. Set a maximum wait and a retry budget, then stop and record the failure when that budget is exhausted. The HTTP standards describe the response semantics; the exact retry limits are implementation choices and should reflect the site, workload and consequences of failure.
Do not retry every error. A transient network failure or service error may recover; an unchanged request that receives a persistent client error usually needs investigation rather than repetition. If a server sends a redirect with Retry-After, account for that timing too. Keep redirect handling separate from retry handling so a redirect chain cannot become an unbounded sequence of requests.
Windows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallOutdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchHow often should a scraper request a site?
There is no universal request interval established by the standards cited here. Set a conservative per-site rate, keep concurrency bounded, and increase it only when you have a sound reason and the site’s guidance permits it. Reuse cached results where appropriate, avoid fetching unchanged pages needlessly, and stop or slow down when you receive 429, 503 or other signs that requests are not being served normally.
Plan a retry budget and a redirect limit as part of the same schedule. Track whether multiple workers share a host: individually modest loops can become a high combined request rate when run in parallel. A crawler should be able to pause a host without stopping unrelated work, and should not treat a successful response as proof that the site’s preferred load level has been respected.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.What should a scraper log?
Keep enough information to reconstruct what the crawler did and diagnose a change in behavior. A useful request record includes:
- Requested URL, method and timestamp.
- User-Agent and the robots.txt decision applied to the URL.
- Response status, final URL and redirect chain.
- Relevant response headers, especially
Retry-AfterandContent-Type. - Elapsed time, retry count and whether the request ended in a network error.
- Parser outcome, such as expected content found, unexpected format, or parse failure.
Keep logs proportionate and protect any sensitive data your system may encounter. These records make it possible to distinguish a changed page from a rate limit, a redirect, a timeout or a parser problem.
Or skip the browser setup
If the task is to capture a website as an image or PDF rather than extract structured records, ScreenshotNeo offers a one-request screenshot API. It is not a general-purpose web scraper; use it for visual captures. The API accepts a URL and returns an image or PDF. For example, this cURL request saves a WebP screenshot:
Best Value
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp
See the ScreenshotNeo API documentation for request options and response details. Before capture, it can accept cookie or consent banners and remove more than 60 known consent platforms, newsletter popups and chat widgets; each step can be turned off. Bot checks, blank pages, timeouts, failed loads and cache hits are not billed, and response headers say which page verdict occurred and whether the request was billed. Its MCP server provides screenshot and PDF tools for AI agents. The free plan includes 1,000 screenshots per month without a card; paid plans start at $5 for 3,000 shots. Every feature is available on every plan.
Sign up for ScreenshotNeo free: 1,000 screenshots a month, no card required.
Frequently Asked Questions
Does robots.txt make a disallowed URL private?
No. Robots.txt is publicly readable crawler guidance, not authentication or a security control. A site must secure private resources separately.
Does an HTTP 200 response guarantee the page is ready to parse?
No. Check the response content type and body as well as the status; a successful HTTP response may still contain an unexpected page or format.
Can ScreenshotNeo replace a scraper that extracts records from HTML?
No. ScreenshotNeo captures pages as images or PDFs. Use an HTTP client and parser when the goal is to extract structured data.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

