Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
Reliable web scraping starts before the first request: check whether an official API or feed can provide the data, review the target site’s crawler guidance, identify your scraper honestly, and collect at a modest rate. Then make failures visible, batch work, and validate the records—not just the HTTP responses. These practices reduce avoidable load and make the resulting data easier to trust.
1. Check for an API or feed before scraping pages
Look for a documented API, downloadable dataset, RSS feed, or other official interface that supplies the fields you need. Compare it with page scraping on permission and terms, field coverage, freshness, quotas, server impact, operational effort, and how you will validate the output. An official interface may be a better fit, but it is not automatically complete or suitable: verify its terms and limits for your use.
If page scraping remains the right choice, define the minimum data and pages needed. A narrow collection plan is easier to review, cheaper to operate, and less likely to send unnecessary requests.
2. Read robots.txt for the exact origin and crawler
Fetch the top-level /robots.txt for the exact scheme, host, and port you plan to access. A rule for one origin does not necessarily apply to another. Match your crawler’s User-Agent product token to the relevant group and follow the parseable rules that apply to it. Google’s explanation of robots.txt parsing is useful context, but the site’s own file and the protocol govern the crawler guidance you inspect: Google’s robots.txt specification guide.
#1 Best Overall
The IETF’s Robots Exclusion Protocol, RFC 9309, states that its rules “are not a form of access authorization.” Robots.txt is crawler guidance, not permission to access restricted material. Review credentials, technical controls, site terms, and applicable legal and privacy obligations separately. RFC 9309 also warns that listing paths in robots.txt can make them discoverable; it is not a security mechanism. Read RFC 9309.
3. Identify your crawler clearly
Send a descriptive User-Agent that identifies the software and purpose. Do not impersonate a browser or another crawler to evade a site’s rules. RFC 9110 §10.1.5 says: “A user agent SHOULD send a User-Agent header field in each request unless specifically configured not to do so.” The same section cautions against needlessly fine-grained detail, which can increase latency and fingerprinting risk. Read RFC 9110.
For example, use a stable product token and a contact channel your team actually monitors, such as ExampleResearchBot/1.0 (+https://example.org/bot-info). Replace the example with your own identity and a real information page; do not claim affiliation with a site or organization you do not represent.
The Tool Desk
Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →4. Start with a conservative per-host request rate
Set limits per host rather than only limiting total requests across your entire job. A fast crawl against one small site can impose more concentrated load than the same request count spread across many hosts. Start cautiously, observe responses and latency, and reduce the rate if the site appears strained or signals that you should slow down.
Amazon Web Services gives illustrative—not universal—examples: one request every 10–15 seconds for small or medium-sized sites, and one to two requests per second for larger sites or sites with explicit crawl permissions. These are vendor examples, not safe thresholds for every site. Site circumstances and explicit guidance take precedence. See AWS’s ethical web crawler guidance.
5. Treat 429 and persistent 403 responses as stop signals
HTTP responses are feedback about the interaction, not merely obstacles to work around. AWS recommends pausing when a server returns 429 Too Many Requests. Do not keep sending requests at the same pace while waiting for a successful response. If the site provides a retry instruction, follow it; otherwise, pause and reassess before resuming at a lower rate.
A 403 Forbidden means the server is refusing the request. AWS advises considering stopping if 403 responses continue. Do not respond to restrictions by increasing traffic, rotating identities to get around them, or repeatedly retrying the same blocked URLs. Record the response and resolve access questions through an authorized route.
6. Use sitemaps to focus discovery
When available, use the site’s sitemap to identify candidate pages instead of repeatedly discovering links across the whole site. A sitemap can make the crawl more targeted, but it does not grant access or override robots.txt, terms, or other restrictions. Filter its URLs to the pages and fields your project actually needs, and check the applicable crawler rules before fetching them.
7. Crawl in small, manageable batches
Split large URL sets into batches rather than launching one enormous job. Batching helps distribute load, makes failures easier to isolate, and reduces the risk that timeouts or resource constraints derail the entire collection. AWS recommends dividing URL sets into smaller batches; choose a batch size appropriate to your host limits, runtime, and recovery needs rather than treating any one size as universal.
Persist progress between batches. A useful record includes the requested URL, collection time, response status, and whether parsing succeeded. That lets a restarted job continue with unfinished work instead of needlessly repeating successful requests.
8. Make retries bounded and failures visible
Track failed URLs and status codes, and set a finite retry policy. A retry loop without a stopping condition can amplify load precisely when a site is struggling. Use increasing pauses between permitted retries as an implementation choice, but do not treat a retry schedule as permission to continue against a 403 or to ignore a 429. Respect site instructions and stop when the response indicates that access is refused.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Keep distinct outcomes distinct: a network timeout, a server error, a rate limit, an access refusal, and a successful response with an unusable page require different handling. Log enough context to diagnose a failure without storing credentials or unnecessary personal data in logs.
Rank #3
9. Validate the data, not just the request
An HTTP success status does not prove that the intended record was collected. Pages can change, fields can be absent, and parsers can return empty or malformed values without raising an obvious error. Add checks that reflect your dataset’s schema and purpose:
- Confirm required fields are present and parseable.
- Check stable keys for duplicates and decide how duplicate records should be handled.
- Track parsing failures and incomplete pagination.
- Compare observed counts with the pages or records you expected to collect, investigating unexplained gaps rather than assuming a universal acceptable threshold.
- Check timestamps for plausibility and record when the collection occurred.
There is no universal validation threshold that makes every scraped dataset reliable. Define checks around the data’s intended use, and preserve enough provenance to revisit questionable records.
10. Recheck rules and extraction assumptions
Web pages, robots.txt files, and site behavior can change. Keep selectors and crawl rules observable, record the collection date, and revalidate them before relying on a new run or an important downstream decision. When a field unexpectedly goes missing or record counts shift, inspect representative pages and logs before accepting the output as complete.
Do not assume that a selector that worked once will remain valid indefinitely. Treat unexplained shifts in failures, page structure, or extracted values as reasons to investigate—not as noise to discard automatically.
11. Choose a rendering method that fits the data
If the information is present in the HTML response, a simple HTTP client may avoid the overhead of rendering a browser page. If the relevant content appears only after JavaScript runs or you need a visual record, a browser-based capture may be more appropriate. In either case, follow the same access, rate, and validation practices: a browser does not make restricted collection permissible or remove the need to control load.
For a do-it-yourself browser workflow, use an authorized browser automation setup, navigate only to allowed pages, wait for the specific content you need rather than an arbitrary long delay, and save the output with its URL and collection timestamp. Verify that the expected content actually rendered before treating a screenshot as evidence. Browser rendering can take more resources and time than fetching static HTML, so reserve it for pages or outputs that require it.
Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server for developers. One GET request can return a PNG, JPEG, WebP, or PDF capture. For example, this cURL request saves a WebP screenshot of Stripe; replace the target URL with a page you are authorized to capture:
Free tools Windows power users keep installed
One-click scans. No signup required.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
See the ScreenshotNeo documentation for request options and setup. Before capture, it can accept cookie or consent banners like a visitor and remove more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers indicate the page verdict and billing status. Its MCP server provides take_screenshot, get_page_info, and capture_pdf tools for AI agents and MCP clients. The free plan includes 1,000 screenshots per month without a card; paid plans start at $5 for 3,000 shots.
Sign up for 1,000 free screenshots a month with no card.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Troubleshooting common scraping failures
Robots.txt is missing or unclear
Confirm you requested the correct origin, including scheme, host, and port. Do not treat an absent or confusing file as authorization. Review the site’s other access conditions and terms, and avoid paths or activity that access controls prohibit.
The server returns 429
Pause requests to that host. Review your per-host rate and resume only cautiously if appropriate, following any instructions the server supplies. Avoid parallel workers that continue sending requests while another part of the job is paused.
Do these 3 things before closing this tab:
1Fix the driver behind crashes, sound loss and screen glitches2Repair Windows errors before they cause bigger problems3Scan for outdated or missing drivers - takes under a minuteThe server keeps returning 403
Stop repeated attempts and determine whether you have permission and the required access. Do not try to defeat the refusal by disguising the crawler or escalating request volume.
Requests succeed but fields are empty
Inspect a small sample of returned pages and compare the actual response with the assumptions in your parser. The content may have moved, may require rendering, or may be unavailable in the response you received. Update extraction only where access is permitted, then rerun the relevant validation checks.
Best Value
The job times out or stops partway through
Reduce the batch size, persist completed records, and capture failed URLs and statuses. Resume from the saved progress rather than repeating the entire collection. If timeouts coincide with server errors or signs of strain, reduce load instead of increasing concurrency.
Performance, reliability, and cost trade-offs
Reliability is not the same as maximum throughput. A slower, bounded crawl with recoverable batches and validated output is often more useful than a fast run that overloads a host or silently produces incomplete records. Browser rendering can handle pages that need client-side execution but adds runtime and resource overhead; static HTTP fetching can be lighter when the needed data is already in the response.
Recommended Free Tools
For larger or scheduled jobs, batching and saved progress can make operational failures easier to contain. AWS notes that Lambda can suit short-lived, event-driven tasks, but cloud infrastructure is not required for ordinary scraping. Choose compute based on the job’s runtime, schedule, and recovery needs, not on an assumption that all crawlers need a cloud platform.
FAQ
Does robots.txt give permission to scrape a page?
No. RFC 9309 explicitly says robots.txt is not access authorization. It expresses crawler guidance; review permission, access controls, terms, and applicable obligations separately.
Is there one request rate that is safe for every website?
No. AWS’s example rates are illustrative guidance, not universal thresholds. Use site-specific directions and begin conservatively.
Should I retry a page after a 403?
Do not retry persistently refused access as if it were a transient error. AWS advises considering stopping if 403 responses continue.
PC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchCan I rely on a successful HTTP status as proof that my dataset is complete?
No. Validate fields, duplicates, pagination, counts, parsing, and timestamps against the needs of your collection.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

