Free tools Windows power users keep installed
One-click scans. No signup required.
Reliable web scraping is less about sending more requests and more about making controlled, observable requests that respect the target site. Check the correct robots.txt, identify your crawler, pace traffic, use sitemaps, batch large jobs, handle failures deliberately, and verify that the rules you read apply to the exact host, protocol, and port you are fetching.
This guide follows the Robots Exclusion Protocol in RFC 9309, operational guidance from Amazon Web Services (AWS), and Google’s documentation about crawler behavior. None of those sources turns a scraping technique into legal permission: terms of service, contracts, privacy obligations, and local law still depend on the target and your situation.
1. Check crawler rules before fetching
Start by requesting the target’s robots.txt and parsing its rules before collecting pages. RFC 9309 defines the Robots Exclusion Protocol (REP) as a crawler coordination mechanism. Its key limitation is explicit: “These rules are not a form of access authorization.” A site can disallow a path in robots.txt, but an allow rule does not replace authentication, contractual review, or legal analysis.
What a successful fetch means
When robots.txt is retrieved successfully, crawlers must follow its parseable rules under RFC 9309. The standard sets a minimum parsing limit of 500 KiB. It also says crawlers should follow at least five consecutive redirects while retrieving the file and should not use a cached copy for more than 24 hours unless the file is unreachable.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
#1 Best Overall
What to do when retrieval fails
- If server or network errors make the file unreachable, RFC 9309 says crawlers must assume complete disallow.
- If the response is unavailable in the 4xx range, the RFC says crawlers may access resources on that server.
- Record the status, retrieval time, final URL, and parsed rules so an operator can audit why a URL was queued.
Those protocol outcomes are distinct from permission. Treat an ambiguous or unstable response conservatively and seek the site owner’s direction when the collection is consequential.
2. Identify your crawler clearly
Send a descriptive HTTP User-Agent rather than pretending to be a browser or using a generic library default. AWS recommends identifying the crawler and commonly including contact information. A useful identity states the project name, an owned information page or email address, and, where appropriate, a way to request reduced traffic.
Transparency is not a bypass
A clear identity helps a site operator diagnose traffic and contact you, but it does not guarantee access. Do not rotate identities to evade a block, forge another organization’s crawler name, or treat a friendly user agent as authorization. Keep the user-agent stable enough that logs and support requests can be correlated.
3. Pace requests and react to load
Request pacing protects both your job and the site. AWS gives contextual examples—not universal limits—of one request every 10–15 seconds for small or medium-sized websites, and 1–2 requests per second for larger websites or sites with explicit crawl permission. Select a rate from the target’s published guidance, expected page cost, and observed responses rather than copying a number blindly.
Signals that require a slower crawl
- HTTP 429: pause the crawl, honor any server-provided retry timing, and reduce concurrency before resuming.
- HTTP 403: check your authorization and crawler rules. AWS advises considering a stop if 403 responses continue; do not keep increasing attempts.
- 5xx responses: treat repeated server errors as a site-health signal, not as permission to hammer the endpoint.
- Rising latency: reduce pressure and record the change. Google documents slower response times, 5xx errors, and rate-limit signals such as 429 as factors that reduce Google’s crawl capacity; they are useful operational indicators for your own monitor too, not a universal limit for every scraper.
Use a bounded queue, a small concurrency limit, and a pause between requests. Make the pause adjustable so an operator can slow or stop a run without redeploying code.
4. Use sitemaps to focus collection
Discover URLs from the site owner’s sitemap when one is available instead of guessing paths or crawling every link indiscriminately. AWS recommends using sitemaps to focus on important pages. This reduces duplicate discovery and gives you a reviewable inventory before downloading content.
Validate the inventory
- Normalize URLs consistently so fragments and obvious duplicates do not create extra requests.
- Keep the sitemap location and retrieval time with each URL record.
- Filter to the content types and paths your project actually needs.
- Re-check
robots.txtscope before fetching URLs discovered from a sitemap.
A sitemap is a discovery aid, not a grant of access. Its presence does not override disallow rules, authentication requirements, or site terms.
5. Divide large jobs into batches
Split a long URL list into smaller batches. AWS recommends batching to distribute load and reduce timeout or resource constraints. Operationally, batches also create checkpoints: you can record which URLs were attempted, succeeded, skipped, or failed and resume without restarting the entire run.
A practical batch record
- Batch identifier and planned URL count
- Start and end times
- Request rate and concurrency settings
- Status-code totals, timeout totals, and bytes received
- Output location and a list of unresolved URLs
Keep batches small enough that a failure is diagnosable, but large enough to avoid excessive setup overhead. If the target begins returning 429, 403, or 5xx responses, stop the current batch and reassess rather than automatically creating more retries.
6. Handle robots.txt outcomes deliberately
Do not reduce robots handling to “file exists” versus “file missing.” Store the response class and apply the protocol’s outcome rules.
Rank #3
Successful, redirected, or oversized files
Follow parseable rules after a successful retrieval, including after redirects within the standard’s guidance. Enforce the 500 KiB parsing minimum from RFC 9309 and record when content exceeds your parser’s supported limit. A parser that silently truncates rules can produce an unsafe allow decision.
Temporary server or network failure
RFC 9309 requires complete disallow when the file is unreachable because of server or network errors. Pause the affected host, retry the file carefully, and retain the last known result for audit purposes without treating it as current authorization.
The Tool Desk
Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →4xx unavailable response
The RFC says crawlers may access resources when robots.txt is unavailable in the 4xx range. That is a protocol behavior, not a statement that the site welcomes your collection. Apply your separate legal, contractual, and security checks before proceeding.
7. Check the rule scope for the exact origin
Robots rules are origin-scoped. Google’s documentation says a file applies only to the host, protocol, and port where it is hosted. A file at https://www.example.com/robots.txt does not automatically cover https://example.com, another subdomain, HTTP, or a different port.
Scope checklist
- Compare the URL’s scheme: HTTP and HTTPS are separate.
- Compare the hostname, including
wwwand other subdomains. - Compare the port when it is non-default.
- Fetch and evaluate the applicable file for each distinct origin.
- Do not assume a sitemap hosted on one origin changes robots rules on another.
Google-specific behavior should not be generalized to every crawler. Google says it generally caches robots.txt for up to 24 hours and may cache longer when refreshing is impossible. Its documentation also describes stopping crawling for the first 12 hours after a fetch failure, then using the last good version for the next 30 days while trying to fetch again. Those are Google crawler behaviors, not universal requirements for your software.
Make reliability observable
Reliability is a property you can inspect. For every request, record the URL, origin, timestamp, status code, response latency, response size, retry count, and final outcome. Separate “not fetched because disallowed” from “failed after fetching” and “fetched but content was incomplete.” Alert on sustained 429, 403, and 5xx rates, unusual latency, and a sudden fall in successful records.
Define completion before you start
Choose acceptance checks appropriate to the collection: expected status classes, minimum content length, required fields, and a stable URL or document identifier. Preserve failed and skipped URL lists rather than silently dropping them. A technically successful HTTP response can still contain an error page, a login wall, or an incomplete document.
When a browser-rendered capture is the simpler option
If the job needs a visual record of a page rather than parsed HTML, a screenshot API can avoid maintaining browser infrastructure. ScreenshotNeo is the first service to try here because it removes consent banners, newsletter popups, and chat widgets before capture, bills only clean shots, and has a $5 paid plan for 3,000 shots.
Or skip the browser setup
One GET request returns a PNG, JPEG, WebP, or PDF. The API accepts options for full-page capture with lazy images loaded, CSS-selector element capture, dark mode, device presets or custom viewports, retina scale, PDF paper settings and page ranges, custom CSS and JavaScript, clicks, waits, blocked requests, headers, cookies, user agents, authorization, timezone, geolocation, transparent backgrounds, resizing, caching TTLs, signed links, asynchronous webhooks, bulk capture of up to 100 URLs per call, usage reporting, and an OpenAPI specification. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing result.
cURL (see the ScreenshotNeo documentation):
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python:
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
ScreenshotNeo also provides an MCP server with take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients. Its Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.
Troubleshooting common failures
“The scraper is suddenly getting 429”
Pause, lower concurrency, lengthen the interval, and inspect whether another worker is sharing the same origin. Resume only after the site is healthy and keep the 429 period in your run log.
Best Value
“403 continues after retries”
Stop rather than escalating retries. Verify credentials, user-agent identity, robots scope, and any published access process. Ask the site owner if your use is authorized.
“robots.txt cannot be fetched”
Classify the failure. Server or network errors require complete disallow under RFC 9309. A 4xx unavailable response has a different protocol outcome, but it still does not settle legal or contractual permission.
“Results are incomplete despite 200 responses”
Check for login pages, consent overlays, client-side rendering, truncated bodies, and timeouts recorded as successful by your client. Validate required fields and retain the raw response or a content hash for investigation.
Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minutePC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11FAQ
Frequently Asked Questions
Is robots.txt legally binding?
No. RFC 9309 describes it as a crawler coordination protocol and expressly says its rules are not access authorization. Review the target’s terms, contracts, laws, and any authentication requirements separately.
What is a universally safe scraping rate?
There is no universal rate. AWS’s 10–15-second and 1–2-request-per-second figures are contextual examples. Use the site’s guidance and slow or stop when latency, 429, 403, or 5xx signals worsen.
Does a sitemap mean every listed URL may be scraped?
No. A sitemap helps discovery and prioritization. Check the applicable robots rules and your independent permission and legal requirements for each origin.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.
Quick wins for a faster PC:
Clear out junk files and repair common Windows errorsFree Scan →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →

