Use a discovery-first checker: read the target host’s robots.txt, collect sitemap URLs, request only approved URL patterns at a controlled rate, and write a report containing the requested URL, final URL, status, headers, timestamp, and a task-specific content test. The templates below provide a small Python implementation and a scalable Scrapy version. They check whether resources respond and whether the response contains what your project needs; they do not turn crawler guidance into authorization.
What a website-resource checker should do
A useful checker separates four jobs:
- Inputs: a starting host or approved URL list, resource types or path patterns, request limits, and an output format such as CSV or JSON Lines.
- Discovery: inspect the host’s root
robots.txtand sitemap references, then expand sitemap indexes if necessary. - Request: fetch only relevant URLs, follow redirects according to your policy, and retain the final URL plus response metadata.
- Report: record the requested URL, final response URL, HTTP status, selected headers, timestamp, and a content check specific to your requirement.
“HTTP 200” means that a server returned a successful response. It does not prove that the page contains a product, API field, canonical tag, image, download, or other resource you require. Keep transport status and content validation as separate fields.
Can I use robots.txt to tell a scraper what not to crawl?
Yes, as crawler guidance. Google Search Central defines it as a file that tells search-engine crawlers which URLs they can access. It is not an access-control system: blocked URLs can still appear in search results, and different crawlers may interpret syntax differently. Do not use it as a substitute for authentication, permission, rate limits, or contractual terms.
Where robots.txt belongs
The file belongs at the root of the exact host, protocol, and port you are checking, for example https://example.com/robots.txt. Rules are UTF-8 text, grouped by crawler identity, and paths are case-sensitive. A sitemap location should be a fully qualified URL. A file on www.example.com does not automatically govern example.com, another protocol, or another port.
Recommended Free Tools
#1 Best Overall
Robots rules and sitemaps have different jobs
Use robots.txt rules to prevent crawling and sitemaps to encourage discovery. A sitemap is not a command to crawl only the listed URLs; links, redirects, feeds, and application routes can expose additional URLs. Treat both as inputs to your policy, then apply your own allowlist, path filters, and request budget.
Template 1: a controlled Python URL and resource checker
This script starts with a host, reads sitemap declarations from robots.txt, supports sitemap indexes, checks selected URLs, follows redirects, and writes JSON Lines. Install its only dependency with python -m pip install requests.
import json
import re
import time
from datetime import datetime, timezone
from urllib.parse import urljoin, urlparse
import requests
START = "https://example.com/"
MAX_URLS = 100
DELAY_SECONDS = 0.5
TIMEOUT = 20
ALLOWED_PATH_PREFIXES = ("/",) # replace with approved prefixes
REQUIRED_TEXT = None # e.g. "pricing"
session = requests.Session()
session.headers.update({"User-Agent": "ResourceChecker/1.0 (contact: [email protected])"})
def now():
return datetime.now(timezone.utc).isoformat()
def same_origin(a, b):
pa, pb = urlparse(a), urlparse(b)
return (pa.scheme, pa.netloc) == (pb.scheme, pb.netloc)
def robots_and_sitemaps(root):
robots_url = urljoin(root, "/robots.txt")
r = session.get(robots_url, timeout=TIMEOUT)
sitemap_urls = []
if r.ok:
for line in r.text.splitlines():
if line.lower().startswith("sitemap:"):
sitemap_urls.append(line.split(":", 1)[1].strip())
return robots_url, r, sitemap_urls
def sitemap_urls(sitemap_url, seen=None):
seen = set() if seen is None else seen
if sitemap_url in seen or len(seen) > 50:
return []
seen.add(sitemap_url)
r = session.get(sitemap_url, timeout=TIMEOUT)
r.raise_for_status()
locs = re.findall(r"s*(.*?)s* ", r.text, re.I)
if "= MAX_URLS:
break
with open("resource-report.jsonl", "w", encoding="utf-8") as out:
for url in unique:
out.write(json.dumps(check(url), ensure_ascii=False) + "n")
out.flush()
time.sleep(DELAY_SECONDS)
Replace START, path prefixes, and the optional text check. For binary assets, avoid reading r.text; inspect status, content type, length, and (where appropriate) a bounded byte signature. Add an explicit maximum response size if you may encounter very large downloads.
Finding URLs when there is no sitemap declaration
A missing declaration does not prove that no sitemap exists. You can add an approved sitemap URL manually, parse links from pages you are allowed to fetch, or use a crawler framework. Keep discovery bounded: deduplicate URLs, restrict origins, normalize fragments, cap depth or count, and record the source that produced each URL.
Rank #2
Template 2: Scrapy SitemapSpider for larger or recurring checks
Scrapy’s SitemapSpider can discover sitemap URLs through robots.txt, process sitemap indexes, and route matching URL patterns to callbacks. Create a project with scrapy startproject resourcecheck, then add this spider under resourcecheck/spiders/site.py:
import scrapy
from scrapy.spiders import SitemapSpider
from datetime import datetime, timezone
class SiteSpider(SitemapSpider):
name = "site_resources"
allowed_domains = ["example.com"]
sitemap_urls = ["https://example.com/robots.txt"]
sitemap_rules = [
(r"/products/", "parse_product"),
(r".(pdf|zip)$", "parse_file"),
]
custom_settings = {
"DOWNLOAD_DELAY": 0.5,
"CONCURRENT_REQUESTS_PER_DOMAIN": 2,
"ROBOTSTXT_OBEY": True,
"FEEDS": {"report.jsonl": {"format": "jsonlines", "overwrite": True}},
}
def base(self, response):
return {
"requested_url": response.request.url,
"final_url": response.url,
"status": response.status,
"content_type": response.headers.get(b"Content-Type", b"").decode("latin-1"),
"checked_at": datetime.now(timezone.utc).isoformat(),
}
def parse_product(self, response):
row = self.base(response)
row["title"] = response.css("title::text").get()
row["has_canonical"] = bool(response.css('link[rel="canonical"]::attr(href)').get())
row["content_check"] = bool(response.css("main").get())
yield row
def parse_file(self, response):
row = self.base(response)
row["content_check"] = response.headers.get(b"Content-Type", b"").startswith(b"application/")
yield row
Run it with scrapy crawl site_resources. Adjust allowed_domains, sitemap rules, delay, concurrency, and extraction checks to your target. Scrapy exposes the response URL, status, headers, and body through the response object, making it practical to produce consistent reports and add retries, caching, pipelines, or scheduling.
How to check whether a website URL is working
For each request, retain at least these columns:
| Field | Purpose |
|---|---|
| requested_url | What discovery produced or what the operator supplied. |
| final_url | Where redirects ended; compare host and path with your policy. |
| status | HTTP result, including redirects, client errors, and server errors. |
| selected headers | Content-Type, Content-Length, Last-Modified, Cache-Control, and relevant authentication or rate-limit signals. |
| checked_at | UTC timestamp for auditability. |
| content_check | A requirement-specific boolean or explanation, never a synonym for status. |
| error | Timeout, DNS, TLS, parsing, or connection failure details. |
Use HEAD only when the target reliably implements it; many sites return different behavior for HEAD. A bounded GET is usually more representative, especially when you must inspect HTML, JSON, or a file signature. Never log credentials or sensitive response bodies by default.
JavaScript, authentication, and permission boundaries
Requests and Scrapy fetch server responses; they do not automatically execute browser JavaScript. If the URL is an application shell whose data arrives through scripts, identify the permitted API or use a rendering-capable workflow. Rendering adds resource cost, timing variability, cookie state, and bot-detection failure modes.
Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchWindows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallAuthenticated pages require explicit authorization and carefully scoped cookies, headers, or tokens. Keep secrets outside source code, redact them from reports, and do not infer permission from a successful response. Respect applicable terms, privacy obligations, and rate limits for every target.
Troubleshooting common failures
robots.txt returns 404, 403, or HTML
Record the status and content type, verify the exact origin and port, and treat ambiguous content conservatively. A browser-accessible file that is not valid plain text should not be silently interpreted as rules.
Sitemap XML will not parse
Check for an index versus a URL set, namespaces, compressed content, redirects, and an HTML error page returned with status 200. Use an XML parser for production rather than relying only on regular expressions, and cap nested sitemap count.
Everything is 200 but the report says missing
The server may return a soft-404, login page, challenge page, or JavaScript shell. Compare final URL, content type, title, body markers, and expected fields; define a task-specific validation rule.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Timeouts, 429, or intermittent 5xx responses
Lower concurrency, add bounded exponential backoff, honor Retry-After, cache completed results, and stop after a clear retry budget. Separate transient failures from permanent authorization or not-found results.
Redirects leave the approved host
Record the final URL and reject or quarantine off-origin destinations unless your policy explicitly allows them. This prevents an apparently safe sitemap entry from becoming an uncontrolled crawl.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Choosing a simple script or Scrapy
| Need | Better starting point | Reason |
|---|---|---|
| A few known URLs and a one-off report | Requests script | Small dependency surface and direct control. |
| Sitemap indexes, recurring jobs, many callbacks, or pipelines | Scrapy SitemapSpider | Built-in discovery, routing, concurrency controls, and response handling. |
| Client-rendered content | Rendering-capable browser workflow | Plain HTTP clients cannot see data created only in the browser. |
| Strict audit output | Either, with a defined schema | Output requirements matter more than library choice. |
There is no universally best library. Scale, page behavior, JavaScript needs, and output requirements determine the appropriate implementation; the available guidance does not establish comparative speed benchmarks.
Or skip the browser setup
ScreenshotNeo provides a website screenshot API and MCP server when the resource you need to verify is visual. A single request can return PNG, JPEG, WebP, or PDF. Before capture it accepts the cookie or consent banner as a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each cleanup step can be turned off. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing status.
Free tools Windows power users keep installed
One-click scans. No signup required.
Using the API documented at https://screenshotneo.com/docs/:
Best Value
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
ScreenshotNeo also offers an MCP server with take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients. Plans include 1,000 screenshots per month free without a card; paid plans start at $5 for 3,000 shots. Features include full-page and element capture, device presets, custom CSS and JavaScript, waits, request blocking, headers and cookies, PDF controls, caching, signed links, asynchronous webhooks, bulk capture, and a usage API.
Create a free ScreenshotNeo account to use the 1,000-shot monthly allowance without entering a card.
Frequently Asked Questions
How do I find all URLs on a website?
Start with robots.txt sitemap declarations, expand sitemap indexes, then add only explicitly approved link or feed discovery with origin, path, count, and depth limits.
Quick wins for a faster PC:
Scan for outdated or missing drivers - takes under a minuteDriver Scan →Repair Windows errors before they cause bigger problemsFix Now →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →How do I check a sitemap with Python?
Fetch the XML, distinguish a sitemap index from a URL set, extract each loc, recurse through indexes with a seen set, and report parse or HTTP errors separately.
Does a sitemap guarantee that Google will crawl every listed URL?
No. It encourages discovery; it does not constrain Google to crawl only those URLs or guarantee indexing.
The Bottom Line
A reliable resource checker is an allowlisted, rate-limited pipeline: discover from robots.txt and sitemaps, request carefully, preserve redirect and response metadata, and validate the actual content requirement separately from HTTP status.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

