Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use a discovery-first checker: read the target host’s robots.txt, collect sitemap URLs, request only approved URL patterns at a controlled rate, and write a report containing the requested URL, final URL, status, headers, timestamp, and a task-specific content test. The templates below provide a small Python implementation and a scalable Scrapy version. They check whether resources respond and whether the response contains what your project needs; they do not turn crawler guidance into authorization.

What a website-resource checker should do

A useful checker separates four jobs:

  1. Inputs: a starting host or approved URL list, resource types or path patterns, request limits, and an output format such as CSV or JSON Lines.
  2. Discovery: inspect the host’s root robots.txt and sitemap references, then expand sitemap indexes if necessary.
  3. Request: fetch only relevant URLs, follow redirects according to your policy, and retain the final URL plus response metadata.
  4. Report: record the requested URL, final response URL, HTTP status, selected headers, timestamp, and a content check specific to your requirement.

“HTTP 200” means that a server returned a successful response. It does not prove that the page contains a product, API field, canonical tag, image, download, or other resource you require. Keep transport status and content validation as separate fields.

Can I use robots.txt to tell a scraper what not to crawl?

Yes, as crawler guidance. Google Search Central defines it as a file that tells search-engine crawlers which URLs they can access. It is not an access-control system: blocked URLs can still appear in search results, and different crawlers may interpret syntax differently. Do not use it as a substitute for authentication, permission, rate limits, or contractual terms.

Where robots.txt belongs

The file belongs at the root of the exact host, protocol, and port you are checking, for example https://example.com/robots.txt. Rules are UTF-8 text, grouped by crawler identity, and paths are case-sensitive. A sitemap location should be a fully qualified URL. A file on www.example.com does not automatically govern example.com, another protocol, or another port.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Robots rules and sitemaps have different jobs

Use robots.txt rules to prevent crawling and sitemaps to encourage discovery. A sitemap is not a command to crawl only the listed URLs; links, redirects, feeds, and application routes can expose additional URLs. Treat both as inputs to your policy, then apply your own allowlist, path filters, and request budget.

Template 1: a controlled Python URL and resource checker

This script starts with a host, reads sitemap declarations from robots.txt, supports sitemap indexes, checks selected URLs, follows redirects, and writes JSON Lines. Install its only dependency with python -m pip install requests.

import json
import re
import time
from datetime import datetime, timezone
from urllib.parse import urljoin, urlparse

import requests

START = "https://example.com/"
MAX_URLS = 100
DELAY_SECONDS = 0.5
TIMEOUT = 20
ALLOWED_PATH_PREFIXES = ("/",)       # replace with approved prefixes
REQUIRED_TEXT = None                  # e.g. "pricing"

session = requests.Session()
session.headers.update({"User-Agent": "ResourceChecker/1.0 (contact: [email protected])"})

def now():
    return datetime.now(timezone.utc).isoformat()

def same_origin(a, b):
    pa, pb = urlparse(a), urlparse(b)
    return (pa.scheme, pa.netloc) == (pb.scheme, pb.netloc)

def robots_and_sitemaps(root):
    robots_url = urljoin(root, "/robots.txt")
    r = session.get(robots_url, timeout=TIMEOUT)
    sitemap_urls = []
    if r.ok:
        for line in r.text.splitlines():
            if line.lower().startswith("sitemap:"):
                sitemap_urls.append(line.split(":", 1)[1].strip())
    return robots_url, r, sitemap_urls

def sitemap_urls(sitemap_url, seen=None):
    seen = set() if seen is None else seen
    if sitemap_url in seen or len(seen) > 50:
        return []
    seen.add(sitemap_url)
    r = session.get(sitemap_url, timeout=TIMEOUT)
    r.raise_for_status()
    locs = re.findall(r"s*(.*?)s*", r.text, re.I)
    if "= MAX_URLS:
        break

with open("resource-report.jsonl", "w", encoding="utf-8") as out:
    for url in unique:
        out.write(json.dumps(check(url), ensure_ascii=False) + "n")
        out.flush()
        time.sleep(DELAY_SECONDS)

Replace START, path prefixes, and the optional text check. For binary assets, avoid reading r.text; inspect status, content type, length, and (where appropriate) a bounded byte signature. Add an explicit maximum response size if you may encounter very large downloads.

Finding URLs when there is no sitemap declaration

A missing declaration does not prove that no sitemap exists. You can add an approved sitemap URL manually, parse links from pages you are allowed to fetch, or use a crawler framework. Keep discovery bounded: deduplicate URLs, restrict origins, normalize fragments, cap depth or count, and record the source that produced each URL.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Template 2: Scrapy SitemapSpider for larger or recurring checks

Scrapy’s SitemapSpider can discover sitemap URLs through robots.txt, process sitemap indexes, and route matching URL patterns to callbacks. Create a project with scrapy startproject resourcecheck, then add this spider under resourcecheck/spiders/site.py:

import scrapy
from scrapy.spiders import SitemapSpider
from datetime import datetime, timezone

class SiteSpider(SitemapSpider):
    name = "site_resources"
    allowed_domains = ["example.com"]
    sitemap_urls = ["https://example.com/robots.txt"]
    sitemap_rules = [
        (r"/products/", "parse_product"),
        (r".(pdf|zip)$", "parse_file"),
    ]

    custom_settings = {
        "DOWNLOAD_DELAY": 0.5,
        "CONCURRENT_REQUESTS_PER_DOMAIN": 2,
        "ROBOTSTXT_OBEY": True,
        "FEEDS": {"report.jsonl": {"format": "jsonlines", "overwrite": True}},
    }

    def base(self, response):
        return {
            "requested_url": response.request.url,
            "final_url": response.url,
            "status": response.status,
            "content_type": response.headers.get(b"Content-Type", b"").decode("latin-1"),
            "checked_at": datetime.now(timezone.utc).isoformat(),
        }

    def parse_product(self, response):
        row = self.base(response)
        row["title"] = response.css("title::text").get()
        row["has_canonical"] = bool(response.css('link[rel="canonical"]::attr(href)').get())
        row["content_check"] = bool(response.css("main").get())
        yield row

    def parse_file(self, response):
        row = self.base(response)
        row["content_check"] = response.headers.get(b"Content-Type", b"").startswith(b"application/")
        yield row

Run it with scrapy crawl site_resources. Adjust allowed_domains, sitemap rules, delay, concurrency, and extraction checks to your target. Scrapy exposes the response URL, status, headers, and body through the response object, making it practical to produce consistent reports and add retries, caching, pipelines, or scheduling.

How to check whether a website URL is working

For each request, retain at least these columns:

Field Purpose
requested_url What discovery produced or what the operator supplied.
final_url Where redirects ended; compare host and path with your policy.
status HTTP result, including redirects, client errors, and server errors.
selected headers Content-Type, Content-Length, Last-Modified, Cache-Control, and relevant authentication or rate-limit signals.
checked_at UTC timestamp for auditability.
content_check A requirement-specific boolean or explanation, never a synonym for status.
error Timeout, DNS, TLS, parsing, or connection failure details.

Use HEAD only when the target reliably implements it; many sites return different behavior for HEAD. A bounded GET is usually more representative, especially when you must inspect HTML, JSON, or a file signature. Never log credentials or sensitive response bodies by default.

JavaScript, authentication, and permission boundaries

Requests and Scrapy fetch server responses; they do not automatically execute browser JavaScript. If the URL is an application shell whose data arrives through scripts, identify the permitted API or use a rendering-capable workflow. Rendering adds resource cost, timing variability, cookie state, and bot-detection failure modes.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Authenticated pages require explicit authorization and carefully scoped cookies, headers, or tokens. Keep secrets outside source code, redact them from reports, and do not infer permission from a successful response. Respect applicable terms, privacy obligations, and rate limits for every target.

Troubleshooting common failures

robots.txt returns 404, 403, or HTML

Record the status and content type, verify the exact origin and port, and treat ambiguous content conservatively. A browser-accessible file that is not valid plain text should not be silently interpreted as rules.

Sitemap XML will not parse

Check for an index versus a URL set, namespaces, compressed content, redirects, and an HTML error page returned with status 200. Use an XML parser for production rather than relying only on regular expressions, and cap nested sitemap count.

Everything is 200 but the report says missing

The server may return a soft-404, login page, challenge page, or JavaScript shell. Compare final URL, content type, title, body markers, and expected fields; define a task-specific validation rule.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Timeouts, 429, or intermittent 5xx responses

Lower concurrency, add bounded exponential backoff, honor Retry-After, cache completed results, and stop after a clear retry budget. Separate transient failures from permanent authorization or not-found results.

Redirects leave the approved host

Record the final URL and reject or quarantine off-origin destinations unless your policy explicitly allows them. This prevents an apparently safe sitemap entry from becoming an uncontrolled crawl.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Choosing a simple script or Scrapy

Need Better starting point Reason
A few known URLs and a one-off report Requests script Small dependency surface and direct control.
Sitemap indexes, recurring jobs, many callbacks, or pipelines Scrapy SitemapSpider Built-in discovery, routing, concurrency controls, and response handling.
Client-rendered content Rendering-capable browser workflow Plain HTTP clients cannot see data created only in the browser.
Strict audit output Either, with a defined schema Output requirements matter more than library choice.

There is no universally best library. Scale, page behavior, JavaScript needs, and output requirements determine the appropriate implementation; the available guidance does not establish comparative speed benchmarks.

Or skip the browser setup

ScreenshotNeo provides a website screenshot API and MCP server when the resource you need to verify is visual. A single request can return PNG, JPEG, WebP, or PDF. Before capture it accepts the cookie or consent banner as a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each cleanup step can be turned off. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing status.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Using the API documented at https://screenshotneo.com/docs/:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

ScreenshotNeo also offers an MCP server with take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients. Plans include 1,000 screenshots per month free without a card; paid plans start at $5 for 3,000 shots. Features include full-page and element capture, device presets, custom CSS and JavaScript, waits, request blocking, headers and cookies, PDF controls, caching, signed links, asynchronous webhooks, bulk capture, and a usage API.

Create a free ScreenshotNeo account to use the 1,000-shot monthly allowance without entering a card.

Frequently Asked Questions

How do I find all URLs on a website?

Start with robots.txt sitemap declarations, expand sitemap indexes, then add only explicitly approved link or feed discovery with origin, path, count, and depth limits.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How do I check a sitemap with Python?

Fetch the XML, distinguish a sitemap index from a URL set, extract each loc, recurse through indexes with a seen set, and report parse or HTTP errors separately.

Does a sitemap guarantee that Google will crawl every listed URL?

No. It encourages discovery; it does not constrain Google to crawl only those URLs or guarantee indexing.

The Bottom Line

A reliable resource checker is an allowlisted, rate-limited pipeline: discover from robots.txt and sitemaps, request carefully, preserve redirect and response metadata, and validate the actual content requirement separately from HTTP status.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.