October DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsPC HealthRecommendedCrashes, freezes, slowdowns? Check your PC nowSpot repairable issues before they interrupt work.Check PCOctober DealsAmazon USDeal season is back - check today's better picksAmazon US: current deals, useful picks and tech finds.See Picks×
Skip to content
World desk9 min

How to Scrape Sitemaps to Discover Scraping Targets (Safely and Reliably)

A practical, safety-conscious guide to discovering URLs from robots.txt and sitemap indexes, parsing them correctly, and validating candidates before crawling.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use a sitemap as a discovery feed, not as a permission slip or a guarantee that every URL works. Start with /robots.txt, follow each Sitemap: declaration, identify whether each XML document is a URL set or a sitemap index, recursively collect namespace-aware <loc> values, then normalize, deduplicate and validate those URLs before any page crawl. A listed URL may be stale, redirected, blocked, unavailable or unsuitable for your job.

What “scraping a sitemap” actually does

A sitemap is an XML document that helps search engines discover URLs. Extracting its <loc> elements gives you candidate targets for a later crawl; it does not prove that a URL is live, crawlable, canonical, public, or authorized for your use. Google’s sitemap guidance explicitly treats sitemaps as discovery aids rather than guarantees of crawling or indexing.

Keep these stages separate:

  • Discovery: obtain sitemap files and extract URL strings.
  • Validation: check DNS, HTTP responses, redirects, content type and canonical signals.
  • Policy: review robots rules, terms, applicable law, authentication requirements and the site owner’s expectations.
  • Fetching: request pages at a controlled rate and stop when errors or restrictions indicate you should.

Step 1: Find the sitemap through robots.txt

Request the origin’s robots file first:

curl -iL --max-time 30 https://example.com/robots.txt

Look for one or more case-insensitive lines beginning with Sitemap:. A declaration normally contains an absolute URL, and a site can publish several declarations for language, product, image or news collections. Do not assume the first one is the only one.

Robots.txt is also where crawler software commonly discovers sitemap locations. If no declaration exists, try a small, documented fallback list such as /sitemap.xml, /sitemap_index.xml and /sitemap-index.xml. There is no universal filename-discovery guarantee, so treat guesses as optional probes rather than an exhaustive search.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Step 2: Fetch and classify the XML

Two root elements matter:

  • <urlset> contains URL records. Each record normally has a <loc>, and may include <lastmod>, <changefreq> or <priority>.
  • <sitemapindex> contains child sitemap records. Each child has a <loc> pointing to another XML file, which can itself be an index.

The common protocol namespace is http://www.sitemaps.org/schemas/sitemap/0.9. XML namespaces are significant: searching for an unqualified tag can return zero results even when the document is valid. XML entities must also be decoded by a real XML parser, not by regular-expression substitution.

Compression is normal. A server may return a gzip-encoded body or expose a file ending in .gz. Send an Accept-Encoding header and let your HTTP client decompress the response; parse the resulting XML bytes.

Step 3: Recursively traverse indexes and collect URLs

The following Python program starts at robots.txt, follows declared sitemaps, handles nested indexes, accepts gzip responses, enforces a sitemap limit, and writes deduplicated candidates. It deliberately does not fetch each page.

from collections import deque
from urllib.parse import urljoin, urlparse
import gzip
import requests
import xml.etree.ElementTree as ET

UA = "SitemapTargetDiscovery/1.0 ([email protected])"
NS = {"sm": "http://www.sitemaps.org/schemas/sitemap/0.9"}
TIMEOUT = 30
MAX_SITEMAPS = 5000


def get(url):
    r = requests.get(url, headers={"User-Agent": UA,
                                    "Accept-Encoding": "gzip, deflate"},
                     timeout=TIMEOUT)
    r.raise_for_status()
    data = r.content
    if url.lower().endswith(".gz") or r.headers.get("Content-Encoding", "").lower() == "gzip":
        try:
            data = gzip.decompress(data)
        except OSError:
            pass  # requests may already have decompressed it
    return r, data


def robots_sitemaps(origin):
    r = requests.get(urljoin(origin, "/robots.txt"),
                     headers={"User-Agent": UA}, timeout=TIMEOUT)
    if not r.ok:
        return []
    found = []
    for line in r.text.splitlines():
        key, sep, value = line.partition(":")
        if sep and key.strip().lower() == "sitemap":
            value = value.strip()
            if value:
                found.append(value)
    return found


def discover(origin):
    queue = deque(robots_sitemaps(origin))
    if not queue:
        queue.extend(urljoin(origin, p) for p in
                     ("/sitemap.xml", "/sitemap_index.xml", "/sitemap-index.xml"))

    seen_maps, urls = set(), set()
    while queue:
        sitemap = queue.popleft()
        if sitemap in seen_maps:
            continue
        if len(seen_maps) >= MAX_SITEMAPS:
            raise RuntimeError("sitemap limit reached; review the site before continuing")
        seen_maps.add(sitemap)
        try:
            _, body = get(sitemap)
            root = ET.fromstring(body)
        except (requests.RequestException, ET.ParseError) as exc:
            print(f"skip {sitemap}: {exc}")
            continue

        tag = root.tag.rsplit("}", 1)[-1]
        if tag == "sitemapindex":
            for node in root.findall("sm:sitemap/sm:loc", NS):
                if node.text:
                    queue.append(node.text.strip())
        elif tag == "urlset":
            for node in root.findall("sm:url/sm:loc", NS):
                if node.text:
                    urls.add(node.text.strip())
        else:
            print(f"skip {sitemap}: root element is {tag!r}")
    return sorted(urls)


if __name__ == "__main__":
    targets = discover("https://example.com")
    with open("targets.txt", "w", encoding="utf-8") as f:
        f.write("n".join(targets) + ("n" if targets else ""))
    print(f"discovered {len(targets)} candidate URLs")

Replace the origin and contact address. For production, persist a queue, log status codes and parse errors, and make retries bounded. A malformed child sitemap should be recorded and skipped rather than terminating an otherwise useful run.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Step 4: Normalize and deduplicate without changing meaning

Use the URL parser, not ad-hoc string surgery. At minimum:

  • Require an http or https scheme and a hostname.
  • Decode XML entities through the parser.
  • Resolve relative values only when your policy explicitly permits them; sitemap guidance recommends fully qualified absolute URLs.
  • Remove exact duplicates while preserving the first-seen value for auditability.
  • Decide deliberately whether fragments should be removed. Fragments are not sent in ordinary HTTP requests, so they usually do not represent separate server resources.
  • Do not automatically lowercase paths, sort query parameters, strip trailing slashes or discard parameters. Those transformations can change resource identity.
  • Optionally restrict hosts to the site you intend to examine, and record out-of-scope URLs instead of silently dropping them.

Keep the original sitemap URL and extraction time beside every candidate. That provenance helps explain later why a target was included.

Step 5: Validate candidates before a crawl

Discovery is not availability. Perform a lightweight validation pass with a strict rate limit:

  1. Resolve the hostname and reject domains outside your allowed scope.
  2. Request the URL with a short timeout. Prefer a normal GET when servers mishandle HEAD; if you use HEAD, fall back to GET on a 405 or misleading response.
  3. Record status, final URL, redirect chain, content type, content length and response time.
  4. Apply your robots and policy decision before downloading page content.
  5. Follow redirects only within an approved host policy, and re-check the final URL.
  6. Classify failures (DNS, timeout, TLS, 4xx, 5xx, bot check, non-HTML) and schedule bounded retries for transient errors.

A 200 response is not proof that the page is canonical or useful. Conversely, a temporary 503 does not prove that the sitemap entry is permanently dead. Store the observation time and revalidate later when freshness matters.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Limits and fields you can trust

Google documents a limit of 50 MB uncompressed or 50,000 URLs per sitemap. A sitemap index can list up to 50,000 sitemap locations. Large sites therefore split files and require traversal. Enforce your own memory, depth and total-request limits as well.

lastmod can help prioritize recrawls when it is consistently accurate. Treat it as publisher-supplied metadata, not proof that content changed. Google says it ignores priority and changefreq; do not build scheduling logic that assumes those fields control search crawling.

Custom parser or crawler framework?

Concern Custom parser Crawler framework with sitemap support
Robots discovery You implement parsing and policy checks. Built-in support may discover declarations automatically.
Nested indexes Explicit queue and recursion are required. Often handled by the sitemap component; verify current behavior.
Namespaces and gzip You control exact XML and decompression logic. Convenient, but confirm supported formats and limits.
Filtering and output Tailored host, path and database rules. Integrates with item pipelines, throttling and retries.
Maintenance Small dependency footprint, more edge cases to own. More features, but APIs and defaults change.

Scrapy’s SitemapSpider documentation describes sitemap discovery from robots.txt and nested sitemap handling, but the cited documentation is for an old release. Check the current Scrapy documentation and APIs before copying settings or method names into a new project.

Operational safety and crawl etiquette

  • Use an identifiable User-Agent and a contact address.
  • Throttle concurrency per host, add jitter, and honor explicit crawl restrictions.
  • Cache sitemap files and use conditional requests where supported; do not repeatedly download unchanged indexes.
  • Set maximum response sizes and XML entity protections. Never enable dangerous external-entity resolution for untrusted XML.
  • Separate discovery credentials from page-fetch credentials, and never expose cookies or Authorization headers in logs.
  • Stop on repeated blocks, authentication prompts or legal complaints.

Common failures and fixes

Robots.txt returns HTML or a redirect

Some hosts redirect HTTP to HTTPS or serve an error page with status 200. Follow redirects, verify the body resembles robots syntax, and continue only when you have a trustworthy sitemap URL.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

“Found zero URLs”

Check the root element and namespace. A namespaced urlset requires namespace-aware queries. Also confirm that you did not mistake a sitemap index for a URL set.

XML parse errors

Save a bounded sample of the response, inspect the status and content type, and check for an HTML block page, truncated transfer or invalid encoding. Retry once for transient transport errors; do not repeatedly hammer a broken endpoint.

Gzip errors

HTTP libraries frequently decompress automatically. Only decompress when the body is still gzip data; otherwise a second decompression raises an error.

Thousands of duplicates

Different files may list the same URL, and URL variants may differ only by a fragment or tracking parameter. Deduplicate exact strings first, then apply a documented canonicalization policy appropriate to your target.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Every target redirects or fails

Sitemaps can be stale or can contain alternate hosts, deleted pages and temporary failures. Preserve the evidence, classify the result and avoid treating the sitemap as a live inventory.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup

If your next step is taking rendered screenshots rather than downloading HTML, ScreenshotNeo turns a URL into a PNG, JPEG, WebP or PDF with one request. Its capture pipeline accepts cookie and consent banners and removes more than 60 known consent platforms, newsletter popups and chat widgets before the shot; each step can be turned off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads and cache hits are not billed, and the response identifies the result with X-Page-Verdict and X-Billed headers. An MCP server provides take_screenshot, get_page_info and capture_pdf tools for Claude, Cursor and other MCP clients.

Use the API documentation at https://screenshotneo.com/docs/ for authentication and options. A minimal call is:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`HTTP ${res.status}`);
const buffer = Buffer.from(await res.arrayBuffer());
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', buffer));

ScreenshotNeo includes full-page and element capture, device presets, custom viewport and retina scale, PDF controls, HTML/CSS rendering, custom JavaScript and CSS, clicks, waits, request blocking, headers, cookies, user agent, timezone, geolocation, transparent backgrounds, resizing, chosen cache TTLs, signed links, asynchronous webhooks, bulk capture of 100 URLs per call, a usage API and an OpenAPI specification. Parameter names used by other screenshot APIs also work, easing migration.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Plan Allowance Price
Free 1,000 shots/month $0, no card
Starter 3,000 shots $5
Growth 15,000 shots $15
Pro 60,000 shots $39
Scale 250,000 shots $99
Business 1,000,000 shots $249

Yearly billing gives two months free, and every feature is on every plan. Start with 1,000 free screenshots a month, with no card required.

A practical output checklist

  • Save every sitemap URL visited and its fetch result.
  • Record whether each XML file was an index or URL set.
  • Store raw and normalized loc values.
  • Track duplicates, redirects, status codes and validation timestamps.
  • Keep discovery separate from authorization and page fetching.
  • Set request, byte, recursion and retry budgets before running at scale.

Frequently Asked Questions

Can I scrape a sitemap that is not listed in robots.txt?

You can try a known, publicly reachable sitemap URL, but its existence does not establish permission to crawl the pages it lists. Apply the site’s rules, terms and applicable law before fetching targets.

Should I use HEAD requests to test every URL?

Not always. Some servers reject or mishandle HEAD. Use it only when appropriate, and fall back to a bounded GET while recording the method and response.

Does a sitemap contain only canonical, indexable pages?

No. It can be incomplete, stale, redirected or contain URLs that fail. Validate responses and canonical signals independently.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

More from the Wire

  1. Shenzhen desk3 min
    HONOR Expands Beyond Smartphones With Humanoid Robot RevealHONOR said it unveiled its first humanoid robot at MWC 2026 and named shopping assistance, workplace inspections, and supportive companionship as intended uses. Later Robotics D1 claims and a reported…
  2. Cupertino desk5 min
    Apple Unveils AirPods Max 2: The Upgrade That Should Have Happened Years AgoAirPods Max 2 adds H2-powered audio features and Apple claims up to 1.5× more effective ANC, but its design, Smart Case, and 20-hour battery rating are unchanged. Wired lossless audio…
  3. Cupertino desk4 min
    Apple’s OLED Touch MacBooks Are Coming—but the Dynamic Island Is the Real GambleApple has not announced an OLED touchscreen MacBook, but reports point to high-end models arriving in late 2026 or early 2027. The reported Mac Dynamic Island could be useful, but…
Recommended PC Tool
Recommended PC Tool
PC Slower Than It Used to Be?Free scan - under a minute
Outdated Drivers Are Slowing You DownFree scan - exact matches

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.