What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Use a sitemap as a discovery feed, not as a permission slip or a guarantee that every URL works. Start with /robots.txt, follow each Sitemap: declaration, identify whether each XML document is a URL set or a sitemap index, recursively collect namespace-aware <loc> values, then normalize, deduplicate and validate those URLs before any page crawl. A listed URL may be stale, redirected, blocked, unavailable or unsuitable for your job.
What “scraping a sitemap” actually does
A sitemap is an XML document that helps search engines discover URLs. Extracting its <loc> elements gives you candidate targets for a later crawl; it does not prove that a URL is live, crawlable, canonical, public, or authorized for your use. Google’s sitemap guidance explicitly treats sitemaps as discovery aids rather than guarantees of crawling or indexing.
Keep these stages separate:
- Discovery: obtain sitemap files and extract URL strings.
- Validation: check DNS, HTTP responses, redirects, content type and canonical signals.
- Policy: review robots rules, terms, applicable law, authentication requirements and the site owner’s expectations.
- Fetching: request pages at a controlled rate and stop when errors or restrictions indicate you should.
Step 1: Find the sitemap through robots.txt
Request the origin’s robots file first:
curl -iL --max-time 30 https://example.com/robots.txt
Look for one or more case-insensitive lines beginning with Sitemap:. A declaration normally contains an absolute URL, and a site can publish several declarations for language, product, image or news collections. Do not assume the first one is the only one.
Robots.txt is also where crawler software commonly discovers sitemap locations. If no declaration exists, try a small, documented fallback list such as /sitemap.xml, /sitemap_index.xml and /sitemap-index.xml. There is no universal filename-discovery guarantee, so treat guesses as optional probes rather than an exhaustive search.
#1 Best Overall
Step 2: Fetch and classify the XML
Two root elements matter:
<urlset>contains URL records. Each record normally has a<loc>, and may include<lastmod>,<changefreq>or<priority>.<sitemapindex>contains child sitemap records. Each child has a<loc>pointing to another XML file, which can itself be an index.
The common protocol namespace is http://www.sitemaps.org/schemas/sitemap/0.9. XML namespaces are significant: searching for an unqualified tag can return zero results even when the document is valid. XML entities must also be decoded by a real XML parser, not by regular-expression substitution.
Compression is normal. A server may return a gzip-encoded body or expose a file ending in .gz. Send an Accept-Encoding header and let your HTTP client decompress the response; parse the resulting XML bytes.
Step 3: Recursively traverse indexes and collect URLs
The following Python program starts at robots.txt, follows declared sitemaps, handles nested indexes, accepts gzip responses, enforces a sitemap limit, and writes deduplicated candidates. It deliberately does not fetch each page.
from collections import deque
from urllib.parse import urljoin, urlparse
import gzip
import requests
import xml.etree.ElementTree as ET
UA = "SitemapTargetDiscovery/1.0 ([email protected])"
NS = {"sm": "http://www.sitemaps.org/schemas/sitemap/0.9"}
TIMEOUT = 30
MAX_SITEMAPS = 5000
def get(url):
r = requests.get(url, headers={"User-Agent": UA,
"Accept-Encoding": "gzip, deflate"},
timeout=TIMEOUT)
r.raise_for_status()
data = r.content
if url.lower().endswith(".gz") or r.headers.get("Content-Encoding", "").lower() == "gzip":
try:
data = gzip.decompress(data)
except OSError:
pass # requests may already have decompressed it
return r, data
def robots_sitemaps(origin):
r = requests.get(urljoin(origin, "/robots.txt"),
headers={"User-Agent": UA}, timeout=TIMEOUT)
if not r.ok:
return []
found = []
for line in r.text.splitlines():
key, sep, value = line.partition(":")
if sep and key.strip().lower() == "sitemap":
value = value.strip()
if value:
found.append(value)
return found
def discover(origin):
queue = deque(robots_sitemaps(origin))
if not queue:
queue.extend(urljoin(origin, p) for p in
("/sitemap.xml", "/sitemap_index.xml", "/sitemap-index.xml"))
seen_maps, urls = set(), set()
while queue:
sitemap = queue.popleft()
if sitemap in seen_maps:
continue
if len(seen_maps) >= MAX_SITEMAPS:
raise RuntimeError("sitemap limit reached; review the site before continuing")
seen_maps.add(sitemap)
try:
_, body = get(sitemap)
root = ET.fromstring(body)
except (requests.RequestException, ET.ParseError) as exc:
print(f"skip {sitemap}: {exc}")
continue
tag = root.tag.rsplit("}", 1)[-1]
if tag == "sitemapindex":
for node in root.findall("sm:sitemap/sm:loc", NS):
if node.text:
queue.append(node.text.strip())
elif tag == "urlset":
for node in root.findall("sm:url/sm:loc", NS):
if node.text:
urls.add(node.text.strip())
else:
print(f"skip {sitemap}: root element is {tag!r}")
return sorted(urls)
if __name__ == "__main__":
targets = discover("https://example.com")
with open("targets.txt", "w", encoding="utf-8") as f:
f.write("n".join(targets) + ("n" if targets else ""))
print(f"discovered {len(targets)} candidate URLs")
Replace the origin and contact address. For production, persist a queue, log status codes and parse errors, and make retries bounded. A malformed child sitemap should be recorded and skipped rather than terminating an otherwise useful run.
Windows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallCrashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minuteStep 4: Normalize and deduplicate without changing meaning
Use the URL parser, not ad-hoc string surgery. At minimum:
- Require an
httporhttpsscheme and a hostname. - Decode XML entities through the parser.
- Resolve relative values only when your policy explicitly permits them; sitemap guidance recommends fully qualified absolute URLs.
- Remove exact duplicates while preserving the first-seen value for auditability.
- Decide deliberately whether fragments should be removed. Fragments are not sent in ordinary HTTP requests, so they usually do not represent separate server resources.
- Do not automatically lowercase paths, sort query parameters, strip trailing slashes or discard parameters. Those transformations can change resource identity.
- Optionally restrict hosts to the site you intend to examine, and record out-of-scope URLs instead of silently dropping them.
Keep the original sitemap URL and extraction time beside every candidate. That provenance helps explain later why a target was included.
Step 5: Validate candidates before a crawl
Discovery is not availability. Perform a lightweight validation pass with a strict rate limit:
- Resolve the hostname and reject domains outside your allowed scope.
- Request the URL with a short timeout. Prefer a normal
GETwhen servers mishandleHEAD; if you useHEAD, fall back toGETon a 405 or misleading response. - Record status, final URL, redirect chain, content type, content length and response time.
- Apply your robots and policy decision before downloading page content.
- Follow redirects only within an approved host policy, and re-check the final URL.
- Classify failures (DNS, timeout, TLS, 4xx, 5xx, bot check, non-HTML) and schedule bounded retries for transient errors.
A 200 response is not proof that the page is canonical or useful. Conversely, a temporary 503 does not prove that the sitemap entry is permanently dead. Store the observation time and revalidate later when freshness matters.
Rank #3
Limits and fields you can trust
Google documents a limit of 50 MB uncompressed or 50,000 URLs per sitemap. A sitemap index can list up to 50,000 sitemap locations. Large sites therefore split files and require traversal. Enforce your own memory, depth and total-request limits as well.
lastmod can help prioritize recrawls when it is consistently accurate. Treat it as publisher-supplied metadata, not proof that content changed. Google says it ignores priority and changefreq; do not build scheduling logic that assumes those fields control search crawling.
Custom parser or crawler framework?
| Concern | Custom parser | Crawler framework with sitemap support |
|---|---|---|
| Robots discovery | You implement parsing and policy checks. | Built-in support may discover declarations automatically. |
| Nested indexes | Explicit queue and recursion are required. | Often handled by the sitemap component; verify current behavior. |
| Namespaces and gzip | You control exact XML and decompression logic. | Convenient, but confirm supported formats and limits. |
| Filtering and output | Tailored host, path and database rules. | Integrates with item pipelines, throttling and retries. |
| Maintenance | Small dependency footprint, more edge cases to own. | More features, but APIs and defaults change. |
Scrapy’s SitemapSpider documentation describes sitemap discovery from robots.txt and nested sitemap handling, but the cited documentation is for an old release. Check the current Scrapy documentation and APIs before copying settings or method names into a new project.
Operational safety and crawl etiquette
- Use an identifiable User-Agent and a contact address.
- Throttle concurrency per host, add jitter, and honor explicit crawl restrictions.
- Cache sitemap files and use conditional requests where supported; do not repeatedly download unchanged indexes.
- Set maximum response sizes and XML entity protections. Never enable dangerous external-entity resolution for untrusted XML.
- Separate discovery credentials from page-fetch credentials, and never expose cookies or Authorization headers in logs.
- Stop on repeated blocks, authentication prompts or legal complaints.
Common failures and fixes
Robots.txt returns HTML or a redirect
Some hosts redirect HTTP to HTTPS or serve an error page with status 200. Follow redirects, verify the body resembles robots syntax, and continue only when you have a trustworthy sitemap URL.
“Found zero URLs”
Check the root element and namespace. A namespaced urlset requires namespace-aware queries. Also confirm that you did not mistake a sitemap index for a URL set.
XML parse errors
Save a bounded sample of the response, inspect the status and content type, and check for an HTML block page, truncated transfer or invalid encoding. Retry once for transient transport errors; do not repeatedly hammer a broken endpoint.
Gzip errors
HTTP libraries frequently decompress automatically. Only decompress when the body is still gzip data; otherwise a second decompression raises an error.
Thousands of duplicates
Different files may list the same URL, and URL variants may differ only by a fragment or tracking parameter. Deduplicate exact strings first, then apply a documented canonicalization policy appropriate to your target.
The Tool Desk
Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Best Value
Every target redirects or fails
Sitemaps can be stale or can contain alternate hosts, deleted pages and temporary failures. Preserve the evidence, classify the result and avoid treating the sitemap as a live inventory.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Or skip the browser setup
If your next step is taking rendered screenshots rather than downloading HTML, ScreenshotNeo turns a URL into a PNG, JPEG, WebP or PDF with one request. Its capture pipeline accepts cookie and consent banners and removes more than 60 known consent platforms, newsletter popups and chat widgets before the shot; each step can be turned off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads and cache hits are not billed, and the response identifies the result with X-Page-Verdict and X-Billed headers. An MCP server provides take_screenshot, get_page_info and capture_pdf tools for Claude, Cursor and other MCP clients.
Use the API documentation at https://screenshotneo.com/docs/ for authentication and options. A minimal call is:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python:
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`HTTP ${res.status}`);
const buffer = Buffer.from(await res.arrayBuffer());
await import('node:fs/promises').then(fs => fs.writeFile('shot.webp', buffer));
ScreenshotNeo includes full-page and element capture, device presets, custom viewport and retina scale, PDF controls, HTML/CSS rendering, custom JavaScript and CSS, clicks, waits, request blocking, headers, cookies, user agent, timezone, geolocation, transparent backgrounds, resizing, chosen cache TTLs, signed links, asynchronous webhooks, bulk capture of 100 URLs per call, a usage API and an OpenAPI specification. Parameter names used by other screenshot APIs also work, easing migration.
Quick wins for a faster PC:
Clear out junk files and repair common Windows errorsFree Scan →Scan for outdated or missing drivers - takes under a minuteDriver Scan →| Plan | Allowance | Price |
|---|---|---|
| Free | 1,000 shots/month | $0, no card |
| Starter | 3,000 shots | $5 |
| Growth | 15,000 shots | $15 |
| Pro | 60,000 shots | $39 |
| Scale | 250,000 shots | $99 |
| Business | 1,000,000 shots | $249 |
Yearly billing gives two months free, and every feature is on every plan. Start with 1,000 free screenshots a month, with no card required.
A practical output checklist
- Save every sitemap URL visited and its fetch result.
- Record whether each XML file was an index or URL set.
- Store raw and normalized
locvalues. - Track duplicates, redirects, status codes and validation timestamps.
- Keep discovery separate from authorization and page fetching.
- Set request, byte, recursion and retry budgets before running at scale.
Frequently Asked Questions
Can I scrape a sitemap that is not listed in robots.txt?
You can try a known, publicly reachable sitemap URL, but its existence does not establish permission to crawl the pages it lists. Apply the site’s rules, terms and applicable law before fetching targets.
Should I use HEAD requests to test every URL?
Not always. Some servers reject or mishandle HEAD. Use it only when appropriate, and fall back to a bounded GET while recording the method and response.
Does a sitemap contain only canonical, indexable pages?
No. It can be incomplete, stale, redirected or contain URLs that fail. Validate responses and canonical signals independently.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




