To extract every URL, download the sitemap, parse XML with the Sitemap namespace, and read each <url><loc> value. If the document is a sitemap index, read its <sitemap><loc> entries and recursively process each child file. The Python example below handles indexes, namespaces, compressed files, deduplication, limits, and common failure cases.
What a sitemap contains
The Sitemap protocol is an XML format (the protocol documentation says, “The Sitemap protocol format consists of XML tags.”). A normal sitemap has a <urlset> root and one <url> element per page. The page address is in <loc>; optional metadata includes <lastmod>, <changefreq>, and <priority>. A sitemap index instead has a <sitemapindex> root and links to other sitemap files through <sitemap><loc>.
Use fully qualified absolute URLs. Keep extracted addresses within the host and protocol scope permitted by the sitemap’s location. A sitemap is a URL declaration, not proof that a search engine indexed a page; <lastmod> is update metadata, not an indexing result.
Python: extract URLs from one sitemap or an index
Install the two dependencies first:
python -m pip install requests lxml
This script checks HTTP responses, parses namespace-qualified XML, follows indexes recursively, supports .xml.gz, prevents loops, enforces depth and URL budgets, and returns unique URLs.
Do these 3 things before closing this tab:
1Clear out junk files and repair common Windows errors2Scan for outdated or missing drivers - takes under a minute3Repair Windows errors before they cause bigger problems#1 Best Overall
from __future__ import annotations
from gzip import decompress
from io import BytesIO
from typing import Iterable
from urllib.parse import urlparse
import requests
from lxml import etree
SITEMAP_NS = "http://www.sitemaps.org/schemas/sitemap/0.9"
NS = {"sm": SITEMAP_NS}
def _xml_root(content: bytes, source_url: str) -> etree._Element:
# XMLParser does not resolve external entities, avoiding XXE-style fetches.
parser = etree.XMLParser(
resolve_entities=False,
no_network=True,
load_dtd=False,
recover=False,
huge_tree=False,
)
try:
return etree.fromstring(content, parser=parser)
except etree.XMLSyntaxError:
if source_url.lower().split("?", 1)[0].endswith(".gz"):
return etree.fromstring(decompress(content), parser=parser)
raise
def extract_urls(
sitemap_url: str,
*,
max_depth: int = 10,
max_urls: int = 1_000_000,
timeout: float = 30,
) -> list[str]:
session = requests.Session()
visited: set[str] = set()
found: list[str] = []
seen_urls: set[str] = set()
def walk(url: str, depth: int) -> None:
if depth > max_depth:
raise ValueError(f"Sitemap index depth exceeds {max_depth}: {url}")
if url in visited:
return
visited.add(url)
response = session.get(
url,
timeout=timeout,
headers={"Accept": "application/xml,text/xml;q=0.9,*/*;q=0.1"},
)
response.raise_for_status()
root = _xml_root(response.content, url)
root_name = etree.QName(root).localname
if root_name == "sitemapindex":
children = root.xpath("/sm:sitemapindex/sm:sitemap/sm:loc/text()", namespaces=NS)
for child in children:
walk(child.strip(), depth + 1)
return
if root_name != "urlset":
raise ValueError(f"{url} has unsupported root element {root_name!r}")
for value in root.xpath("/sm:urlset/sm:url/sm:loc/text()", namespaces=NS):
address = value.strip()
if not address:
continue
parsed = urlparse(address)
if parsed.scheme not in {"http", "https"} or not parsed.netloc:
continue
if address not in seen_urls:
seen_urls.add(address)
found.append(address)
if len(found) > max_urls:
raise ValueError(f"More than {max_urls:,} URLs found")
walk(sitemap_url, 0)
return found
if __name__ == "__main__":
import sys
for address in extract_urls(sys.argv[1]):
print(address)
Run it with either a regular file or an index:
python extract_sitemap.py https://example.com/sitemap.xml > urls.txt
python extract_sitemap.py https://example.com/sitemap-index.xml > urls.txt
Why the namespace matters
Sitemap elements normally belong to http://www.sitemaps.org/schemas/sitemap/0.9. In namespace-aware XPath, an arbitrary prefix such as sm is bound to that URI. An expression such as //url/loc therefore returns nothing on a correctly namespaced document. The script uses /sm:urlset/sm:url/sm:loc and /sm:sitemapindex/sm:sitemap/sm:loc so the namespace cannot be silently ignored.
What the script deliberately does not change
- It trims surrounding whitespace but preserves URL spelling, query strings, fragments, and encoding.
- It does not infer whether a URL is indexed, canonical, reachable, or duplicate under a site’s own normalization rules.
- It rejects non-HTTP(S) and malformed addresses instead of emitting unusable output.
- It keeps a visited set for sitemap files and a separate set for page URLs, so repeated references do not create loops or duplicate lines.
Handling compressed files, limits, and scope
Gzip
Sitemaps are often published as sitemap.xml.gz. The response may be transparently decompressed by the server/client, or the bytes may remain compressed. The helper first parses normally and then tries gzip decompression when the URL ends in .gz. If your server sends a gzip content encoding rather than a gzip sitemap file, requests normally decodes that transfer encoding automatically.
Per-file size and URL limits
Google Search Central’s 2026 documentation specifies a maximum of 50 MB uncompressed or 50,000 URLs per sitemap file. Larger collections should be split into multiple files and referenced by an index. The script’s max_urls is an application safety budget across the whole traversal; set it to a value appropriate for your job rather than assuming Google’s per-file limit is a total-site limit.
Host and protocol scope
A sitemap should list URLs allowed by the host and protocol scope of its location. If you are auditing a site, validate each parsed URL against an allowlist before crawling or storing it. Do not automatically “fix” http to https, remove parameters, or convert trailing slashes unless that is an explicit project rule.
Recommended Free Tools
Rank #2
Preserving metadata when you need it
If your export needs modification dates, select each <url> element and read its child values rather than extracting only text:
for item in root.xpath("/sm:urlset/sm:url", namespaces=NS):
loc = item.xpath("string(sm:loc)", namespaces=NS).strip()
lastmod = item.xpath("string(sm:lastmod)", namespaces=NS).strip() or None
print({"url": loc, "lastmod": lastmod})
Store lastmod as supplied, with its original precision. It is a publisher-provided change date and should not be presented as evidence that Google crawled or indexed the page.
Alternative extraction approaches
Scrapy
Scrapy’s SitemapSpider accepts one or more sitemap URLs, follows sitemap indexes, and yields matching entries for a crawl. Scrapy removes XML namespaces from tags in its item representation, which can make selectors simpler. Use it when extraction is part of a larger crawl with concurrency, retries, pipelines, and robots-policy handling; use the small Python script when you only need a deterministic URL file.
Hosted extraction API
A hosted service can be useful when you do not want to maintain XML fetching, recursion, retries, and exports. SitemapKit documents an authenticated endpoint with sitemap-index recursion up to depth 5 and a maxUrls parameter capped at 50,000. Verify its current authentication, output format, limits, and retention terms before depending on it.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Generate from your own database
For a site you operate, Google Search Central recommends extracting the canonical URL set from your database or having the site software generate the sitemap. That avoids treating a published sitemap as a database of record and lets you apply your own canonicalization and publication rules before XML generation.
Or skip the browser setup
If your next step is to inspect each URL visually rather than parse XML, ScreenshotNeo provides a website screenshot API and MCP server. One GET request returns a PNG, JPEG, WebP, or PDF. Before capture it accepts the cookie/consent banner like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each cleanup step can be disabled. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing state.
For a single URL:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python:
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`${res.status} ${res.statusText}`);
const fs = await import('node:fs/promises');
await fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer()));
See the ScreenshotNeo API documentation for all options. It supports full-page captures with lazy images loaded, CSS-selector element capture, dark mode, 12 device presets or custom viewports, retina scale, PDF paper settings and page ranges, custom CSS/JavaScript, pre-capture clicks, hidden selectors, selector/delay/network-idle waits, request and resource blocking, headers/cookies/user agents/Authorization, timezone and geolocation, transparent backgrounds, resizing, TTL caching, signed image links, asynchronous jobs with signed webhooks, bulk capture of 100 URLs per call, a usage API, an OpenAPI specification, and familiar parameter names for easier migration.
An MCP server exposes take_screenshot, get_page_info, and capture_pdf to Claude, Cursor, and other MCP clients. The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.
The Tool Desk
Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Troubleshooting
“No URLs found”
Check the root element and namespace. Print etree.QName(root).localname; if it is sitemapindex, you must recurse into child files. If it is urlset, use namespace-qualified XPath. Also confirm that you fetched XML rather than an HTML login page or an error document.
HTTP 403, 404, or 5xx
Inspect the response URL, status, redirects, and content type. A 404 usually means the sitemap path is wrong; 403 may require an allowed user agent or authenticated access; 5xx calls for retry with backoff and a bounded attempt count. Do not parse a failed response body as XML.
XML syntax errors
Save the response bytes and inspect the first line and encoding. Common causes are an HTML error page, truncated transfer, invalid characters, or a file advertised as XML that is actually compressed. The parser configuration above disables external entity resolution and network access; repair or regenerate malformed source XML rather than enabling unsafe recovery.
Recursion never finishes
Use both a visited sitemap set and a maximum depth. Cyclic indexes are invalid but can occur accidentally; a visited set makes the traversal terminate. Keep a total URL budget to protect memory and runtime when a source unexpectedly points to a very large collection.
Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchWindows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallDuplicate or apparently different URLs
Exact-string deduplication catches repeated entries. Query parameters, percent-encoding, case, default ports, fragments, and trailing slashes may represent equivalent resources or intentionally different ones. Apply normalization only with rules agreed for your project, and document every transformation.
Best Value
Operational checklist
- Start with an absolute sitemap or sitemap-index URL.
- Check status and content before parsing.
- Use an XML parser with external entities and network access disabled.
- Inspect the root local name and branch between
urlsetandsitemapindex. - Use the standard namespace in every XPath.
- Trim
loctext, validate HTTP(S), and deduplicate. - Follow
.xml.gzfiles and enforce depth, file, and total-URL budgets. - Respect host/protocol scope and the 50 MB/50,000-URL per-file limits.
- Preserve
lastmodonly when your workflow needs it.
Frequently Asked Questions
Can I extract URLs without downloading the linked pages?
Yes. Sitemap extraction downloads the XML sitemap files only; fetching each page is a separate crawl or screenshot step.
What if a site has several sitemap indexes?
Begin with the advertised index and recurse through every child index, while tracking visited files and enforcing a depth and URL budget.
Should I trust every URL in a sitemap?
Treat entries as publisher-supplied data. Validate scheme, host scope, and your own canonicalization rules before crawling or importing them.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

