Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

To extract every URL, download the sitemap, parse XML with the Sitemap namespace, and read each <url><loc> value. If the document is a sitemap index, read its <sitemap><loc> entries and recursively process each child file. The Python example below handles indexes, namespaces, compressed files, deduplication, limits, and common failure cases.

What a sitemap contains

The Sitemap protocol is an XML format (the protocol documentation says, “The Sitemap protocol format consists of XML tags.”). A normal sitemap has a <urlset> root and one <url> element per page. The page address is in <loc>; optional metadata includes <lastmod>, <changefreq>, and <priority>. A sitemap index instead has a <sitemapindex> root and links to other sitemap files through <sitemap><loc>.

Use fully qualified absolute URLs. Keep extracted addresses within the host and protocol scope permitted by the sitemap’s location. A sitemap is a URL declaration, not proof that a search engine indexed a page; <lastmod> is update metadata, not an indexing result.

Python: extract URLs from one sitemap or an index

Install the two dependencies first:

python -m pip install requests lxml

This script checks HTTP responses, parses namespace-qualified XML, follows indexes recursively, supports .xml.gz, prevents loops, enforces depth and URL budgets, and returns unique URLs.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
from __future__ import annotations

from gzip import decompress
from io import BytesIO
from typing import Iterable
from urllib.parse import urlparse

import requests
from lxml import etree

SITEMAP_NS = "http://www.sitemaps.org/schemas/sitemap/0.9"
NS = {"sm": SITEMAP_NS}


def _xml_root(content: bytes, source_url: str) -> etree._Element:
    # XMLParser does not resolve external entities, avoiding XXE-style fetches.
    parser = etree.XMLParser(
        resolve_entities=False,
        no_network=True,
        load_dtd=False,
        recover=False,
        huge_tree=False,
    )
    try:
        return etree.fromstring(content, parser=parser)
    except etree.XMLSyntaxError:
        if source_url.lower().split("?", 1)[0].endswith(".gz"):
            return etree.fromstring(decompress(content), parser=parser)
        raise


def extract_urls(
    sitemap_url: str,
    *,
    max_depth: int = 10,
    max_urls: int = 1_000_000,
    timeout: float = 30,
) -> list[str]:
    session = requests.Session()
    visited: set[str] = set()
    found: list[str] = []
    seen_urls: set[str] = set()

    def walk(url: str, depth: int) -> None:
        if depth > max_depth:
            raise ValueError(f"Sitemap index depth exceeds {max_depth}: {url}")
        if url in visited:
            return
        visited.add(url)

        response = session.get(
            url,
            timeout=timeout,
            headers={"Accept": "application/xml,text/xml;q=0.9,*/*;q=0.1"},
        )
        response.raise_for_status()
        root = _xml_root(response.content, url)
        root_name = etree.QName(root).localname

        if root_name == "sitemapindex":
            children = root.xpath("/sm:sitemapindex/sm:sitemap/sm:loc/text()", namespaces=NS)
            for child in children:
                walk(child.strip(), depth + 1)
            return

        if root_name != "urlset":
            raise ValueError(f"{url} has unsupported root element {root_name!r}")

        for value in root.xpath("/sm:urlset/sm:url/sm:loc/text()", namespaces=NS):
            address = value.strip()
            if not address:
                continue
            parsed = urlparse(address)
            if parsed.scheme not in {"http", "https"} or not parsed.netloc:
                continue
            if address not in seen_urls:
                seen_urls.add(address)
                found.append(address)
                if len(found) > max_urls:
                    raise ValueError(f"More than {max_urls:,} URLs found")

    walk(sitemap_url, 0)
    return found


if __name__ == "__main__":
    import sys

    for address in extract_urls(sys.argv[1]):
        print(address)

Run it with either a regular file or an index:

python extract_sitemap.py https://example.com/sitemap.xml > urls.txt
python extract_sitemap.py https://example.com/sitemap-index.xml > urls.txt

Why the namespace matters

Sitemap elements normally belong to http://www.sitemaps.org/schemas/sitemap/0.9. In namespace-aware XPath, an arbitrary prefix such as sm is bound to that URI. An expression such as //url/loc therefore returns nothing on a correctly namespaced document. The script uses /sm:urlset/sm:url/sm:loc and /sm:sitemapindex/sm:sitemap/sm:loc so the namespace cannot be silently ignored.

What the script deliberately does not change

  • It trims surrounding whitespace but preserves URL spelling, query strings, fragments, and encoding.
  • It does not infer whether a URL is indexed, canonical, reachable, or duplicate under a site’s own normalization rules.
  • It rejects non-HTTP(S) and malformed addresses instead of emitting unusable output.
  • It keeps a visited set for sitemap files and a separate set for page URLs, so repeated references do not create loops or duplicate lines.

Handling compressed files, limits, and scope

Gzip

Sitemaps are often published as sitemap.xml.gz. The response may be transparently decompressed by the server/client, or the bytes may remain compressed. The helper first parses normally and then tries gzip decompression when the URL ends in .gz. If your server sends a gzip content encoding rather than a gzip sitemap file, requests normally decodes that transfer encoding automatically.

Per-file size and URL limits

Google Search Central’s 2026 documentation specifies a maximum of 50 MB uncompressed or 50,000 URLs per sitemap file. Larger collections should be split into multiple files and referenced by an index. The script’s max_urls is an application safety budget across the whole traversal; set it to a value appropriate for your job rather than assuming Google’s per-file limit is a total-site limit.

Host and protocol scope

A sitemap should list URLs allowed by the host and protocol scope of its location. If you are auditing a site, validate each parsed URL against an allowlist before crawling or storing it. Do not automatically “fix” http to https, remove parameters, or convert trailing slashes unless that is an explicit project rule.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Preserving metadata when you need it

If your export needs modification dates, select each <url> element and read its child values rather than extracting only text:

for item in root.xpath("/sm:urlset/sm:url", namespaces=NS):
    loc = item.xpath("string(sm:loc)", namespaces=NS).strip()
    lastmod = item.xpath("string(sm:lastmod)", namespaces=NS).strip() or None
    print({"url": loc, "lastmod": lastmod})

Store lastmod as supplied, with its original precision. It is a publisher-provided change date and should not be presented as evidence that Google crawled or indexed the page.

Alternative extraction approaches

Scrapy

Scrapy’s SitemapSpider accepts one or more sitemap URLs, follows sitemap indexes, and yields matching entries for a crawl. Scrapy removes XML namespaces from tags in its item representation, which can make selectors simpler. Use it when extraction is part of a larger crawl with concurrency, retries, pipelines, and robots-policy handling; use the small Python script when you only need a deterministic URL file.

Hosted extraction API

A hosted service can be useful when you do not want to maintain XML fetching, recursion, retries, and exports. SitemapKit documents an authenticated endpoint with sitemap-index recursion up to depth 5 and a maxUrls parameter capped at 50,000. Verify its current authentication, output format, limits, and retention terms before depending on it.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Generate from your own database

For a site you operate, Google Search Central recommends extracting the canonical URL set from your database or having the site software generate the sitemap. That avoids treating a published sitemap as a database of record and lets you apply your own canonicalization and publication rules before XML generation.

Or skip the browser setup

If your next step is to inspect each URL visually rather than parse XML, ScreenshotNeo provides a website screenshot API and MCP server. One GET request returns a PNG, JPEG, WebP, or PDF. Before capture it accepts the cookie/consent banner like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each cleanup step can be disabled. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing state.

For a single URL:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`${res.status} ${res.statusText}`);
const fs = await import('node:fs/promises');
await fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer()));

See the ScreenshotNeo API documentation for all options. It supports full-page captures with lazy images loaded, CSS-selector element capture, dark mode, 12 device presets or custom viewports, retina scale, PDF paper settings and page ranges, custom CSS/JavaScript, pre-capture clicks, hidden selectors, selector/delay/network-idle waits, request and resource blocking, headers/cookies/user agents/Authorization, timezone and geolocation, transparent backgrounds, resizing, TTL caching, signed image links, asynchronous jobs with signed webhooks, bulk capture of 100 URLs per call, a usage API, an OpenAPI specification, and familiar parameter names for easier migration.

An MCP server exposes take_screenshot, get_page_info, and capture_pdf to Claude, Cursor, and other MCP clients. The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Troubleshooting

“No URLs found”

Check the root element and namespace. Print etree.QName(root).localname; if it is sitemapindex, you must recurse into child files. If it is urlset, use namespace-qualified XPath. Also confirm that you fetched XML rather than an HTML login page or an error document.

HTTP 403, 404, or 5xx

Inspect the response URL, status, redirects, and content type. A 404 usually means the sitemap path is wrong; 403 may require an allowed user agent or authenticated access; 5xx calls for retry with backoff and a bounded attempt count. Do not parse a failed response body as XML.

XML syntax errors

Save the response bytes and inspect the first line and encoding. Common causes are an HTML error page, truncated transfer, invalid characters, or a file advertised as XML that is actually compressed. The parser configuration above disables external entity resolution and network access; repair or regenerate malformed source XML rather than enabling unsafe recovery.

Recursion never finishes

Use both a visited sitemap set and a maximum depth. Cyclic indexes are invalid but can occur accidentally; a visited set makes the traversal terminate. Keep a total URL budget to protect memory and runtime when a source unexpectedly points to a very large collection.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Duplicate or apparently different URLs

Exact-string deduplication catches repeated entries. Query parameters, percent-encoding, case, default ports, fragments, and trailing slashes may represent equivalent resources or intentionally different ones. Apply normalization only with rules agreed for your project, and document every transformation.

Operational checklist

  • Start with an absolute sitemap or sitemap-index URL.
  • Check status and content before parsing.
  • Use an XML parser with external entities and network access disabled.
  • Inspect the root local name and branch between urlset and sitemapindex.
  • Use the standard namespace in every XPath.
  • Trim loc text, validate HTTP(S), and deduplicate.
  • Follow .xml.gz files and enforce depth, file, and total-URL budgets.
  • Respect host/protocol scope and the 50 MB/50,000-URL per-file limits.
  • Preserve lastmod only when your workflow needs it.

Frequently Asked Questions

Can I extract URLs without downloading the linked pages?

Yes. Sitemap extraction downloads the XML sitemap files only; fetching each page is a separate crawl or screenshot step.

What if a site has several sitemap indexes?

Begin with the advertised index and recurse through every child index, while tracking visited files and enforcing a depth and URL budget.

Should I trust every URL in a sitemap?

Treat entries as publisher-supplied data. Validate scheme, host scope, and your own canonicalization rules before crawling or importing them.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.