Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minutePC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Short answer: fetch a site’s top-level /robots.txt, collect every Sitemap: record, then fetch each sitemap and recursively follow any sitemap indexes. That produces a list of URLs declared in the site’s sitemaps—not a guaranteed list of every page on the site. Some pages may not be listed, and listed URLs may be unavailable, noncanonical, blocked from crawling, or never indexed.
What you can—and cannot—get from robots.txt
A robots.txt file can point crawlers and other clients to one or more sitemap files. It does not ordinarily enumerate a website’s pages itself. To build an inventory, you must follow the sitemap URLs and read the URLs in their XML files.
That distinction matters when someone asks for “every page.” The output is the set of URLs declared in the sitemaps you successfully retrieved and parsed. It cannot establish that the set includes every page the site has published, or that each listed URL currently works, is unique in the site’s preferred sense, is crawlable, or appears in search results. Google Search Central says: “A sitemap helps search engines discover URLs on your site, but it doesn’t guarantee that all the items in your sitemap will be crawled and indexed.” Google also notes that a URL disallowed by robots.txt may still be indexed if other pages link to it.
Robots.txt is crawler guidance, not access authorization. Do not treat it as a permission mechanism or use sitemap discovery as a reason to access pages you are not authorized to retrieve.
Do these 3 things before closing this tab:
1Repair Windows errors before they cause bigger problems2Fix the driver behind crashes, sound loss and screen glitches3Clear out junk files and repair common Windows errors#1 Best Overall
How sitemap discovery works
- Request the origin’s robots.txt. Use the site’s supported protocol and request the top-level
/robots.txtpath—for example,https://example.com/robots.txt. RFC 9309 specifies that robots.txt is served at that top-level path, encoded as UTF-8, with thetext/plainmedia type. - Read all Sitemap records. Google documents the field as
Sitemap: [absoluteURL]. Multiple records are allowed; the field is independent of anyUser-agentgroup, and the sitemap URL may point to another host. - Fetch each declared URL. A sitemap URL may identify a URL set or a sitemap index. A URL set contains page locations; an index points to child sitemap files. Follow the index entries, then collect the page locations from the URL sets.
- Record what happened. Preserve the source robots.txt URL and each sitemap URL, retrieval time, HTTP status, and parse result. Record failures rather than silently treating an incomplete run as a complete inventory.
The Sitemap: record is an additional robots.txt record; it must not interfere with parsing the core User-agent, Allow, and Disallow rules. Parse sitemap records separately from crawler-rule evaluation.
Run a small Python extractor
This example uses only Python’s standard library. It requests the top-level robots.txt, collects case-insensitive Sitemap field names, removes trailing comments, follows sitemap indexes, and prints discovered locations. It intentionally fails visibly on HTTP, XML, or network errors rather than claiming a complete result after a failed fetch.
Rank #2
- HTML CSS Design and Build Web Sites
- Comes with secure packaging
- It can be a gift option
from urllib.error import HTTPError, URLError
from urllib.parse import urlparse
from urllib.request import Request, urlopen
import xml.etree.ElementTree as ET
def fetch(url):
request = Request(url, headers={"User-Agent": "SitemapInventory/1.0"})
with urlopen(request, timeout=20) as response:
status = response.status
final_url = response.geturl()
data = response.read()
return status, final_url, data
def sitemap_urls_from_robots(robots_url):
status, final_url, data = fetch(robots_url)
text = data.decode("utf-8-sig")
found = []
for line in text.splitlines():
field, separator, value = line.partition(":")
if separator and field.strip().lower() == "sitemap":
value = value.split("#", 1)[0].strip()
parsed = urlparse(value)
if parsed.scheme in ("http", "https") and parsed.netloc:
found.append(value)
else:
print(f"Ignoring non-absolute Sitemap value: {value!r}")
return status, final_url, found
def local_name(tag):
return tag.rsplit("}", 1)[-1].lower()
def loc_values(root):
return [node.text.strip() for node in root.iter()
if local_name(node.tag) == "loc" and node.text and node.text.strip()]
def collect_sitemaps(start_urls):
pending = list(start_urls)
seen_sitemaps = set()
page_urls = []
seen_pages = set()
while pending:
sitemap_url = pending.pop(0)
if sitemap_url in seen_sitemaps:
continue
seen_sitemaps.add(sitemap_url)
status, final_url, data = fetch(sitemap_url)
root = ET.fromstring(data)
kind = local_name(root.tag)
locations = loc_values(root)
if kind == "sitemapindex":
pending.extend(locations)
elif kind == "urlset":
for url in locations:
if url not in seen_pages:
seen_pages.add(url)
page_urls.append(url)
else:
raise ValueError(f"Unexpected XML root {root.tag!r} at {final_url}")
print(f"Fetched {sitemap_url}: HTTP {status}; parsed {kind}; "
f"{len(locations)} loc element(s)")
return page_urls
robots = "https://example.com/robots.txt"
try:
status, robots_final_url, declared = sitemap_urls_from_robots(robots)
print(f"Robots file: {robots_final_url} (HTTP {status})")
print(f"Declared sitemap URLs: {len(declared)}")
pages = collect_sitemaps(declared)
for page in pages:
print(page)
print(f"Unique exact URL strings collected: {len(pages)}")
except (HTTPError, URLError, TimeoutError, UnicodeDecodeError,
ET.ParseError, ValueError) as error:
raise SystemExit(f"Extraction incomplete: {error}")
Save it as extract_sitemaps.py, replace https://example.com with the site’s origin, and run python extract_sitemaps.py. The example preserves each <loc> value as text and de-duplicates exact repeats; it does not canonicalize URLs or decide whether two different strings represent the same page.
What to change before using it at scale
- Keep provenance in structured output. The sample prints the final robots URL and per-sitemap retrieval details. For an audit, also persist the original sitemap URL, timestamp, status, parser result, and any error alongside each result.
- Set operational limits. Add explicit limits for response size, number of child sitemaps, redirects, elapsed time, and retries appropriate to your environment. A sitemap index can lead to many requests; bounded processing prevents an unexpected crawl from consuming unbounded time or memory.
- Decide how to handle compression and content types. The sample parses XML bytes and does not implement special handling for compressed sitemap files or validate response content types. If your target publishes compressed files or inconsistent headers, add and test those behaviors rather than assuming the basic script handles them.
- Choose duplicate semantics deliberately. Exact-string de-duplication is predictable and reversible. URL normalization can merge values that differ in case, escaping, query strings, or trailing slashes; apply it only if your inventory specification defines those transformations.
- Preserve partial results on failure. For batch jobs, catch errors per sitemap, write successful results and an error log, and mark the run incomplete. A single failed child sitemap should not be mistaken for proof that no other URLs exist.
Parsing decisions that affect the result
Robots.txt line handling
Field names should be handled case-insensitively where your parser supports that behavior. Ignore blank lines and comments without accidentally treating comment text as part of a sitemap URL. Validate that a Sitemap value is an absolute URL before requesting it. The Python example uses the text before a # as the value and accepts HTTP or HTTPS URLs; adjust comment handling only if your implementation has a clearly defined syntax policy.
Rank #3
Indexes, cycles, and repeated files
Follow sitemap indexes recursively, not just the first level. Keep a set of sitemap URLs already visited so repeated entries or cycles do not cause repeated fetching forever. The example de-duplicates the exact sitemap URL string before fetching it; it does not merge redirect aliases. A production crawler should set a maximum traversal depth or count as an additional safety boundary.
XML and malformed input
A malformed XML response should be reported with the sitemap URL and retrieval details. Silently skipping it makes the final list look more complete than it is. The standard-library parser shown here stops at malformed XML; it does not attempt recovery. Be explicit about whether your application rejects the whole run, keeps prior results and records the error, or uses a tolerant parser.
Rank #4
- Brand: Wiley
- Set of 2 Volumes
- A handy two-book set that uniquely combines related technologies Highly visual format and accessible language makes these books highly effective learning tools Perfect for beginning web designers and front-end developers
Cross-host sitemap URLs
Google’s documentation allows a Sitemap record to point to a URL on another host. Do not assume the sitemap is hosted on the same origin as robots.txt. At the same time, apply your own network and security policy before fetching arbitrary URLs: validate schemes, restrict destinations if needed, and avoid turning an extractor into an unrestricted URL-fetching service.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.How to interpret and export the inventory
Call the output a sitemap-declared URL inventory, and attach a run status such as complete or incomplete. “Complete” should mean that all discovered sitemap documents in scope were fetched and parsed successfully—not that the inventory contains every page the website has ever exposed.
The Tool Desk
Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Best Value
- Keep source columns: page URL, source sitemap URL, robots.txt URL, retrieval timestamp, HTTP status, and parse outcome.
- Separate observations from conclusions: a URL appearing in XML establishes that it was listed in that retrieved document; it does not establish that the URL resolves, is canonical, is crawlable, or is indexed.
- Make transformations visible: if you normalize or filter URLs, retain the original value and record the transformation so results can be reproduced.
- Report errors alongside counts: a URL count without the failed sitemap list and run status can conceal missing branches of an index.
Export to the format your downstream process needs, such as newline-delimited text for simple pipelines or CSV/JSON when provenance fields matter. The format does not change the meaning of the inventory.
Troubleshooting common failures
- No Sitemap records found: the file may not declare any. Check that you requested the origin’s top-level robots.txt, followed its redirect, and parsed the response body rather than an error page. A missing declaration does not prove the site has no sitemap; it means this discovery route did not provide one.
- Robots request returns an error: check the requested scheme and host, response status, redirects, and network access. Record the failure and do not report an empty inventory as a successful extraction.
- XML parsing fails: verify the fetched body is actually sitemap XML and not an HTML error or access-denied page. Preserve the response URL and status, then inspect the body safely; do not discard the error from the audit trail.
- Some pages seem absent: determine whether all sitemap-index children were traversed and whether any fetch or parse errors occurred. Sitemaps are a discovery aid, not a guarantee of total site coverage.
- Duplicate-looking URLs remain: the sample removes only byte-for-byte-equivalent URL strings. Establish a normalization policy before merging URLs that differ in representation.
- A disallowed URL appears in results: sitemap membership and robots crawling rules answer different questions. A listed URL can still be disallowed for crawling; conversely, a disallowed URL may still be indexed through links from other pages.
Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server, not a sitemap extractor; it does not replace the retrieval and XML traversal above. If your adjacent task is capturing what a page looks like, one GET request returns an image or PDF. For example, this captures a visual screenshot of the robots.txt URL—it does not turn its contents into a URL inventory:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com/robots.txt -o shot.webp
See the ScreenshotNeo API documentation for request options. Cookie banners, popups, and chat widgets are removed before the shot; bot checks, blank pages, and failed loads are never billed. Its MCP server lets AI agents take screenshots, and 1,000 screenshots a month are free with no card; paid plans start at $5 for 3,000. Learn about ScreenshotNeo, or sign up for 1,000 free screenshots a month with no card.
Frequently Asked Questions
Does robots.txt list every page on a website?
No. It can declare sitemap locations, and those sitemaps list URLs, but neither file guarantees a complete list of all site pages.
Free tools Windows power users keep installed
One-click scans. No signup required.
Can a sitemap be hosted on a different domain?
Yes. Google documents that a Sitemap record may point to a sitemap on another host.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

