Build a sitemap link extractor in n8n with this workflow: HTTP Request → XML → branch for sitemap indexes or page sitemaps → Split Out or Code → deduplicate → export. A flat sitemap contains page URLs under urlset; a sitemap index contains links to other sitemap files under sitemapindex, so it must be fetched and parsed recursively before you have the page URLs.
The examples below show how to accept a sitemap URL, parse both document types, retain lastmod and source-sitemap context, and prepare results for CSV, Google Sheets, a database, or later link checks.
What the workflow extracts
A standard XML sitemap identifies pages with a urlset root and one or more url entries. Each entry has a required loc child containing the URL; lastmod is optional metadata. A sitemap index instead has a sitemapindex root and sitemap entries whose loc values point to child sitemap files. You must fetch those files before extracting their page-level URLs. See the Sitemaps protocol.
In n8n, use an HTTP Request node to retrieve XML, then the native XML node to convert it into structured data. Branch based on the parsed root, flatten the appropriate array to one item per URL, and send the result wherever it is needed. The exact field paths can vary with the XML node output shape, so inspect a real parsed item before mapping it.
Free tools Windows power users keep installed
One-click scans. No signup required.
#1 Best Overall
Build the basic flat-sitemap workflow
1. Supply the sitemap URL
Start with a Manual Trigger for a one-off run, or accept a URL using a Webhook, Chat Trigger, or a Set/Edit Fields node. Store it in a clearly named field such as sitemap_url. For a repeatable workflow, validate that the value is an HTTP or HTTPS URL before making a request.
2. Fetch the XML
Add an HTTP Request node and configure it to use the incoming sitemap_url as its request URL. Set the response format to text so the XML response can be passed to the XML node. The HTTP Request node is n8n’s general-purpose node for making REST and API requests; its official documentation is at HTTP Request.
Give failed requests an explicit error path rather than treating an empty response as an empty sitemap. A timeout, blocked request, or server error is not proof that the sitemap has no URLs.
3. Parse the response with XML
Add the native XML node after the HTTP Request node and configure it to convert the response text into structured data. Check the node’s output on a known sitemap. You should see a root corresponding to either urlset or sitemapindex, followed by child records.
Do these 3 things before closing this tab:
1Scan for outdated or missing drivers - takes under a minute2Repair Windows errors before they cause bigger problems3Fix the driver behind crashes, sound loss and screen glitchesXML parsers may represent a single child differently from multiple children, or use arrays at nested levels. Avoid assuming the path or data type until you have checked the actual output. The official n8n XML node reference is XML.
Rank #2
4. Flatten page records
For a flat sitemap, map each url entry to an n8n item containing its loc value and, if present, lastmod. Use Split Out for a straightforward array, or a Code node if the XML shape needs normalization. A convenient output shape is one item per page with fields named url, lastmod, and source_sitemap.
This Code node example is a fallback after XML parsing. It assumes the parsed object has the indicated shape; adapt the path to the output you inspected:
const root = $json;
const rows = root.urlset?.url ?? [];
return rows
.map(entry => ({
json: {
url: entry.loc,
lastmod: entry.lastmod ?? null,
},
}))
.filter(item => typeof item.json.url === 'string' && item.json.url.length > 0);
If the parser yields a single url as an object rather than an array, normalize it to an array before mapping. Preserve lastmod as supplied; it is optional source metadata, not a guarantee that a page changed at that time.
Handle sitemap indexes recursively
A sitemap index is not a list of the site’s page URLs. Its sitemap.loc values identify child sitemap documents. The workflow needs two levels: extract child sitemap URLs, fetch and parse each child document, and then extract page-level url.loc values. Some child documents can themselves be indexes, so a recursive implementation should repeat the same root check until it reaches page sitemaps, subject to a depth or request limit you choose.
- Branch on the parsed root. If the object contains
urlset, send it to page extraction. If it containssitemapindex, send it to child-sitemap extraction. Treat any other root as an unsupported or invalid response. - Split child sitemap locations. Convert each
sitemap.locinto one n8n item. Include the index URL assource_sitemapor equivalent context. - Fetch each child. Loop the child items into the HTTP Request and XML nodes, using each child location as the request URL.
- Check each child root. For a
urlset, extract page records. For anothersitemapindex, continue the loop if nested indexes are within the workflow’s intended scope. - Carry provenance through the loop. Preserve the child sitemap URL alongside each page URL so exported rows and later HTTP checks can be traced to their source.
- Stop safely. Set practical limits for nesting depth, total sitemap requests, and execution time. Route malformed XML, unknown roots, and failed child requests to an error branch rather than silently discarding them.
For an index, the first Code node should emit one item per root.sitemapindex.sitemap[*].loc; iterate over those sitemap URLs and apply page-sitemap extraction to each response. Match the field path to the XML node’s actual output rather than copying a presumed structure blindly.
Rank #3
Clean, filter, and deduplicate the extracted URLs
Before exporting, standardize the record shape and decide what counts as a duplicate. A practical deduplication key is the normalized URL string; if URLs differ only by tracking parameters or trailing-slash conventions, decide explicitly whether those differences should be preserved or normalized for your task. Do not silently rewrite canonical URLs when the goal is an exact sitemap inventory.
- Keep non-empty string values from
loc; optionally reject values that are not valid HTTP or HTTPS URLs. - Remove duplicates after combining the page records from all child sitemaps.
- Optionally filter by hostname, path prefix, scheme, or file extension for a scoped audit.
- Retain
lastmodwhen present and leave it empty or null when absent. - Keep
source_sitemapif you need to diagnose omissions or connect a page to its original sitemap file.
If a downstream HTTP check replaces the current item’s fields with the response, map the original URL and provenance forward explicitly, or use a Merge/Code step to combine the response status with the input item. Otherwise, reports can lose the URL they were meant to check.
The Tool Desk
Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Export to CSV, Google Sheets, or another destination
Once each n8n item represents one page URL, connect it to the destination that fits the job. A CSV export is useful for an offline inventory or handoff; Google Sheets is convenient for collaborative review; a database can support recurring snapshots; and a crawler or HTTP Request stage can turn the list into link checks, redirect analysis, or content-scraping input.
- CSV: map the normalized fields to columns such as
url,lastmod, andsource_sitemap, then use n8n’s file and spreadsheet nodes to produce a file. - Google Sheets: map one extracted item to a row and choose whether repeated runs append rows or update a maintained inventory.
- Database or reporting: store the URL and source context with a run identifier or timestamp if you need to compare repeated exports.
- Link checks: deduplicate first, select the URLs to test, then batch requests and preserve the input URL alongside each returned status.
n8n’s official template gallery includes examples for sitemap-to-CSV extraction, Google Sheets integration, scraping preparation, and broken-link or redirect reporting: n8n workflows.
Respect sitemap limits and n8n capacity
The Sitemap protocol limits one sitemap file to 50,000 URLs and 50MB (52,428,800 bytes) uncompressed. These are per-file limits, not a maximum total for a site: larger inventories should be divided among files and referenced from a sitemap index. The limits are stated by Sitemaps.org.
Rank #4
Large XML responses and large item arrays consume memory, and the amount available depends on your n8n deployment and hosting environment. The n8n sitemap-to-CSV template cautions that files above 50,000 URLs may require more memory depending on hosting. For large sites, process child sitemaps separately, avoid unnecessary copies of full response bodies, and batch downstream page requests instead of firing them all at once. If a run grows too large for one execution, persist intermediate results and continue in smaller batches.
Recommended Free Tools
When a list feeds crawling or scraping, limit concurrency and crawl depth according to your use case. Extraction itself only inventories sitemap entries; it does not establish that every listed page is reachable, indexable, or returning a successful status.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Troubleshooting common failures
The XML node produces no usable fields
Inspect the raw HTTP response and the XML node output. The response may be an HTML error page, a compressed or otherwise unexpected payload, or a valid sitemap whose namespace or nesting differs from your assumed path. Confirm the root name and adapt mappings to the parsed structure.
Only one URL is emitted, or a mapping error appears
A single XML child can be represented as an object while multiple children are represented as an array. Normalize either form to an array before using Split Out or mapping in Code. Test against both one-entry and multi-entry examples.
An index returns sitemap files instead of web pages
That is expected: sitemapindex contains child sitemap locations. Fetch and parse each child, then extract the children of its urlset root. Do not export index locations as though they were page URLs.
Quick wins for a faster PC:
Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Repair Windows errors before they cause bigger problemsFix Now →Scan for outdated or missing drivers - takes under a minuteDriver Scan →Best Value
- Book - powershell for sysadmins: workflow automation made easy
- Language: english
- Binding: paperback
The request fails or returns an empty body
Check that the incoming URL is populated and reachable from the n8n host, and inspect the HTTP status and error output. DNS, TLS, access controls, server errors, and timeouts can prevent a fetch. Add retries only for transient failures and route persistent failures into a report so a partial export is not mistaken for a complete one.
URLs or status fields disappear after checks
HTTP response nodes may change the item shape. Carry the original page URL and source sitemap through the request explicitly, or merge the request result with the input data before building the report.
The workflow runs out of memory or takes too long
Break an index into child files, process smaller batches, and reduce parallel downstream requests. If the sitemap itself is close to protocol limits, the n8n instance still needs enough memory to hold and parse its response; the protocol limit does not guarantee that a particular deployment can process the file in one execution.
Use the extracted URLs for screenshots
If the next step is capturing page previews rather than only exporting or checking URLs, ScreenshotNeo is a website screenshot API and MCP server. It accepts one GET request with a URL and returns a screenshot or PDF. The API can be called from a later n8n HTTP Request node for each extracted URL; keep the list deduplicated and batch the work to fit the size and pace of your workflow.
Or skip the browser setup
Call the ScreenshotNeo API with a page URL; the API documentation is at ScreenshotNeo docs. Example using cURL:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python:
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
ScreenshotNeo removes supported cookie and consent banners, newsletter popups, and chat widgets before capture. Bot checks, blank pages, failed loads, and cache hits are not billed; response headers indicate the page verdict and billing status. Its MCP server exposes take_screenshot, get_page_info, and capture_pdf for AI agents. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 shots. Sign up for 1,000 free screenshots a month, with no card required.
FAQ
Can I extract URLs without a Code node?
Yes. For a simple flat sitemap, the XML node followed by Split Out can flatten the page entries. Use Code when you need to normalize inconsistent array shapes, branch recursively, or apply custom cleanup.
Does lastmod have to be present?
No. It is optional sitemap metadata. Preserve it when supplied, but do not reject a valid URL entry because it is missing.
Can this workflow also find broken links?
Yes. Send the deduplicated page URL items to HTTP checks or a crawler, while retaining each input URL and its source sitemap for the resulting report.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

