To audit page titles and meta descriptions across a website, first collect URLs from its XML sitemap, fetch each page’s HTML, and extract the <title> element and <meta name="description"> content. Save the URL, redirect destination, status, extracted values, and quality flags in a CSV or database. If metadata is missing from the initial HTML because a page adds it with JavaScript, render that page in a browser and extract it from the resulting DOM.
A sitemap is a strong starting inventory, not a guarantee that every page is included. Treat the crawl as an audit: record how URLs were discovered, respect site access rules, and review missing, duplicated, or misleading metadata rather than relying on a universal character-count cutoff.
What you need to extract
A page title is the text in the document’s <title> element. A meta description is the value of the content attribute on a tag such as <meta name="description" content="Page summary">. These are different fields in the page markup; neither should be inferred from the URL or visible heading when the corresponding tag is absent.
Collect both the original value and a normalized value. The raw field preserves what the page sent; the normalized field makes comparison easier by trimming and collapsing whitespace. Also retain status and source information so an empty value can be distinguished from a page that failed to load.
#1 Best Overall
Build a URL inventory before crawling
Start with the sitemap
Request the site’s XML sitemap, commonly available at /sitemap.xml, and parse every <loc> entry. If it is a sitemap index, retrieve each referenced sitemap and collect its entries too. Keep <lastmod> when present: it can help prioritize recently changed pages, but it is not proof that a page’s metadata changed on that date.
Sitemaps help search engines discover URLs; they do not replace ordinary crawling or guarantee full coverage. Supplement the inventory with canonical internal links from the home page or other known pages, and add manual seeds for important URLs that neither source found. Store a discovery-source value such as sitemap, internal_link, or manual_seed.
Normalize and constrain URLs
Before fetching, remove fragments because they identify locations within a page rather than separate HTTP resources. Decide how to handle tracking query parameters and URL variants, and apply that same rule consistently. Keep only allowed same-domain URLs for a site-wide crawl unless you explicitly intend to include subdomains or external destinations. Deduplicate after normalization, but retain the originally discovered URL for traceability.
Rank #2
Do not assume every URL in a sitemap is a page to parse: it may redirect, return an error, serve a non-HTML file, or be excluded from crawling. Record those outcomes instead of silently dropping them.
Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchWindows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallFetch pages safely and preserve provenance
For each request, record the requested URL, final URL after redirects, HTTP status, response content type, fetch timestamp, and whether the response is HTML. Use a clear user agent, bounded concurrency, and retries with backoff for transient failures. These are operational safeguards; a suitable request rate depends on the site and its access rules, not a universal vendor limit.
Respect robots directives and other site access controls. A robots meta instruction can only be read if your crawler can access the page, so a blocked URL should not be reported as though its metadata were inspected. Keep fetch failures and blocked pages in separate issue categories.
Run a small-site extraction with Python
For a modest site whose metadata is present in the response HTML, an HTTP client and Beautiful Soup are enough. Install the dependencies with python -m pip install requests beautifulsoup4. The example below reads a sitemap, follows sitemap indexes, filters URLs to the sitemap host, fetches HTML pages with a delay, and writes a CSV.
Save it as extract_metadata.py and run python extract_metadata.py https://example.com/sitemap.xml metadata.csv, replacing the example domain and sitemap path with the site you are authorized to crawl.
import csv
import sys
import time
from datetime import datetime, timezone
from urllib.parse import urljoin, urlparse, urldefrag
import requests
from bs4 import BeautifulSoup
HEADERS = {"User-Agent": "MetadataAudit/1.0 (contact: [email protected])"}
TIMEOUT = 20
DELAY_SECONDS = 0.5
def fetch(url):
return requests.get(url, headers=HEADERS, timeout=TIMEOUT)
def sitemap_entries(sitemap_url, allowed_host, seen=None):
seen = seen or set()
if sitemap_url in seen:
return []
seen.add(sitemap_url)
response = fetch(sitemap_url)
response.raise_for_status()
soup = BeautifulSoup(response.content, "xml")
entries = []
# A sitemap index points to other sitemap files; a urlset lists pages.
for node in soup.find_all("sitemap"):
loc = node.find("loc")
if loc and loc.text.strip():
child = urljoin(sitemap_url, loc.text.strip())
if urlparse(child).hostname == allowed_host:
entries.extend(sitemap_entries(child, allowed_host, seen))
for node in soup.find_all("url"):
loc = node.find("loc")
if not loc or not loc.text.strip():
continue
url, _fragment = urldefrag(loc.text.strip())
if urlparse(url).hostname == allowed_host:
lastmod = node.find("lastmod")
entries.append((url, lastmod.text.strip() if lastmod else ""))
return entries
def normalized(value):
return " ".join((value or "").split())
def extract(url, lastmod):
fetched_at = datetime.now(timezone.utc).isoformat()
row = {
"url": url, "lastmod": lastmod, "final_url": "", "status": "",
"content_type": "", "title_raw": "", "title_normalized": "",
"description_raw": "", "description_normalized": "",
"metadata_source": "initial_html", "issue_flags": "", "fetched_at": fetched_at,
}
flags = []
try:
response = fetch(url)
row["final_url"] = response.url
row["status"] = response.status_code
row["content_type"] = response.headers.get("Content-Type", "")
if response.status_code >= 400:
flags.append("http_error")
if "html" not in row["content_type"].lower():
flags.append("not_html")
else:
soup = BeautifulSoup(response.text, "html.parser")
title_tags = soup.find_all("title")
descriptions = soup.find_all(
"meta", attrs={"name": lambda value: value and value.lower() == "description"}
)
title = title_tags[0].get_text(" ", strip=True) if title_tags else ""
description = descriptions[0].get("content", "").strip() if descriptions else ""
row["title_raw"] = title
row["title_normalized"] = normalized(title)
row["description_raw"] = description
row["description_normalized"] = normalized(description)
if not title:
flags.append("missing_title")
if not description:
flags.append("missing_description")
if len(title_tags) > 1:
flags.append("multiple_title_tags")
if len(descriptions) > 1:
flags.append("multiple_description_tags")
except requests.RequestException as error:
row["status"] = "request_failed"
flags.append("fetch_error:" + type(error).__name__)
row["issue_flags"] = ";".join(flags)
return row
def main():
if len(sys.argv) != 3:
raise SystemExit("Usage: python extract_metadata.py SITEMAP_URL OUTPUT.csv")
sitemap_url, output_path = sys.argv[1:]
host = urlparse(sitemap_url).hostname
if not host:
raise SystemExit("Sitemap URL must include a hostname")
urls = sitemap_entries(sitemap_url, host)
columns = ["url", "lastmod", "final_url", "status", "content_type",
"title_raw", "title_normalized", "description_raw",
"description_normalized", "metadata_source", "issue_flags", "fetched_at"]
with open(output_path, "w", newline="", encoding="utf-8") as file:
writer = csv.DictWriter(file, fieldnames=columns)
writer.writeheader()
for url, lastmod in urls:
writer.writerow(extract(url, lastmod))
time.sleep(DELAY_SECONDS)
if __name__ == "__main__":
main()
This deliberately conservative script is a starting point, not a full crawler. It stays on the sitemap host, fetches sequentially, and captures the first title and description candidate while flagging multiple tags. It does not discover internal links, inspect robots policy automatically, render JavaScript, or calculate duplicate groups; add those capabilities if your audit needs them. Review the site’s robots rules and permissions before running it, and adjust the user agent and delay appropriately.
Rank #4
When initial HTML is not enough
Parse the initial response first. If the title or description is missing, or you know the application inserts or changes metadata after load, send that URL to a browser-rendering queue and inspect the rendered DOM. Store metadata_source=initial_html or metadata_source=rendered_dom so later reviewers can tell why the values differ.
Do not render every page by default if only a subset needs it. A selective second pass lowers browser work and makes it easier to identify the pages whose metadata depends on JavaScript. If the initial and rendered values differ, keep both for diagnosis rather than overwriting the first result.
Turn extracted values into an actionable audit
Extraction gives you data; it does not decide whether the data is good. Create flags and review queues for missing titles, missing descriptions, duplicates, boilerplate, and metadata that does not match the visible page content. A practical export can include:
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Best Value
urlandfinal_urlto identify the requested page and redirect destination;status,content_type, andfetched_atto explain the fetch result;title_raw,title_normalized,description_raw, anddescription_normalized;metadata_source, canonical URL, robots information, and sitemaplastmodwhen available;duplicate_group,issue_flags, and a representative URL for each group.
Group exact duplicates by lowercasing and collapsing whitespace in the normalized values. Near-duplicates require a similarity rule and editorial review; do not treat a shared brand suffix as proof that two page titles are interchangeable. Keep separate queues for missing metadata, duplicate metadata, boilerplate variants, JavaScript-only metadata, and pages whose metadata may describe the wrong content.
Review for usefulness, not a fixed character limit
Titles should be descriptive, concise, and distinct. Repeated boilerplate titles are a quality problem, as are vague values such as “Home” when they do not identify the page. Descriptions should describe the particular page; identical or similar descriptions across pages are not helpful when individual pages appear in search results.
Do not label a title or description wrong solely because it exceeds a universal character threshold. Search results can truncate title links and snippets as needed, so length is a review signal, not a universal hard limit. Check whether the wording is accurate, distinct, and useful for the page instead.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Choose a method that fits the site
- Small site or one-off audit: use an HTTP client and HTML parser when metadata is present in the response HTML.
- Large crawl or recursive discovery: Scrapy provides crawling and extraction primitives; it can be paired with Beautiful Soup for parsing responses.
- JavaScript-heavy application: keep direct HTML parsing as the first pass, then add browser rendering only for URLs that require it.
- No-code audit tool: compare URL coverage and JavaScript-rendering behavior against a sample you have checked yourself before relying on the full report.
When evaluating a crawler or audit service, compare URL discovery coverage, rendering, duplicate detection, robots handling, retry and concurrency controls, export quality, scheduling, and total cost. Those are engineering trade-offs; choose based on your site’s size and needs rather than assuming one approach is always best.
The Tool Desk
Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Troubleshoot common extraction problems
- The sitemap request returns an error: confirm the sitemap URL and response status, check whether the site uses a sitemap index or a different sitemap location, and verify that your request is permitted.
- The sitemap is valid but misses known pages: supplement it with canonical internal links and manual seeds. Store the discovery source so the report shows how each URL entered the crawl.
- The title or description is blank: inspect the original response HTML. If the application adds metadata after load, send that page through the browser-rendering pass and record the rendered source.
- The script reports non-HTML or an HTTP error: check the recorded content type, status, and final URL. The entry may be a redirect, an inaccessible page, or a non-page asset; do not interpret it as a successful metadata extraction.
- Many URLs appear more than once: normalize fragments and agreed tracking parameters, then deduplicate before fetching. Preserve the discovered form if you need to trace a variant back to its source.
- Values differ between runs: compare timestamps, final URLs, and metadata source. The page may have changed, redirected, or generated metadata client-side; retain provenance rather than silently replacing the earlier value.
- The crawl is slow or puts load on the site: reduce concurrency, increase the delay, and use bounded retries with backoff. A large sitemap is not a reason to send unbounded requests.
Or skip the browser setup
If you need rendered screenshots while checking how pages appear, ScreenshotNeo is a website screenshot API and MCP server; it is not a replacement for parsing title and description tags into a crawl report. One GET request returns a screenshot or PDF, and browser rendering can help inspect a page whose appearance depends on JavaScript. The page can also be captured after consent banners, popups, and chat widgets are removed. See the ScreenshotNeo API documentation for request options.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp
Bot checks, blank pages, timeouts, and failed loads are not billed; cache hits are not billed either. An MCP server provides take_screenshot, get_page_info, and capture_pdf for AI agents and MCP clients. The free plan includes 1,000 screenshots a month with no card, and paid plans start at $5 for 3,000. Create a free ScreenshotNeo account to try it.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




