What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
To extract website metadata, inspect the page’s HTML <head>, then check the response headers and, when JavaScript may change the page, its rendered DOM. A useful extraction records the value, where it came from, and whether it appeared in the original response or only after rendering. Metadata is not one tag, and extracting a value does not guarantee that a search engine or social platform will display it.
What counts as website metadata?
The <head> is the first place to look. It contains the document’s <title>, many <meta> elements, and often link relations such as a canonical URL. Google describes the head as the primary element for specifying metadata about an HTML page. These items are related, but they are not all meta tags.
- Document metadata: the page title and
<meta>name/content pairs, such asdescriptionorrobots. - Social preview metadata: Open Graph properties such as
og:title,og:description, andog:image, plus Twitter/X card fields when present. - Link relations: elements such as
<link rel="canonical">, which are in the head but are not<meta>elements. - Structured data: JSON-LD blocks and Microdata or RDFa markup that describe entities and their properties. These are not ordinary meta tags.
- HTTP response headers: for example,
X-Robots-Tag, which is delivered with the response rather than inside the HTML.
Keep those categories separate in notes or an export. Flattening them into one list of “meta tags” loses the context needed to interpret the data.
Inspect metadata manually in a browser
Check the original HTML source
- Open the page you want to inspect.
- Open its source using the browser’s “View page source” command or the
view-source:prefix where supported. - Search the source for
<title>,name="description",name="robots", andproperty="og:. Also look fortwitter:,rel="canonical", andapplication/ld+json. - Record the exact value and its location. If a field occurs more than once, record each occurrence rather than silently choosing one.
Source inspection shows the HTML response you received. It does not execute page scripts. It is a good first check because it helps distinguish values sent by the server from changes made later in the browser.
Quick wins for a faster PC:
Repair Windows errors before they cause bigger problemsFix Now →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →#1 Best Overall
- C Instruments
- Pages: 160
- Instrumentation: C Instruments
Check the live DOM
Open the browser’s developer tools and use the Elements or Inspector panel to inspect the document’s <head> after the page has loaded. Compare it with the original source if a title, description, or social value is absent or different. A site may insert or update metadata with JavaScript, so the source response and live DOM can disagree. If the value appears only in the live DOM, note that it was rendered, not present in the initial HTML.
Extract metadata with Python
This standard-library example fetches a URL, follows normal redirects handled by urllib, records the final URL, status and response headers, and parses the returned HTML without executing scripts. It keeps duplicate meta and link entries, captures their source line numbers, collects title text, and retains JSON-LD blocks as separate values. It is a starting point for inspecting server-returned markup, not a complete HTML conformance validator.
from html.parser import HTMLParser
from urllib.request import Request, urlopen
import json
import sys
class MetadataParser(HTMLParser):
def __init__(self):
super().__init__(convert_charrefs=True)
self.items = []
self.title_parts = []
self.in_title = False
self.jsonld_parts = []
self.in_jsonld = False
self.current_jsonld = []
def handle_starttag(self, tag, attrs):
attrs = dict(attrs)
line = self.getpos()[0]
if tag == "title":
self.in_title = True
elif tag == "meta":
self.items.append({"type": "meta", "line": line, **attrs})
elif tag == "link":
self.items.append({"type": "link", "line": line, **attrs})
elif tag == "script" and attrs.get("type", "").lower() == "application/ld+json":
self.in_jsonld = True
self.current_jsonld = []
def handle_endtag(self, tag):
if tag == "title":
self.in_title = False
elif tag == "script" and self.in_jsonld:
self.jsonld_parts.append("".join(self.current_jsonld))
self.in_jsonld = False
self.current_jsonld = []
def handle_data(self, data):
if self.in_title:
self.title_parts.append(data)
if self.in_jsonld:
self.current_jsonld.append(data)
url = sys.argv[1] if len(sys.argv) > 1 else "https://example.com/"
request = Request(url, headers={"User-Agent": "MetadataInspector/1.0"})
try:
with urlopen(request, timeout=20) as response:
raw = response.read()
charset = response.headers.get_content_charset() or "utf-8"
html = raw.decode(charset, errors="replace")
parser = MetadataParser()
parser.feed(html)
result = {
"requested_url": url,
"final_url": response.geturl(),
"status": response.status,
"content_type": response.headers.get("Content-Type"),
"x_robots_tag": response.headers.get_all("X-Robots-Tag", []),
"title": "".join(parser.title_parts).strip(),
"head_items": parser.items,
"json_ld": parser.jsonld_parts,
}
print(json.dumps(result, ensure_ascii=False, indent=2))
except Exception as exc:
print(f"Fetch or parse failed: {exc}", file=sys.stderr)
raise
Save it as extract_metadata.py and run python extract_metadata.py https://example.com/. Replace the example URL with the page to inspect. The output distinguishes link elements, meta elements, JSON-LD and the X-Robots-Tag response header. It deliberately preserves duplicate fields as separate entries. The parser’s line values refer to positions in the response text that Python parsed.
The example decodes using the response’s declared charset, falling back to UTF-8 and replacing undecodable bytes. HTML5 requires a UTF-8 encoding declaration to be located entirely within the first 1024 bytes; a missing or incorrect charset declaration can still complicate extraction. For a production crawler, also set a suitable user agent, limit response size and redirects, handle compressed responses and content types deliberately, and log fetch failures instead of treating them as pages with empty metadata.
Windows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallCrashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minuteWhat to extract and how to interpret it
Title, description and crawler directives
Capture the text inside <title> and the values of relevant <meta> elements, especially name="description" and name="robots". Preserve the original value; trimming whitespace for a normalized display is useful, but do not overwrite the raw value. If the page has duplicates, report them as duplicates. Do not assume the first or last one is the value a crawler will use.
A robots meta directive is an instruction, not evidence that a crawler has followed it. A crawler must be able to access a page to read its directives. Check response access and any applicable X-Robots-Tag header separately, particularly for non-HTML resources such as images or PDFs.
Rank #3
Canonical and social fields
Record the canonical link’s href and other link relations separately from meta values. For Open Graph and Twitter/X fields, preserve the property or name and its content value exactly as supplied. Social preview fields describe what a page declares; they do not establish what a particular platform will show after fetching, caching, or interpreting the page.
Structured data
Keep each JSON-LD script block intact as its own entry before attempting to parse its JSON. A block can contain nested objects or multiple entities; flattening it into a string of name/value pairs discards relationships. Microdata and RDFa are expressed in the HTML rather than in JSON-LD script blocks, so they need their own parsing logic. Google supports JSON-LD, Microdata and RDFa and generally recommends JSON-LD when it is practical to implement and maintain. Valid structured data does not guarantee a rich result: eligibility depends on the relevant Google feature’s documentation and guidelines.
The Tool Desk
Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Build a reliable metadata record
For a one-off inspection, a list of values may be enough. For repeatable extraction or an audit, preserve enough context to reproduce and interpret each result:
Rank #4
- Requested URL, final URL after redirects, fetch time, HTTP status and content type.
- Raw field name, raw value, field type and source location, such as a meta element, link, JSON-LD block or response header.
- Duplicate occurrences rather than a guessed “winning” value.
- Whether the value appeared in the original HTML or only in the rendered DOM.
- Fetch errors, missing fields, malformed markup and decoding issues as explicit outcomes.
Do not fill missing fields with assumed defaults or interpret absence as a specific search result. The extractor reports what it observed, not what a search engine will decide to use.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.When fetching HTML is not enough
A simple HTTP fetch sees the response body; it does not run JavaScript. If the site creates metadata client-side, inspect a rendered DOM with browser developer tools or a browser-rendering workflow. Compare both views: the initial HTML is useful for understanding what the server returned, while the rendered DOM shows changes made after scripts ran. Neither view alone answers what every crawler or social platform will observe, since their rendering and fetching behavior can differ.
Also inspect headers independently of page markup. A parser that reads only HTML cannot see X-Robots-Tag. Likewise, an HTML metadata extractor should not claim that a URL’s content is available just because a request returned a page: record status, content type, redirects and failure conditions along with extracted fields.
Common problems and fixes
- The source has no value, but the page shows one: inspect the live DOM after load; a script may have inserted or changed it. Record the source and rendered results separately.
- The page appears to have conflicting titles or descriptions: retain all occurrences and their locations. Do not silently collapse duplicates into one value.
- There is no HTML to parse: check the HTTP status, redirect destination, content type and fetch error. A timeout, access denial or non-HTML response is not the same as a successful page with no metadata.
- Text is garbled or characters are replaced: check the response charset and the document’s encoding declaration. Decode according to the response where possible and preserve the raw response if exact byte-level analysis matters.
- Structured data is missing from the extracted fields: search for JSON-LD script blocks and inspect Microdata or RDFa separately. The Python example collects JSON-LD only; it does not parse the latter two formats.
- A crawler directive appears ineffective: verify that the crawler can fetch the resource and check both HTML robots metadata and the response’s
X-Robots-Tag. A directive cannot be read from a resource the crawler cannot access. - Metadata in a result page differs from the extracted value: treat this as expected possibility, not an extraction failure. Search engines may generate title links and snippets from more than the page’s declared title and description.
Why extracted values may not match Google’s display
Google may use a page’s description meta tag for a search snippet in some cases, but it can instead select page text. Its title link is generated automatically from multiple signals; the title element is one input, not a guarantee of the displayed wording. Extraction is useful for checking a page’s declared metadata and diagnosing inconsistencies, but it cannot predict an exact search result display.
Similarly, a meta name="keywords" field may still appear in source, but it should not be treated as a reliable SEO field: search engines ignore the keywords meta element. A present field is evidence that markup exists, not proof that it has search value.
Or skip the browser setup
ScreenshotNeo can capture a clean screenshot of a page, which is useful as visual evidence alongside metadata inspection; it does not extract metadata values, so use the Python method or browser tools above for that job. Its API accepts one GET request and returns an image or PDF. For a screenshot of a rendered page, try:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
See the ScreenshotNeo API documentation for request options. ScreenshotNeo accepts cookie and consent banners and removes more than 60 known consent platforms, newsletter popups and chat widgets before capture; each of those steps can be turned off. Bot checks, blank pages, timeouts, failed loads and cache hits are not billed, and responses identify the page verdict and billing status in headers. An MCP server provides take_screenshot, get_page_info and capture_pdf tools for AI agents and MCP clients. The free plan includes 1,000 screenshots per month without a card; paid plans start at $5 for 3,000 shots. See ScreenshotNeo for product details, or sign up for free.
Recommended Free Tools
Frequently Asked Questions
Do I need permission to inspect a public page’s metadata?
A public page can generally be fetched without site credentials, but access controls, rate limits, and the site’s terms still apply. Do not treat a failed or blocked request as proof that the page has no metadata.
Does extracted metadata prove what every crawler sees?
No. The result describes the response or rendered DOM you inspected. Crawlers and social platforms may fetch, render, or interpret pages differently.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




