To extract metadata from a website, fetch its HTML, inspect the document’s <head>, and parse each metadata layer separately: title and standard meta tags, link elements such as canonical and language alternates, social-preview tags, robots directives, and JSON-LD structured data. A command-line fetch is often enough for a server-rendered page; if metadata appears only after JavaScript runs, inspect the rendered DOM or use a rendering-capable service.
What website metadata includes
Metadata is not one field or one format. Different elements in and around a page’s head serve search engines, social networks, browsers, and software consuming structured information. Extracting the title alone does not tell you whether a canonical URL, social image, robots rule, or JSON-LD entity is present.
| Layer | What to look for | Typical use |
|---|---|---|
| Core HTML | <title>, <meta name="description">, charset, viewport |
Page identity, summary, and browser behavior |
| Links in the head | rel="canonical", rel="alternate" |
Preferred page URL and alternate versions, including language variants |
| Robots directives | meta name="robots", bot-specific meta directives, and the HTTP X-Robots-Tag header |
Crawl, indexing, and search-result presentation controls |
| Social metadata | Open Graph properties such as og:title and Twitter Card fields |
Information used to build link previews |
| Structured data | script type="application/ld+json" |
Machine-readable entities and relationships described using vocabularies such as Schema.org |
Google describes meta tags as HTML tags that provide information to search engines and other clients. The <head> is the primary place for page metadata, but robots instructions can also arrive in an HTTP response header. Keep descriptive metadata, structured data, and crawler instructions distinct when you store or analyze results.
Fetch the page and preserve the response
Start with the exact URL you want to investigate. Preserve the requested URL and final URL, because redirects can move a request to a different page. Record the HTTP status, content type, retrieval time, and raw HTML too; those details help explain whether a missing field was absent from the response or missed during parsing.
Free tools Windows power users keep installed
One-click scans. No signup required.
#1 Best Overall
Quick check with curl
For an initial inspection, save the response body and ask curl to show response headers and redirect information:
curl -L -D response-headers.txt -o page.html "https://example.com/page"
-L follows redirects, while -D writes the received headers to a file. Review the final response’s status and content type in that file, then open page.html and inspect its head. This is a useful spot check, not a complete audit: it does not execute page JavaScript, and following redirects means you should still note the final destination.
Check the head, not just a text search
Look for the actual document structure around the metadata. Google lists title, meta, link, script, style, base, noscript, and template among the valid elements for a page head. Invalid elements in the head can lead to later metadata being ignored. See Google’s guidance on valid page metadata.
Extract the core tags and link elements
For each page, collect the title text, description, character encoding, viewport setting, canonical URL, and any language or other alternate links. Preserve both the literal value and a normalized value where useful; for example, retain the canonical link exactly as found, then separately resolve it to an absolute URL if the markup uses a relative path.
Quick wins for a faster PC:
Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Clear out junk files and repair common Windows errorsFree Scan →Scan for outdated or missing drivers - takes under a minuteDriver Scan →When parsing HTML, use an HTML parser rather than a regular expression. HTML can contain unusual whitespace, quoted or unquoted attributes, entity-escaped characters, and malformed markup that a simple pattern will mishandle. Select elements by tag and attributes, and account for repeated fields rather than assuming there is exactly one description, canonical, or robots element.
Rank #2
- HTML CSS Design and Build Web Sites
- Comes with secure packaging
- It can be a gift option
Google’s documentation covers description and robots meta tags in its special tags reference. The presence of a tag is evidence of what the page declares, not proof of how every consumer will interpret or display it.
Extract Open Graph and Twitter Card data
Social metadata is a separate layer. Common Open Graph properties include og:title, og:description, og:type, og:url, and og:image. Twitter Card fields may specify a card type and preview content. Record the raw property names and values, including duplicates; different consumers may not resolve conflicting values in the same way.
Check whether the URL and image values are absolute and reachable in the context where previews are generated. Compare social titles and images with the page’s visible content, and note when a social value intentionally differs from the HTML title or description.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Parse every JSON-LD block
Find each script whose type is application/ld+json, not just the first one. Parse its text as JSON and retain the full object or array, including @context, @type, @id, URLs, and nested entities. A JSON parse that succeeds only establishes syntactic validity; it does not establish that the types and properties are semantically appropriate or that the claims match the visible page.
Use Schema.org’s vocabulary definitions to check what a type or property means, and its JSON-LD context as the machine-readable vocabulary context. Validate that structured data describes content a reader can actually find on the page. Preserve parse failures and the offending script text in audit output instead of silently dropping them.
Rank #3
Distinguish robots controls from descriptive metadata
Values such as noindex, nofollow, and nosnippet communicate crawling, indexing, or result-presentation rules. They are not substitutes for page descriptions, social fields, or structured data. Extract both page-level robots meta directives and the X-Robots-Tag response header, then report them in a separate part of your result.
A crawler has to be allowed to fetch a page or resource to discover its robots directives. Google explains this distinction in its robots meta tag and X-Robots-Tag documentation. Do not infer that a page is indexable merely because no robots meta tag appears in the HTML you fetched; access restrictions or response headers may matter.
Handle JavaScript-rendered metadata
Compare the raw response with the browser’s rendered DOM when a field visible in the browser is missing from the downloaded HTML. A client-side application may inject or change tags after JavaScript executes, while a raw HTTP fetch sees only the initial response. Use browser developer tools to inspect the live document head, or choose a service that renders JavaScript when processing pages at scale.
OpenGraph.io documents an endpoint that returns Open Graph, Twitter Card, and HTML metadata, as well as full_render and proxy options for rendered-page use cases: OpenGraph.io documentation and service. Rendering and proxying can help with pages that need browser execution, but they do not remove the need to preserve response details and validate extracted values.
Choose an approach for the size of the job
| Approach | Best fit | Important limitation |
|---|---|---|
| Browser developer tools | One-off inspection and comparison of initial HTML with the live DOM | Manual and not suited to a large URL inventory |
curl plus an HTML parser |
Repeatable extraction from pages whose metadata is in the HTTP response | Does not run JavaScript or interpret metadata semantics by itself |
| Hosted metadata API | Repeated URL inventories or workflows needing a documented extraction endpoint | Confirm rendering behavior and returned fields against the service documentation |
| Browser-rendering service | Pages where metadata is added or altered after JavaScript runs | Rendering adds browser work and still requires validation of the output |
For a one-page question, start with browser tools or a raw fetch. For recurring inventories, automate parsing and preserve failures as data. For client-rendered sites, compare raw and rendered results so your pipeline does not confuse “not in the response” with “not present after load.”
Rank #4
- Brand: Wiley
- Set of 2 Volumes
- A handy two-book set that uniquely combines related technologies Highly visual format and accessible language makes these books highly effective learning tools Perfect for beginning web designers and front-end developers
Or skip the browser setup
For a rendered page, a screenshot can help confirm what a visitor sees alongside the metadata you extracted. ScreenshotNeo is a website screenshot API and MCP server for developers; it can capture a rendered page and offers 63 capture options. One GET request can return a PNG, JPEG, WebP, or PDF. Here is the cURL form:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp
See the ScreenshotNeo API documentation for the request parameters. Cookie banners, newsletter popups, and chat widgets are removed before the shot; each cleanup step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing, with response headers identifying the page verdict and billing status. Its MCP server lets AI agents use screenshot tools. The Free plan includes 1,000 shots per month with no card; paid plans start at $5 for 3,000 shots. Screenshot evidence complements metadata extraction; it does not replace parsing the HTML or validating JSON-LD. Sign up for 1,000 free screenshots a month with no card.
Validate the extracted result
- Keep requested and final URLs, status, content type, retrieval time, and response headers with the extracted fields.
- Resolve relative canonical, alternate, and image URLs carefully; retain the original values too.
- Flag duplicate or conflicting title, description, canonical, social, and robots values rather than silently choosing one.
- Parse every JSON-LD block, record syntax errors, and check vocabulary meaning and alignment with visible page content.
- Compare raw HTML with rendered DOM when expected metadata is absent from the initial response.
- Keep crawler directives separate from descriptive fields and structured data.
Troubleshooting missing or confusing metadata
The browser shows a title, but the fetched HTML does not
The page may set or change its head after JavaScript executes. Inspect the live DOM and compare it with the saved response; use a rendering-capable workflow if that distinction matters across many URLs.
The description, canonical, or social field appears more than once
Do not discard duplicates without recording them. Preserve all values and their document order, then investigate which system generated the markup and which value your downstream consumer uses.
A canonical URL or social image points somewhere unexpected
Check whether the value is relative, whether redirects changed the requested page, and whether the page deliberately identifies another preferred URL. Store both the original value and its resolved destination.
Recommended Free Tools
JSON-LD is present but your parser fails
Retain the script text and report the parse error with its location. Pages can contain multiple JSON-LD scripts or arrays, so make sure the parser processes each block and accepts both object and array roots.
Best Value
Robots instructions seem absent
Inspect the response headers as well as the HTML. Also check whether the page could be fetched at all: a crawler blocked from retrieving a resource cannot discover directives inside it.
Later head metadata seems to be ignored
Check the document’s head structure for invalid elements or markup errors before the affected tags. Google’s valid-head guidance explains why valid head content matters; correct the markup and fetch the response again.
Build a useful metadata record
A reliable extractor should return structured fields rather than a single unlabelled text blob. A practical record includes the URL evidence, core tags, link relations, robots controls, social properties, and an array of parsed JSON-LD blocks, plus warnings and parse errors. Preserve raw values alongside normalized values so later consumers can make their own choices without losing what the page actually returned.
PC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minuteFor example, keep canonical_raw and canonical_resolved separately, and represent duplicate descriptions as an array rather than overwriting one with another. For JSON-LD, retain each block’s parsed value and original text. For rendered extraction, record that rendering was used and which values were absent in the raw response. These details make bulk audits explainable and help distinguish a parser bug from a genuine page difference.
Frequently Asked Questions
Can I extract metadata without running JavaScript?
Yes. A raw HTTP fetch works when the page includes the relevant metadata in its initial HTML. If the fields are injected or changed after load, inspect the rendered DOM or use a rendering-capable workflow.
Does valid JSON-LD mean the page’s structured data is correct?
No. Parsing confirms JSON syntax only. Check the Schema.org type and property meanings and compare the claims with the page’s visible content.
Are robots tags the same thing as metadata like a page description?
No. Robots directives control crawling, indexing, or search-result presentation; they do not provide a descriptive summary or structured data.
Do these 3 things before closing this tab:
1Clear out junk files and repair common Windows errors2Fix the driver behind crashes, sound loss and screen glitches3Repair Windows errors before they cause bigger problemsQuick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




