The Tool Desk
Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
Metascraper extracts and normalizes fields such as a page’s title, description, image, author and publication date—but it does not fetch the page for you. Supply both the target URL and its HTML, configure the property bundles you need, then pass the result to Metascraper. Use a browser-rendered HTML source only when a normal HTTP response does not contain the metadata you need.
What Metascraper does—and what it does not do
Metascraper is a Node.js library for turning a page’s metadata signals into a unified object. Its documented sources include Open Graph, regular HTML metadata, Microdata, RDFa, Twitter Cards and JSON-LD, with additional rule bundles for other content and services. The maintainers describe it as “A library to easily extract unified metadata from websites using Open Graph, Microdata, RDFa, Twitter Cards, JSON-LD, HTML, and more.” Metascraper project documentation.
The distinction that determines your implementation is that Metascraper needs two inputs: the target URL and the HTML markup behind it. It parses markup; a separate HTTP client or browser retrieves it. The URL helps resolve relative links and can be used as a fallback by some rules. A screenshot, by contrast, is a visual record of a page, not a substitute for its HTML metadata.
Install Metascraper and its property bundles
Install the core library and only the bundles for fields you plan to extract. The official example uses CommonJS and these packages:
#1 Best Overall
npm install metascraper metascraper-author metascraper-date metascraper-description metascraper-image metascraper-logo metascraper-publisher metascraper-title metascraper-url html-get browserless
The example below follows the project’s documented html-get and browserless pattern. It is most useful when the target requires browser-rendered HTML; for a page whose HTTP response already contains accurate metadata, use a simpler HTTP retrieval method instead. This is an adaptation of the documentation example, not an assertion that it has been independently executed here. Check the project’s current README for package and API changes before deploying: Metascraper documentation.
Extract title, description, image, author and date
Create a file such as extract.js and run it with Node.js. Replace the example URL with a page you are allowed to access.
const getHTML = require('html-get')
const browserless = require('browserless')()
const metascraper = require('metascraper')([
require('metascraper-author')(),
require('metascraper-date')(),
require('metascraper-description')(),
require('metascraper-image')(),
require('metascraper-logo')(),
require('metascraper-publisher')(),
require('metascraper-title')(),
require('metascraper-url')()
])
async function getContent(url) {
const browserContext = browserless.createContext()
try {
return await getHTML(url, { getBrowserless: () => browserContext })
} finally {
await browserContext.destroyContext()
}
}
async function main() {
const url = 'https://example.com'
try {
const html = await getContent(url)
const metadata = await metascraper({ url, html })
console.log(JSON.stringify(metadata, null, 2))
} finally {
await browserless.close()
}
}
main().catch(error => {
console.error(error)
process.exitCode = 1
})
The extraction call is metascraper({ url, html }). Its output is an object whose properties depend on the configured bundles and the metadata available in the page. The example configures author, date, description, image, logo, publisher, title and URL; it does not guarantee that every page supplies every value.
Recommended Free Tools
Use an HTTP response when it is sufficient
Do not launch a browser by default just because the target is a website. If a regular request returns the page HTML with the relevant metadata in it, pass that HTML along with the target URL to Metascraper. This is generally the lighter retrieval path. The project’s example uses a browser context to handle cases where rendered content is needed, not to claim that all pages require one.
Use browser-rendered HTML when markup differs
Some sites populate or alter metadata with JavaScript. A plain HTTP response may therefore be missing a field that appears in the browser, or may represent a different page state. In that case retrieve browser-rendered HTML and parse that markup. Browser-based retrieval adds operational work—browser processes, memory, latency and cleanup—so apply it to targets that need it rather than indiscriminately.
Rank #2
Choose bundles and fields deliberately
Metascraper’s architecture is modular: property bundles provide rules for particular output fields. The README lists author, date, description, image, language, logo, publisher, title, URL, audio and video, as well as bundles for citation metadata, feeds, readability, media providers, manifests and vendor-specific sources such as Amazon, Instagram, Reddit, Spotify, TikTok, X and YouTube. See the project documentation for the current bundle catalog.
For a link preview, title, description, image and URL may be enough. For an article record, author and date may matter too. Install and configure only what your application uses, and treat absent properties as possible output rather than an exceptional condition: pages can omit metadata or expose incomplete signals.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Limit output to selected properties
The API accepts pickPropNames to run selected properties. For example, once you have an HTML string in html, you can request only the three fields needed for a card:
const metadata = await metascraper({
url: 'https://example.com/article',
html,
pickPropNames: new Set(['title', 'description', 'image'])
})
The documented controls also include omitPropNames, htmlDom, rules and validateUrl. The README states that pickPropNames takes precedence over omitPropNames. URL validation defaults to true and checks WHATWG URL compliance. Use the project documentation for the exact current call shape when adding less common controls.
Understand fallback and conflicting tags
Pages often publish overlapping or inconsistent values: an Open Graph title may differ from the HTML title, or multiple tags may point to different images. Metascraper’s rules run from more specific to more generic; the first successful rule supplies the value and later rules act as fallbacks. You can add custom bundles or pass additional rules at execution time. This is why a configured fallback chain is useful: the extractor can still find a candidate when a preferred signal is absent.
Rank #3
Do not interpret a normalized value as proof that the page’s metadata is objectively correct. It is the result of the configured rules and the markup that was supplied. If the distinction matters—for editorial review, ingestion audits or debugging—store the source URL and, where appropriate, preserve the raw HTML or relevant original tag values alongside the normalized result. That gives you a way to explain why a candidate was selected.
Windows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallOutdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchWhen a field is absent or wrong
- Confirm the HTML passed to Metascraper is from the intended URL and represents the content shown to users.
- Inspect the markup for the relevant Open Graph, HTML, JSON-LD or other metadata signal; do not assume every page publishes all fields.
- If the signal appears only after JavaScript runs, switch the retrieval stage to browser-rendered HTML.
- If the page has a site-specific convention, add or adjust rules rather than assuming a generic fallback will understand it.
- For values that must be auditable, retain the page URL and record which source signal your own application accepted.
Retrieval, scale and accuracy limits
Parsing is only one part of a production scraper. Your fetch layer must handle timeouts, redirects, response errors, content types and site access policies; browser-based retrieval additionally needs reliable context cleanup and resource limits. At larger scale, headless browsers, proxies, anti-bot workarounds, paywalls and restricted platforms can become operational burdens. Metascraper’s documentation points to the managed Microlink API as an option for those needs and describes it as pay-as-you-go and starting free. Check the live Microlink service for current prices, quotas, regional availability and terms rather than relying on an older description: Metascraper project documentation.
The Metascraper README reports a Microlink benchmark of 95.54% correct, 1.79% incorrect and 2.68% missed. The README does not state the benchmark year, dataset or methodology, so these are project-reported figures, not a universal accuracy guarantee or a prediction for your URLs. Measure against representative pages from your own target sites before treating extracted fields as authoritative.
Troubleshooting common extraction problems
All fields are empty or unexpectedly missing
Check that the fetch succeeded and that the HTML is nonempty, then inspect whether the desired tags exist in that exact markup. If the page fills them in through JavaScript, retrieve rendered HTML. Also verify that you installed and configured the bundle for the property you expect.
The output reflects a different page than the one requested
Inspect redirects and the final document returned by your retrieval layer. Pass the intended target URL to Metascraper so URL-based resolution and fallbacks use the correct context. If your application needs the final redirected URL, retain it separately from the original request URL.
Rank #4
A relative image URL does not point where expected
Pass the page URL with the HTML. The URL is used to resolve relative links; passing an incorrect URL can therefore affect the resolved result.
Browser retrieval hangs or consumes too many resources
Keep browser use scoped to pages that need rendering, ensure contexts are destroyed and the browser is closed even on errors, and apply timeouts and concurrency limits in your retrieval layer. The documented example demonstrates browser-context creation and cleanup, but operational limits depend on your deployment and are not specified as universal Metascraper settings.
URL validation rejects an input
The documented validateUrl setting defaults to true and checks WHATWG URL compliance. Confirm that the input is a valid absolute URL in the format your application expects. Change validation behavior only if you understand the consequences for untrusted or malformed inputs.
Or skip the browser setup
If your goal is a visual capture rather than parsed metadata—or you need browser capture infrastructure without wiring it up yourself—ScreenshotNeo provides a website screenshot API and MCP server. It returns PNG, JPEG or WebP screenshots, or PDFs. A screenshot does not replace Metascraper’s HTML parsing; it is useful when the required output is the rendered page itself.
One GET request can capture a URL. The API can return a WebP file like this:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
See the ScreenshotNeo API documentation for the request options. Cookie and consent banners are accepted like a visitor and more than 60 known consent platforms, newsletter popups and chat widgets can be removed before capture; each of those steps can be turned off. Bot checks, blank pages, timeouts, failed loads and cache hits are not billed, and response headers report the page verdict and billing status. Its MCP server exposes screenshot, page-info and PDF tools to AI agents. The free plan includes 1,000 shots per month without a card; paid plans start at $5 for 3,000 shots.
Sign up free for 1,000 screenshots a month—no card required.
Frequently Asked Questions
Does Metascraper fetch the URL itself?
No. Your application retrieves the HTML and supplies it with the target URL.
Can Metascraper extract metadata from JavaScript-rendered pages?
It can parse browser-rendered HTML when your retrieval layer supplies that markup; it does not itself provide the browser retrieval step.
Does a screenshot provide the same data as metadata extraction?
No. A screenshot records the rendered appearance; Metascraper parses HTML signals into metadata fields.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

