Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

Metascraper extracts and normalizes fields such as a page’s title, description, image, author and publication date—but it does not fetch the page for you. Supply both the target URL and its HTML, configure the property bundles you need, then pass the result to Metascraper. Use a browser-rendered HTML source only when a normal HTTP response does not contain the metadata you need.

What Metascraper does—and what it does not do

Metascraper is a Node.js library for turning a page’s metadata signals into a unified object. Its documented sources include Open Graph, regular HTML metadata, Microdata, RDFa, Twitter Cards and JSON-LD, with additional rule bundles for other content and services. The maintainers describe it as “A library to easily extract unified metadata from websites using Open Graph, Microdata, RDFa, Twitter Cards, JSON-LD, HTML, and more.” Metascraper project documentation.

The distinction that determines your implementation is that Metascraper needs two inputs: the target URL and the HTML markup behind it. It parses markup; a separate HTTP client or browser retrieves it. The URL helps resolve relative links and can be used as a fallback by some rules. A screenshot, by contrast, is a visual record of a page, not a substitute for its HTML metadata.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Install Metascraper and its property bundles

Install the core library and only the bundles for fields you plan to extract. The official example uses CommonJS and these packages:

npm install metascraper metascraper-author metascraper-date metascraper-description metascraper-image metascraper-logo metascraper-publisher metascraper-title metascraper-url html-get browserless

The example below follows the project’s documented html-get and browserless pattern. It is most useful when the target requires browser-rendered HTML; for a page whose HTTP response already contains accurate metadata, use a simpler HTTP retrieval method instead. This is an adaptation of the documentation example, not an assertion that it has been independently executed here. Check the project’s current README for package and API changes before deploying: Metascraper documentation.

Extract title, description, image, author and date

Create a file such as extract.js and run it with Node.js. Replace the example URL with a page you are allowed to access.

const getHTML = require('html-get')
const browserless = require('browserless')()
const metascraper = require('metascraper')([
  require('metascraper-author')(),
  require('metascraper-date')(),
  require('metascraper-description')(),
  require('metascraper-image')(),
  require('metascraper-logo')(),
  require('metascraper-publisher')(),
  require('metascraper-title')(),
  require('metascraper-url')()
])

async function getContent(url) {
  const browserContext = browserless.createContext()
  try {
    return await getHTML(url, { getBrowserless: () => browserContext })
  } finally {
    await browserContext.destroyContext()
  }
}

async function main() {
  const url = 'https://example.com'
  try {
    const html = await getContent(url)
    const metadata = await metascraper({ url, html })
    console.log(JSON.stringify(metadata, null, 2))
  } finally {
    await browserless.close()
  }
}

main().catch(error => {
  console.error(error)
  process.exitCode = 1
})

The extraction call is metascraper({ url, html }). Its output is an object whose properties depend on the configured bundles and the metadata available in the page. The example configures author, date, description, image, logo, publisher, title and URL; it does not guarantee that every page supplies every value.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use an HTTP response when it is sufficient

Do not launch a browser by default just because the target is a website. If a regular request returns the page HTML with the relevant metadata in it, pass that HTML along with the target URL to Metascraper. This is generally the lighter retrieval path. The project’s example uses a browser context to handle cases where rendered content is needed, not to claim that all pages require one.

Use browser-rendered HTML when markup differs

Some sites populate or alter metadata with JavaScript. A plain HTTP response may therefore be missing a field that appears in the browser, or may represent a different page state. In that case retrieve browser-rendered HTML and parse that markup. Browser-based retrieval adds operational work—browser processes, memory, latency and cleanup—so apply it to targets that need it rather than indiscriminately.

Choose bundles and fields deliberately

Metascraper’s architecture is modular: property bundles provide rules for particular output fields. The README lists author, date, description, image, language, logo, publisher, title, URL, audio and video, as well as bundles for citation metadata, feeds, readability, media providers, manifests and vendor-specific sources such as Amazon, Instagram, Reddit, Spotify, TikTok, X and YouTube. See the project documentation for the current bundle catalog.

For a link preview, title, description, image and URL may be enough. For an article record, author and date may matter too. Install and configure only what your application uses, and treat absent properties as possible output rather than an exceptional condition: pages can omit metadata or expose incomplete signals.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Limit output to selected properties

The API accepts pickPropNames to run selected properties. For example, once you have an HTML string in html, you can request only the three fields needed for a card:

const metadata = await metascraper({
  url: 'https://example.com/article',
  html,
  pickPropNames: new Set(['title', 'description', 'image'])
})

The documented controls also include omitPropNames, htmlDom, rules and validateUrl. The README states that pickPropNames takes precedence over omitPropNames. URL validation defaults to true and checks WHATWG URL compliance. Use the project documentation for the exact current call shape when adding less common controls.

Understand fallback and conflicting tags

Pages often publish overlapping or inconsistent values: an Open Graph title may differ from the HTML title, or multiple tags may point to different images. Metascraper’s rules run from more specific to more generic; the first successful rule supplies the value and later rules act as fallbacks. You can add custom bundles or pass additional rules at execution time. This is why a configured fallback chain is useful: the extractor can still find a candidate when a preferred signal is absent.

Do not interpret a normalized value as proof that the page’s metadata is objectively correct. It is the result of the configured rules and the markup that was supplied. If the distinction matters—for editorial review, ingestion audits or debugging—store the source URL and, where appropriate, preserve the raw HTML or relevant original tag values alongside the normalized result. That gives you a way to explain why a candidate was selected.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

When a field is absent or wrong

  • Confirm the HTML passed to Metascraper is from the intended URL and represents the content shown to users.
  • Inspect the markup for the relevant Open Graph, HTML, JSON-LD or other metadata signal; do not assume every page publishes all fields.
  • If the signal appears only after JavaScript runs, switch the retrieval stage to browser-rendered HTML.
  • If the page has a site-specific convention, add or adjust rules rather than assuming a generic fallback will understand it.
  • For values that must be auditable, retain the page URL and record which source signal your own application accepted.

Retrieval, scale and accuracy limits

Parsing is only one part of a production scraper. Your fetch layer must handle timeouts, redirects, response errors, content types and site access policies; browser-based retrieval additionally needs reliable context cleanup and resource limits. At larger scale, headless browsers, proxies, anti-bot workarounds, paywalls and restricted platforms can become operational burdens. Metascraper’s documentation points to the managed Microlink API as an option for those needs and describes it as pay-as-you-go and starting free. Check the live Microlink service for current prices, quotas, regional availability and terms rather than relying on an older description: Metascraper project documentation.

The Metascraper README reports a Microlink benchmark of 95.54% correct, 1.79% incorrect and 2.68% missed. The README does not state the benchmark year, dataset or methodology, so these are project-reported figures, not a universal accuracy guarantee or a prediction for your URLs. Measure against representative pages from your own target sites before treating extracted fields as authoritative.

Troubleshooting common extraction problems

All fields are empty or unexpectedly missing

Check that the fetch succeeded and that the HTML is nonempty, then inspect whether the desired tags exist in that exact markup. If the page fills them in through JavaScript, retrieve rendered HTML. Also verify that you installed and configured the bundle for the property you expect.

The output reflects a different page than the one requested

Inspect redirects and the final document returned by your retrieval layer. Pass the intended target URL to Metascraper so URL-based resolution and fallbacks use the correct context. If your application needs the final redirected URL, retain it separately from the original request URL.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A relative image URL does not point where expected

Pass the page URL with the HTML. The URL is used to resolve relative links; passing an incorrect URL can therefore affect the resolved result.

Browser retrieval hangs or consumes too many resources

Keep browser use scoped to pages that need rendering, ensure contexts are destroyed and the browser is closed even on errors, and apply timeouts and concurrency limits in your retrieval layer. The documented example demonstrates browser-context creation and cleanup, but operational limits depend on your deployment and are not specified as universal Metascraper settings.

URL validation rejects an input

The documented validateUrl setting defaults to true and checks WHATWG URL compliance. Confirm that the input is a valid absolute URL in the format your application expects. Change validation behavior only if you understand the consequences for untrusted or malformed inputs.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup

If your goal is a visual capture rather than parsed metadata—or you need browser capture infrastructure without wiring it up yourself—ScreenshotNeo provides a website screenshot API and MCP server. It returns PNG, JPEG or WebP screenshots, or PDFs. A screenshot does not replace Metascraper’s HTML parsing; it is useful when the required output is the rendered page itself.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

One GET request can capture a URL. The API can return a WebP file like this:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

See the ScreenshotNeo API documentation for the request options. Cookie and consent banners are accepted like a visitor and more than 60 known consent platforms, newsletter popups and chat widgets can be removed before capture; each of those steps can be turned off. Bot checks, blank pages, timeouts, failed loads and cache hits are not billed, and response headers report the page verdict and billing status. Its MCP server exposes screenshot, page-info and PDF tools to AI agents. The free plan includes 1,000 shots per month without a card; paid plans start at $5 for 3,000 shots.

Sign up free for 1,000 screenshots a month—no card required.

Frequently Asked Questions

Does Metascraper fetch the URL itself?

No. Your application retrieves the HTML and supplies it with the target URL.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Can Metascraper extract metadata from JavaScript-rendered pages?

It can parse browser-rendered HTML when your retrieval layer supplies that markup; it does not itself provide the browser retrieval step.

Does a screenshot provide the same data as metadata extraction?

No. A screenshot records the rendered appearance; Metascraper parses HTML signals into metadata fields.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.