Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

Use Playwright to load the page, inspect its rendered <head> for Open Graph <meta> elements, and read each element’s content value. Save a screenshot separately if you need a visual artifact: an image of the page does not contain the structured metadata in a reliable, machine-readable form.

How do I extract Open Graph metadata with Playwright?

The example below uses Node.js and Playwright. It waits for the document to load, then waits for og:title to appear before collecting all Open Graph properties in document order. It writes the metadata as JSON and saves a full-page screenshot as a separate file.

  1. Install Playwright: run npm init -y, then npm install playwright and npx playwright install chromium.
  2. Save the script: put the code below in extract-og.js.
  3. Run it: use node extract-og.js https://example.com, replacing the URL with the page to inspect.
const { chromium } = require('playwright');

async function main() {
  const target = process.argv[2];
  if (!target) throw new Error('Usage: node extract-og.js https://example.com');

  const browser = await chromium.launch({ headless: true });
  const page = await browser.newPage({ viewport: { width: 1440, height: 1000 } });

  try {
    const response = await page.goto(target, {
      waitUntil: 'domcontentloaded',
      timeout: 30000
    });

    // Use a meaningful page-specific condition when the page inserts tags client-side.
    await page.locator('meta[property="og:title"]').waitFor({ timeout: 10000 }).catch(() => {});

    const result = await page.evaluate(() => {
      const properties = {};
      for (const meta of document.head.querySelectorAll('meta[property]')) {
        const property = meta.getAttribute('property');
        if (!property.toLowerCase().startsWith('og:')) continue;
        (properties[property] ??= []).push(meta.getAttribute('content') ?? '');
      }

      return {
        pageUrl: location.href,
        documentTitle: document.title,
        extractedAt: new Date().toISOString(),
        properties
      };
    });

    result.httpStatus = response?.status() ?? null;
    result.navigationError = response ? null : 'No HTTP response was returned';
    console.log(JSON.stringify(result, null, 2));

    await page.screenshot({ path: 'page.png', fullPage: true });
  } finally {
    await browser.close();
  }
}

main().catch(error => {
  console.error(error);
  process.exitCode = 1;
});

The script deliberately stores values as ordered arrays rather than overwriting repeated property names. A page can contain multiple og:image entries, for example, and order can affect how a consumer treats conflicting values. Empty or missing content values are represented as empty strings; downstream code can decide whether to discard or report them.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How do I get og:title and og:image after a page loads?

Open Graph values are commonly declared in the document head as tags such as <meta property="og:title" content="Example title">. The property attribute identifies the Open Graph field; content carries its value. See the Open Graph protocol and MDN’s HTML meta reference.

For a page whose framework updates metadata after JavaScript runs, extract from the rendered DOM only after a condition that reflects the page’s actual readiness. The example waits briefly for og:title; if the page uses a different signal, substitute a selector or application-state check known to occur after its metadata update. A timeout is a bound on waiting, not proof that every site has finished rendering.

The collector can also include conventional metadata declared with name, such as a description. Add a second loop if needed:

const named = {};
for (const meta of document.head.querySelectorAll('meta[name]')) {
  const name = meta.getAttribute('name');
  (named[name] ??= []).push(meta.getAttribute('content') ?? '');
}

Keep the Open Graph and conventional metadata maps distinct: they use different identifying attributes, and the fields need not match.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Which Open Graph fields should you collect?

The protocol identifies og:title, og:type, og:image, and og:url as core properties. A practical extractor often also retains og:description, og:site_name, and og:locale. Image-related structured properties include og:image:secure_url, og:image:type, og:image:width, og:image:height, and og:image:alt; the protocol recommends an alt description when an image is specified.

Do not flatten image metadata into unrelated fields if a consumer needs to associate dimensions, type, or alt text with a particular image. Structured image properties belong to the preceding image root property; a new image root begins the next group. Preserve sequence, then apply the protocol’s grouping rules when building a richer object.

Rank #2
Sale
HTML and CSS: Design and Build Websites
  • HTML CSS Design and Build Web Sites
  • Comes with secure packaging
  • It can be a gift option

When repeated properties conflict, the protocol gives preference to the first property from top to bottom. Keeping all values lets your application preserve that order, inspect alternate images, or apply a different downstream policy without losing source information.

How should you handle relative image URLs?

A page may return an image value that is relative rather than an absolute URL. If your application needs a fetchable URL, resolve it against the final document URL and retain the original string for debugging. This is an implementation choice, not a URL-resolution rule specified by the Open Graph protocol. Record the final location after redirects rather than assuming the requested URL is still the document URL.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Do not silently discard malformed values or assume every page supplies every core property. Return the values that are present, and consider including navigation and parsing errors alongside the output so a missing field is distinguishable from a failed load.

How do I take a screenshot of a page and read its meta tags?

Read the DOM and capture the image as separate operations, as in the script above. Playwright supports a viewport screenshot, a full-page screenshot, and an element-specific screenshot. It can write the image to a path or return image bytes for further processing. The official Playwright screenshot documentation describes these capture modes.

  • Viewport: omit fullPage to capture the current viewport.
  • Full page: set fullPage: true to capture the scrollable page.
  • One element: use await page.locator('selector').screenshot({ path: 'element.png' }).
  • Image bytes: call const bytes = await page.screenshot() and pass the returned buffer to your image-processing code.

A screenshot helps inspect visual layout, but it cannot substitute for reading metadata attributes. For repeatable captures, record browser, viewport, and relevant page settings. Do not assume screenshots will be pixel-identical across machines; no cross-platform identity guarantee is established here.

Choose a readiness condition that fits the page

Playwright navigation supports commit, domcontentloaded, load, and networkidle wait conditions. They represent different milestones; none guarantees that a particular application has inserted its final metadata. Playwright’s Page API marks networkidle as discouraged for testing and advises using assertions to assess readiness. For client-rendered tags, wait for the expected meta element or another meaningful page-specific signal.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Prefer the navigation call’s waitUntil option over deprecated page.waitForNavigation patterns. The Page API’s warning that “This method is inherently racy, please use page.waitForURL() instead” is specifically about the deprecated page.waitForNavigation method; it is not a general warning against all navigation waits. See the Playwright Page API.

Or skip the browser setup

For a one-request screenshot, ScreenshotNeo returns an image or PDF, but the screenshot itself is not a structured Open Graph extraction result. Keep that distinction in mind if your application needs both outputs.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

See the ScreenshotNeo API documentation for request options. ScreenshotNeo removes cookie banners, newsletter popups, and chat widgets before the shot; each of those steps can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing status. Its MCP server offers take_screenshot, get_page_info, and capture_pdf for AI agents. The free plan includes 1,000 shots per month without a card; paid plans start at $5 for 3,000 shots.

Sign up for ScreenshotNeo’s free plan to get 1,000 screenshots a month with no card.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Troubleshooting extraction and capture

No Open Graph values appear

Check whether the page actually has meta[property^="og:"] elements in its rendered head. It may omit Open Graph metadata, add it only under a later application state, or expose conventional fields using name instead. Inspect the final DOM and adjust the readiness condition to match the page.

Rank #4
Sale
Web Design with HTML, CSS, JavaScript and jQuery Set
  • Brand: Wiley
  • Set of 2 Volumes
  • A handy two-book set that uniquely combines related technologies Highly visual format and accessible language makes these books highly effective learning tools Perfect for beginning web designers and front-end developers

The script times out waiting for og:title

The selector may not exist, or the site may use another metadata field. Treat the wait as an optional readiness signal when pages vary: log that it timed out, continue collecting what is present, and distinguish an absent field from a navigation failure. Increase the timeout only when the target’s behavior justifies it.

Navigation returns no response or an error status

Check the URL, network access, redirects, and target availability. A missing response is not evidence that metadata was successfully read. When a response exists, retain its status alongside the extracted values so callers can make an explicit decision about non-success responses.

Repeated image tags disappear in the output

Use arrays keyed by property name rather than a single-value object assignment. Preserve document order and keep image structured properties grouped with their preceding image root when transforming the raw map.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The screenshot is blank or incomplete

Try a readiness condition tied to visible page content, wait for a specific element, or add a measured delay for a known animation or deferred asset. Avoid treating networkidle as a universal fix: pages with continuing network activity may not reach it, while a quiet network does not necessarily mean the desired content is ready.

The image file is too large or captures the wrong area

Choose viewport, full-page, or locator capture according to the output’s purpose. Full-page images can include substantially more content than viewport images; element screenshots constrain capture to a selected region. Use the output format and path that suit your pipeline, and handle screenshot bytes directly if you do not need a file.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Performance, reliability, and output design

A browser render costs more work than parsing initial HTML, but it can observe metadata added or changed by client-side code. If the initial response already contains the tags and you do not need a screenshot, browser automation may be unnecessary; when you need the rendered state or visual evidence, it keeps those outputs in one browser session.

For batch jobs, set explicit navigation and selector timeouts, close pages and browsers in cleanup paths, and report failures per URL rather than abandoning an entire batch. Store the requested URL and final page URL, extraction timestamp, response status, ordered values, and any navigation or readiness error. This makes redirects, changing metadata, and partial results diagnosable. Do not publish a universal timeout or speed expectation: site behavior varies, and no performance benchmark is established here.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Playwright is the DIY choice when you need to control browser state, inspect the DOM, or capture a particular element. ScreenshotNeo is a separate screenshot service; use it for the screenshot artifact rather than assuming it returns the ordered Open Graph map shown in the Playwright example.

Frequently Asked Questions

Does a screenshot contain Open Graph metadata?

No. A screenshot is a visual capture; read the document’s meta elements to obtain structured Open Graph values.

Should I keep multiple og:image values?

Yes, when downstream behavior may need image alternatives or structured image properties. Preserve their document order.

Can I use this approach for pages without JavaScript?

Yes. Playwright can read metadata already present in the document head; choose the least costly method that meets your needs.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.