Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

Use the xpath package with @xmldom/xmldom when the HTML or XML is already available to Node.js. Parse the response into a DOM, then call xpath.select for a collection, select1 for one node, or evaluate when you need a typed XPath result. If the page builds its content with JavaScript, first load it in Playwright or Puppeteer and run XPath in that browser context. The choice of parser, result type, namespaces, frames and shadow roots determines whether a selector works reliably.

Choose the right XPath workflow

There are two fundamentally different scraping cases:

Target page What Node initially receives Recommended approach
Static HTML or XML The response already contains the elements and text Parse with @xmldom/xmldom, query with xpath
JavaScript-rendered HTML An initial shell; data appears after scripts run Use Playwright or Puppeteer, wait for the content, then query in the browser

An HTTP request alone does not execute page JavaScript. Downloading a single-page application and applying XPath to that response can therefore produce zero matches even though a human sees the content in a browser.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Install the XPath engine and DOM parser

For a static document, install both packages:

npm install xpath @xmldom/xmldom

The xpath package implements XPath 1.0 for Node.js. @xmldom/xmldom supplies a DOM that the engine can search. In an ES module, the smallest complete example is:

import xpath from 'xpath';
import { DOMParser } from '@xmldom/xmldom';

const html = '<article><h1>XPath guide</h1><a href="/docs">Docs</a></article>';
const doc = new DOMParser().parseFromString(html, 'text/html');

const headings = xpath.select('//article//h1', doc);
const href = xpath.select1('//article//a/@href', doc)?.value;

console.log(headings[0]?.textContent); // XPath guide
console.log(href);                     // /docs

If your project uses CommonJS, replace the imports with const xpath = require('xpath') and const { DOMParser } = require('@xmldom/xmldom'), subject to your Node.js module configuration.

Select one node, many nodes, or a typed result

select: collect every match

xpath.select(expression, contextNode) returns matching nodes (or a scalar when the expression itself returns one). Treat the result as a collection when selecting elements:

const links = xpath.select('//article//a', doc);
for (const link of links) {
  console.log({ text: link.textContent.trim(), href: link.getAttribute('href') });
}

Use a relative expression with a node as the context when processing repeated records:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
const cards = xpath.select('//article[contains(@class, "card")]', doc);
const rows = cards.map(card => ({
  title: xpath.select('string(.//h2)', card).trim(),
  url: xpath.select1('.//a/@href', card)?.value ?? null
}));

select1: take the first matching node

select1 is useful when you need one element or attribute. It returns undefined when there is no match, so use optional chaining or an explicit check:

const titleNode = xpath.select1('//article//h1', doc);
if (!titleNode) throw new Error('Article heading was not found');
console.log(titleNode.textContent.trim());

const canonical = xpath.select1('//link[@rel="canonical"]/@href', doc)?.value ?? null;

Scalar functions: extract text without traversing nodes

XPath functions can return a string, number or boolean directly. The package documentation demonstrates string(//title); the same pattern avoids manual text-node handling:

const title = xpath.select('string(//article//h1)', doc).trim();
const price = Number(xpath.select('number(//meta[@property="product:price:amount"]/@content)', doc));
const hasVideo = xpath.select('boolean(//video)', doc);

Be careful with missing values: XPath converts an absent node to an empty string for string() and to NaN for number().

evaluate: request XPathResult-style control

When you need a specific result type or iterator, use the browser-like evaluate signature:

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
const result = xpath.evaluate(
  '//article//a',
  doc,
  null,
  xpath.XPathResult.ORDERED_NODE_ITERATOR_TYPE,
  null
);

for (let node = result.iterateNext(); node; node = result.iterateNext()) {
  console.log(node.textContent.trim(), node.getAttribute('href'));
}

The arguments are the expression, context node, namespace resolver, requested result type and an optional reusable result object. This is preferable to coercing a large result into an array when you want controlled iteration.

Build selectors that survive markup changes

Prefer semantic anchors

Start with a short selector tied to an element name, stable attribute or meaningful text:

//main[@id="content"]//article//h2
//a[@data-testid="product-link"]/@href
//button[normalize-space(.)="Next"]

Avoid generated class names and absolute paths such as /html/body/div[2]/div[1]/.... Those expressions encode the current implementation rather than the content you want and tend to break when a frontend inserts a wrapper.

Normalize text and whitespace

Use normalize-space when indentation or nested spans make exact text matching unreliable:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
//h2[normalize-space(.)="Shipping details"]

For case-insensitive matching in XPath 1.0, translate ASCII capitals to lowercase:

//a[translate(normalize-space(.), 'ABCDEFGHIJKLMNOPQRSTUVWXYZ', 'abcdefghijklmnopqrstuvwxyz')='read more']

Check the match before extraction

During development, log the expression, count and a short text sample. A zero count can indicate client rendering, a frame, a namespace, a malformed parse or a shadow root—not only a typo.

function inspect(expression, context = doc) {
  const nodes = xpath.select(expression, context);
  console.log({ expression, count: nodes.length,
    sample: nodes.slice(0, 3).map(node => node.textContent.trim().slice(0, 120)) });
  return nodes;
}

const products = inspect('//article[@data-product]');

Handle XML and HTML namespaces

Namespace-qualified XML requires a resolver. Bind a convenient prefix to the document’s namespace URI, then use that prefix in the expression:

const select = xpath.useNamespaces({ book: 'http://example.com/book' });
const titles = select('//book:title/text()', doc);
console.log(titles.map(node => node.data));

The prefix in your XPath does not have to match the source document’s prefix; the URI mapping is what matters. A default XML namespace still needs an explicit prefix in your expression.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

When the namespace prefix is unknown

If incoming documents use unpredictable prefixes, test the local name and namespace URI instead:

const titles = xpath.select(
  '//*[local-name()="title" and namespace-uri()="http://example.com/book"]',
  doc
);

This fallback is more tolerant of prefix changes, but it is less specific than a bound prefix and can match unintended elements if the namespace check is omitted.

Scrape JavaScript-rendered pages with Playwright

Use a browser automation context when the required nodes are created after scripts run. Install Playwright and launch a browser:

npm install playwright
import { chromium } from 'playwright';

const browser = await chromium.launch();
const page = await browser.newPage();
try {
  await page.goto('https://example.com/catalog', { waitUntil: 'domcontentloaded' });
  await page.locator('//article//h2').first().waitFor();
  const headings = await page.locator('xpath=//article//h2').allTextContents();
  console.log(headings.map(text => text.trim()));
} finally {
  await browser.close();
}

Playwright supports CSS and XPath through page.locator(). Strings beginning with // or .. are also auto-detected as XPath, although the explicit xpath= prefix makes intent clear. Locators wait for elements according to Playwright’s normal action and assertion rules; still wait for a meaningful application state when data arrives asynchronously.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Frames

XPath runs in the document you query. If the target is inside an iframe, select the frame first and query its document rather than the top-level page:

const frame = page.frame({ name: 'results' });
if (!frame) throw new Error('Results frame is unavailable');
const text = await frame.locator('xpath=//article//h2').allTextContents();

Shadow DOM

Playwright’s XPath locator does not pierce shadow roots. For an open shadow root, use a supported locator strategy or enter the relevant shadow root before applying a selector. A correct XPath against the light DOM will otherwise return zero nodes.

Use Puppeteer when it is your browser stack

Puppeteer evaluates XPath with the browser’s native Document.evaluate. Its selector syntax uses the ::-p-xpath() prefix:

import puppeteer from 'puppeteer';

const browser = await puppeteer.launch();
const page = await browser.newPage();
try {
  await page.goto('https://example.com/catalog', { waitUntil: 'networkidle0' });
  const heading = await page.waitForSelector('::-p-xpath(//article//h2)');
  console.log(await heading.evaluate(el => el.textContent.trim()));
} finally {
  await browser.close();
}

Do not copy Playwright’s page.locator syntax into Puppeteer or vice versa. Both use XPath, but the APIs and waiting behavior differ.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Parse fetched HTML safely

For a static endpoint, fetch the response, check the HTTP status and parse the body before selecting:

import xpath from 'xpath';
import { DOMParser } from '@xmldom/xmldom';

const response = await fetch('https://example.com/article');
if (!response.ok) throw new Error(`HTTP ${response.status}`);
const html = await response.text();
const doc = new DOMParser().parseFromString(html, 'text/html');
const title = xpath.select('string(//main//h1)', doc).trim();
if (!title) throw new Error('No title in server response');
console.log(title);

Keep network errors, non-2xx responses and parse warnings separate from “selector matched nothing.” That distinction makes retries and selector maintenance much easier.

Performance, reliability and responsible operation

  • Parse a response once and reuse the DOM for all expressions; reparsing for every field wastes CPU.
  • Use one broad query to identify records, then relative queries against each record instead of repeatedly scanning the whole document.
  • For browser scraping, reuse a browser process and create isolated pages or contexts rather than launching a new process for every URL.
  • Wait for a specific selector or application state instead of an arbitrary long sleep; this reduces latency while avoiding early extraction.
  • Set navigation and operation timeouts, close pages in finally blocks, and retry transient network failures with limits.
  • Respect the target site’s terms, robots guidance and rate limits. Avoid bypassing authentication, bot checks or access controls.
  • Cache stable responses where permitted, and record the URL, status, match count and extraction errors so a markup change is visible immediately.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Common failures and fixes

“Cannot find module” or import errors

Install both packages in the project that runs the script. If imports fail, align the code with your module mode: ES modules use import; CommonJS uses require.

Every expression returns zero nodes

Inspect the raw response. If it contains only an application shell, switch to Playwright or Puppeteer. Otherwise check whether the content is in an iframe, namespace-qualified, inside a shadow root, or represented differently by the parser.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The selector worked yesterday

Log the current markup and replace generated classes or positional paths with semantic attributes, stable text or a shorter structural path. Add a fixture test containing the expected elements.

Text contains unexpected whitespace

Use normalize-space(.) in the XPath and trim() in JavaScript. For nested content, select the element’s string value rather than a single text child.

Browser code times out

Confirm that the URL loads, increase the navigation timeout only when justified, and wait for the selector that signals finished rendering. Check for an iframe or a bot-check page before changing the XPath.

Namespace queries fail

Bind the namespace URI with useNamespaces. If prefixes vary between documents, use local-name() together with namespace-uri().

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Or skip the browser setup

If your goal is a clean image or PDF rather than DOM data, ScreenshotNeo accepts a URL and handles the capture in one request. Cookie and consent banners are accepted and removed before the shot, along with more than 60 known consent platforms, newsletter popups and chat widgets. Bot checks, blank pages, timeouts, failed loads and cache hits are not billed, and response headers identify the page verdict and billing status. Its MCP server provides take_screenshot, get_page_info and capture_pdf tools for Claude, Cursor and other MCP clients.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

See the ScreenshotNeo API documentation for options such as full-page or element capture, device and retina settings, waits, custom headers and cookies, JavaScript, blocking rules, PDFs, caching, signed links, asynchronous jobs and bulk requests.

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

The Free plan includes 1,000 screenshots per month with no card. Paid plans start at $5 for 3,000 shots; every feature is available on every plan. Create a free ScreenshotNeo account.

FAQ

Does the Node.js xpath package support XPath 2.0?

No. The package implements XPath 1.0, so use XPath 1.0 functions and syntax in these examples.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Can I use one XPath expression for both static HTML and Playwright?

Usually yes: the expression describes the document. The surrounding APIs differ, so static code uses select while Playwright uses a locator and browser waiting.

Why does a browser show an element that XPath cannot find?

It may be inside a different frame or shadow root, or the browser may have rendered it after the document you queried was captured. Query the correct context after rendering.

Is XPath always better than CSS selectors?

No. XPath is useful for relationships, text conditions and XML namespaces. CSS is often simpler for straightforward element and attribute matches; choose the selector that remains stable for the page you maintain.

Frequently Asked Questions

Does the Node.js xpath package support XPath 2.0?

No. It implements XPath 1.0.

Can one XPath expression be reused in static parsing and browser automation?

Usually yes, but the surrounding APIs and waiting behavior differ.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Why can a browser show an element that XPath cannot find?

The element may be rendered later, inside an iframe, or inside a shadow root.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.