Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use DOMXPath and the XPath union operator (|) when you need several HTML tag names in one result: //h1 | //h2 | //p. PHP evaluates the expression against a DOMDocument and returns a DOMNodeList in document order. For a one-off tag, getElementsByTagName() remains simpler, but it accepts only one tag name per call.

Complete PHP example: select h1, h2 and p elements

This runnable example parses an HTML string, executes one XPath query, checks for a malformed expression, and prints each matching element:

<?php
$html = <<<'HTML'
<!doctype html>
<html><body>
  <h1>Page title</h1>
  <p>Intro</p>
  <h2>Section</h2>
</body></html>
HTML;

$doc = new DOMDocument();
libxml_use_internal_errors(true);
$doc->loadHTML($html);
libxml_clear_errors();

$xpath = new DOMXPath($doc);
$nodes = $xpath->query('//h1 | //h2 | //p');

if ($nodes === false) {
    throw new RuntimeException('Invalid XPath expression');
}

foreach ($nodes as $node) {
    echo $node->nodeName . ': ' . trim($node->textContent) . PHP_EOL;
}

The output is:

h1: Page title
p: Intro
h2: Section

DOMXPath provides XPath 1.0 queries over HTML or XML. Its query() method yields a DOMNodeList for node results and returns false when the expression or context node is invalid. Always test the result before iterating.

How the XPath union works

Fixed list of tag names

Separate location paths with |:

//h1 | //h2 | //p

Each path means “find descendants anywhere in the document.” The union combines the matches and removes duplicates, with the resulting nodes presented in document order. This is generally the clearest solution when the tag list is known in advance.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A single tag predicate

You can express the same selection with a wildcard and a name test:

//*[self::h1 or self::h2 or self::p]

This form is useful when you will add a shared predicate. For example, to keep only headings carrying a particular class:

//*[self::h1 or self::h2][@class='article-heading']

The union form is usually easier to read for a short, fixed list; the predicate form scales better when the condition is generated or shared.

Limit the search to a container

To avoid matching navigation, footers, or unrelated markup, anchor the query at a known container:

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
//main//*[self::h1 or self::h2 or self::p]

If you already have a context element, use a relative descendant path beginning with a dot:

.//h1 | .//h2 | .//p

Without the dot, //h1 is evaluated from the document root even when a context node is supplied. This is a frequent reason a query appears to ignore the element you passed.

Position predicates and parentheses

Parentheses determine whether a positional predicate applies to the complete union or to each branch:

(//h1 | //h2)[1]

This selects the first matching heading overall. By contrast, //h1[1] | //h2[1] selects the first h1 and the first h2 separately.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Using getElementsByTagName() instead

DOMDocument::getElementsByTagName() accepts one local tag name. There is no comma-separated multi-tag syntax, so call it once per name:

$wanted = ['h1', 'h2', 'p'];

foreach ($wanted as $tag) {
    $nodes = $doc->getElementsByTagName($tag);
    foreach ($nodes as $node) {
        echo $node->nodeName . ': ' . trim($node->textContent) . PHP_EOL;
    }
}

This is straightforward for independent processing. It performs a separate lookup for every name, however, and the output is grouped by tag rather than naturally interleaved in document order. If you need one ordered collection, collect the nodes and sort them by document position, or use one XPath union instead.

Choosing between XPath and multiple DOM lookups

Need Best fit Reason
One tag, simple traversal getElementsByTagName('p') Smallest and most direct API.
Several fixed tags in document order XPath union One expression, one traversal result, no manual merge.
Shared class, attribute, text, or ancestor condition XPath Predicates and ancestry are built into the query.
Different logic for each tag Separate tag lookups Each result can be handled independently without a complex expression.
Namespace-aware XML/XHTML Namespace-registered XPath Prefixes make namespace matching explicit.

For normal HTML parsed by loadHTML(), the union expression is the most readable multi-tag solution. The practical performance difference is usually less important than avoiding accidental matches and keeping the selection understandable.

Parsing HTML safely before querying

Suppress parser warnings for imperfect fragments

Real-world HTML fragments are often missing closing tags or document wrappers. Surround loadHTML() with libxml’s internal-error mode when warnings should be handled by your application:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
$previous = libxml_use_internal_errors(true);
try {
    if ($doc->loadHTML($html) === false) {
        throw new RuntimeException('HTML could not be parsed');
    }
} finally {
    libxml_clear_errors();
    libxml_use_internal_errors($previous);
}

This prevents parser warnings from being emitted, but it does not silently make incorrect markup correct. Validate or sanitize untrusted input according to your application’s security requirements.

HTML name casing

After HTML parsing, element and attribute names are matched in lower case. Query //h1, not //H1. The same applies to attribute names in predicates such as [@class='article-heading'].

XML and XHTML namespaces

Namespace-aware XML is different from ordinary HTML. Register the document’s namespace with a prefix and use that prefix in every path:

$doc = new DOMDocument();
$doc->loadXML($xml);
$xpath = new DOMXPath($doc);
$xpath->registerNamespace('xhtml', 'http://www.w3.org/1999/xhtml');

$nodes = $xpath->query('//xhtml:h1 | //xhtml:h2 | //xhtml:p');
if ($nodes === false) {
    throw new RuntimeException('Invalid namespaced XPath');
}

Using unprefixed //h1 against a namespaced document normally returns no nodes because the unprefixed name means “no namespace.”

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Useful production patterns

Read attributes and normalized text

$nodes = $xpath->query('//h1 | //h2 | //p');
if ($nodes === false) {
    throw new RuntimeException('Invalid XPath expression');
}

foreach ($nodes as $node) {
    $id = $node instanceof DOMElement ? $node->getAttribute('id') : '';
    $text = trim(preg_replace('/s+/', ' ', $node->textContent));
    printf("%s%s: %sn", $node->nodeName, $id !== '' ? "#$id" : '', $text);
}

Match attributes precisely

Use an attribute predicate when every selected tag must carry a value:

//h1[@data-published='true'] | //h2[@data-published='true']

For a class token, do not use a bare equality test if the element may have several classes. A token-safe XPath 1.0 test is:

//*[self::h1 or self::h2][contains(concat(' ', normalize-space(@class), ' '), ' article-heading ')]

Use a compiled expression when querying repeatedly

If the same XPath runs many times, create one DOMXPath object per document and reuse the expression. Avoid constructing a new document or parser for every node; parse once, then issue the required queries.

Troubleshooting an empty or invalid result

  • Empty DOMNodeList: confirm the HTML was actually loaded, inspect the parsed document, and remember that HTML names are lower case.
  • Namespaced XHTML/XML: register the namespace and use its prefix, such as //xhtml:h1.
  • query() returns false: the XPath syntax is malformed or the context node is invalid. Check the expression and test the return value before foreach.
  • Only descendants of a supplied element are wanted: use .//, not //, in the context-relative expression.
  • Unexpected order: multiple getElementsByTagName() calls produce tag-grouped output. Use a union when source order matters.
  • Warnings from loadHTML(): enable libxml internal errors, clear them after parsing, and treat malformed input deliberately rather than assuming it was repaired.
  • Position selects too many nodes: wrap a union in parentheses, for example (//h1 | //h2)[1].

Performance, safety and maintainability notes

  • Scope queries narrowly: //main//p avoids scanning unrelated regions when a container is known.
  • Prefer one expressive query: a union avoids manual merging and preserves document order.
  • Do not trust extracted text: escape it when inserting into HTML output, for example with htmlspecialchars().
  • Control external loading: parse supplied HTML as data and avoid enabling network-dependent features unless your application explicitly needs them.
  • Handle large documents: retain only the fields you need instead of storing every node or full textContent string.
  • Test the parser version you deploy: libxml behavior and available DOM features depend on the PHP/libxml runtime, while the XPath expressions above use XPath 1.0 features supported by DOMXPath.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup

If your PHP workflow ultimately needs an image or PDF of a rendered URL, ScreenshotNeo provides a single HTTP request instead of maintaining a headless-browser setup. It accepts cookie and consent banners before capture and removes more than 60 known consent platforms, newsletter popups and chat widgets; each step can be disabled. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads and cache hits are not billed, and response headers identify the page verdict and billing status.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use the API directly (see the ScreenshotNeo API documentation):

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python and Node.js clients can use the same endpoint:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

ScreenshotNeo also offers an MCP server with take_screenshot, get_page_info and capture_pdf tools for Claude, Cursor and other MCP clients. Every feature is on every plan: 1,000 screenshots per month are free with no card, and paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.

Practical decision checklist

  • Choose XPath union for several tags that should come back in source order.
  • Choose a predicate when tag names share an attribute, text, or ancestor rule.
  • Use a leading dot for descendant searches from a context element.
  • Register namespaces for XML or XHTML.
  • Check for false before iterating over a query result.
  • Use separate getElementsByTagName() calls only when one-name lookups or tag-specific processing are clearer.

Frequently Asked Questions

Can I pass several names to getElementsByTagName(), such as ‘h1,h2,p’?

No. The method accepts one local tag name. Use separate calls or an XPath union such as //h1 | //h2 | //p.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Does an XPath union preserve the order in the original HTML?

Yes. The node-set returned by the union is presented in document order, unlike output concatenated from separate tag lookups.

Why does my query work for HTML but not for XHTML?

XHTML commonly places elements in a namespace. Register that namespace on DOMXPath and query with a prefix such as //xhtml:h1.

What should I do when query() returns false?

Treat it as an XPath or context error, inspect the expression, and handle the failure before iterating. An empty DOMNodeList is different: it means the expression was valid but matched nothing.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.