Free tools Windows power users keep installed
One-click scans. No signup required.
Use DOMXPath and the XPath union operator (|) when you need several HTML tag names in one result: //h1 | //h2 | //p. PHP evaluates the expression against a DOMDocument and returns a DOMNodeList in document order. For a one-off tag, getElementsByTagName() remains simpler, but it accepts only one tag name per call.
Complete PHP example: select h1, h2 and p elements
This runnable example parses an HTML string, executes one XPath query, checks for a malformed expression, and prints each matching element:
<?php
$html = <<<'HTML'
<!doctype html>
<html><body>
<h1>Page title</h1>
<p>Intro</p>
<h2>Section</h2>
</body></html>
HTML;
$doc = new DOMDocument();
libxml_use_internal_errors(true);
$doc->loadHTML($html);
libxml_clear_errors();
$xpath = new DOMXPath($doc);
$nodes = $xpath->query('//h1 | //h2 | //p');
if ($nodes === false) {
throw new RuntimeException('Invalid XPath expression');
}
foreach ($nodes as $node) {
echo $node->nodeName . ': ' . trim($node->textContent) . PHP_EOL;
}
The output is:
h1: Page title
p: Intro
h2: Section
DOMXPath provides XPath 1.0 queries over HTML or XML. Its query() method yields a DOMNodeList for node results and returns false when the expression or context node is invalid. Always test the result before iterating.
How the XPath union works
Fixed list of tag names
Separate location paths with |:
//h1 | //h2 | //p
Each path means “find descendants anywhere in the document.” The union combines the matches and removes duplicates, with the resulting nodes presented in document order. This is generally the clearest solution when the tag list is known in advance.
#1 Best Overall
A single tag predicate
You can express the same selection with a wildcard and a name test:
//*[self::h1 or self::h2 or self::p]
This form is useful when you will add a shared predicate. For example, to keep only headings carrying a particular class:
//*[self::h1 or self::h2][@class='article-heading']
The union form is usually easier to read for a short, fixed list; the predicate form scales better when the condition is generated or shared.
Limit the search to a container
To avoid matching navigation, footers, or unrelated markup, anchor the query at a known container:
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
//main//*[self::h1 or self::h2 or self::p]
If you already have a context element, use a relative descendant path beginning with a dot:
Rank #2
.//h1 | .//h2 | .//p
Without the dot, //h1 is evaluated from the document root even when a context node is supplied. This is a frequent reason a query appears to ignore the element you passed.
Position predicates and parentheses
Parentheses determine whether a positional predicate applies to the complete union or to each branch:
(//h1 | //h2)[1]
This selects the first matching heading overall. By contrast, //h1[1] | //h2[1] selects the first h1 and the first h2 separately.
Quick wins for a faster PC:
Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Repair Windows errors before they cause bigger problemsFix Now →Using getElementsByTagName() instead
DOMDocument::getElementsByTagName() accepts one local tag name. There is no comma-separated multi-tag syntax, so call it once per name:
$wanted = ['h1', 'h2', 'p'];
foreach ($wanted as $tag) {
$nodes = $doc->getElementsByTagName($tag);
foreach ($nodes as $node) {
echo $node->nodeName . ': ' . trim($node->textContent) . PHP_EOL;
}
}
This is straightforward for independent processing. It performs a separate lookup for every name, however, and the output is grouped by tag rather than naturally interleaved in document order. If you need one ordered collection, collect the nodes and sort them by document position, or use one XPath union instead.
Choosing between XPath and multiple DOM lookups
| Need | Best fit | Reason |
|---|---|---|
| One tag, simple traversal | getElementsByTagName('p') |
Smallest and most direct API. |
| Several fixed tags in document order | XPath union | One expression, one traversal result, no manual merge. |
| Shared class, attribute, text, or ancestor condition | XPath | Predicates and ancestry are built into the query. |
| Different logic for each tag | Separate tag lookups | Each result can be handled independently without a complex expression. |
| Namespace-aware XML/XHTML | Namespace-registered XPath | Prefixes make namespace matching explicit. |
For normal HTML parsed by loadHTML(), the union expression is the most readable multi-tag solution. The practical performance difference is usually less important than avoiding accidental matches and keeping the selection understandable.
Parsing HTML safely before querying
Suppress parser warnings for imperfect fragments
Real-world HTML fragments are often missing closing tags or document wrappers. Surround loadHTML() with libxml’s internal-error mode when warnings should be handled by your application:
$previous = libxml_use_internal_errors(true);
try {
if ($doc->loadHTML($html) === false) {
throw new RuntimeException('HTML could not be parsed');
}
} finally {
libxml_clear_errors();
libxml_use_internal_errors($previous);
}
This prevents parser warnings from being emitted, but it does not silently make incorrect markup correct. Validate or sanitize untrusted input according to your application’s security requirements.
HTML name casing
After HTML parsing, element and attribute names are matched in lower case. Query //h1, not //H1. The same applies to attribute names in predicates such as [@class='article-heading'].
XML and XHTML namespaces
Namespace-aware XML is different from ordinary HTML. Register the document’s namespace with a prefix and use that prefix in every path:
Rank #4
$doc = new DOMDocument();
$doc->loadXML($xml);
$xpath = new DOMXPath($doc);
$xpath->registerNamespace('xhtml', 'http://www.w3.org/1999/xhtml');
$nodes = $xpath->query('//xhtml:h1 | //xhtml:h2 | //xhtml:p');
if ($nodes === false) {
throw new RuntimeException('Invalid namespaced XPath');
}
Using unprefixed //h1 against a namespaced document normally returns no nodes because the unprefixed name means “no namespace.”
Useful production patterns
Read attributes and normalized text
$nodes = $xpath->query('//h1 | //h2 | //p');
if ($nodes === false) {
throw new RuntimeException('Invalid XPath expression');
}
foreach ($nodes as $node) {
$id = $node instanceof DOMElement ? $node->getAttribute('id') : '';
$text = trim(preg_replace('/s+/', ' ', $node->textContent));
printf("%s%s: %sn", $node->nodeName, $id !== '' ? "#$id" : '', $text);
}
Match attributes precisely
Use an attribute predicate when every selected tag must carry a value:
//h1[@data-published='true'] | //h2[@data-published='true']
For a class token, do not use a bare equality test if the element may have several classes. A token-safe XPath 1.0 test is:
//*[self::h1 or self::h2][contains(concat(' ', normalize-space(@class), ' '), ' article-heading ')]
Use a compiled expression when querying repeatedly
If the same XPath runs many times, create one DOMXPath object per document and reuse the expression. Avoid constructing a new document or parser for every node; parse once, then issue the required queries.
Troubleshooting an empty or invalid result
- Empty
DOMNodeList: confirm the HTML was actually loaded, inspect the parsed document, and remember that HTML names are lower case. - Namespaced XHTML/XML: register the namespace and use its prefix, such as
//xhtml:h1. query()returnsfalse: the XPath syntax is malformed or the context node is invalid. Check the expression and test the return value beforeforeach.- Only descendants of a supplied element are wanted: use
.//, not//, in the context-relative expression. - Unexpected order: multiple
getElementsByTagName()calls produce tag-grouped output. Use a union when source order matters. - Warnings from
loadHTML(): enable libxml internal errors, clear them after parsing, and treat malformed input deliberately rather than assuming it was repaired. - Position selects too many nodes: wrap a union in parentheses, for example
(//h1 | //h2)[1].
Performance, safety and maintainability notes
- Scope queries narrowly:
//main//pavoids scanning unrelated regions when a container is known. - Prefer one expressive query: a union avoids manual merging and preserves document order.
- Do not trust extracted text: escape it when inserting into HTML output, for example with
htmlspecialchars(). - Control external loading: parse supplied HTML as data and avoid enabling network-dependent features unless your application explicitly needs them.
- Handle large documents: retain only the fields you need instead of storing every node or full
textContentstring. - Test the parser version you deploy: libxml behavior and available DOM features depend on the PHP/libxml runtime, while the XPath expressions above use XPath 1.0 features supported by
DOMXPath.
Or skip the browser setup
If your PHP workflow ultimately needs an image or PDF of a rendered URL, ScreenshotNeo provides a single HTTP request instead of maintaining a headless-browser setup. It accepts cookie and consent banners before capture and removes more than 60 known consent platforms, newsletter popups and chat widgets; each step can be disabled. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads and cache hits are not billed, and response headers identify the page verdict and billing status.
Windows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallOutdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchUse the API directly (see the ScreenshotNeo API documentation):
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python and Node.js clients can use the same endpoint:
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
ScreenshotNeo also offers an MCP server with take_screenshot, get_page_info and capture_pdf tools for Claude, Cursor and other MCP clients. Every feature is on every plan: 1,000 screenshots per month are free with no card, and paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.
Practical decision checklist
- Choose XPath union for several tags that should come back in source order.
- Choose a predicate when tag names share an attribute, text, or ancestor rule.
- Use a leading dot for descendant searches from a context element.
- Register namespaces for XML or XHTML.
- Check for
falsebefore iterating over a query result. - Use separate
getElementsByTagName()calls only when one-name lookups or tag-specific processing are clearer.
Frequently Asked Questions
Can I pass several names to getElementsByTagName(), such as ‘h1,h2,p’?
No. The method accepts one local tag name. Use separate calls or an XPath union such as //h1 | //h2 | //p.
Does an XPath union preserve the order in the original HTML?
Yes. The node-set returned by the union is presented in document order, unlike output concatenated from separate tag lookups.
Why does my query work for HTML but not for XHTML?
XHTML commonly places elements in a namespace. Register that namespace on DOMXPath and query with a prefix such as //xhtml:h1.
What should I do when query() returns false?
Treat it as an XPath or context error, inspect the expression, and handle the failure before iterating. An empty DOMNodeList is different: it means the expression was valid but matched nothing.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.
The Tool Desk
Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →

