Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

Use PHP’s DOM extension to load the HTML, then query it with XPath. For example, //a[@href] finds links that have an href attribute, while //a[@href="/about"] finds links whose value is exactly /about. Iterate over the matches and call getAttribute() to read an attribute’s value.

Find elements by attribute with DOMXPath

DOMXPath is PHP’s built-in way to run XPath 1.0 queries against HTML or XML documents. In an XPath expression, @ refers to an attribute. Put an attribute test in square brackets after the element name to select elements based on whether that attribute exists or what its value is.

This complete example finds every anchor with an href attribute, including one whose value is empty. It then prints the value of each match:

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
<?php
$html = '<main>
    <a href="/about">About</a>
    <a href="">Empty link</a>
    <a>Missing href</a>
</main>';

$doc = new DOMDocument();
$doc->loadHTML($html);
$xpath = new DOMXPath($doc);

$links = $xpath->query('//a[@href]');
if ($links === false) {
    throw new RuntimeException('Invalid XPath expression');
}

foreach ($links as $link) {
    echo $link->getAttribute('href'), PHP_EOL;
}

The output is /about followed by an empty line. The third anchor is not selected because it has no href attribute.

Check that an attribute exists

Use [@attribute] to require that an attribute is present. For example, //*[@data-id] selects any element with a data-id attribute. To restrict the result to a tag, name the tag: //button[@type] finds buttons that have a type attribute.

Match an exact value

Use an equality test inside the predicate: //*[@data-id="42"] selects elements whose data-id value is exactly 42. Combine a tag and a value test when that is more specific: //button[@type="submit"] finds submit buttons.

XPath string values use quotes. If the value you need to match contains quotes, build the XPath string literal carefully rather than concatenating untrusted input into an expression. Untrusted text inserted into XPath can change the meaning of a query. For a fixed or application-controlled value, a correctly quoted literal is usually sufficient; for arbitrary input, use a safe XPath-literal construction strategy or select a narrower set of nodes and compare the attribute in PHP.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Read an attribute value after selecting an element

Finding the element and retrieving its value are separate operations. XPath returns matching nodes; on a matched DOMElement, getAttribute('data-id') reads the value:

$nodes = $xpath->query('//*[@data-id]');
if ($nodes === false) {
    throw new RuntimeException('Invalid XPath expression');
}

foreach ($nodes as $node) {
    if ($node instanceof DOMElement) {
        $value = $node->getAttribute('data-id');
        echo $node->tagName, ': ', $value, PHP_EOL;
    }
}

getAttribute() returns an empty string when the requested attribute is absent. That means an empty result alone cannot tell you whether the attribute was missing or present with an empty value. If that distinction matters, check hasAttribute() first:

if ($node instanceof DOMElement && $node->hasAttribute('data-id')) {
    $value = $node->getAttribute('data-id');
    // The attribute exists; its value may still be an empty string.
}

Handle XPath results and errors

DOMXPath::query() returns a DOMNodeList for a node-selecting expression. When the XPath is valid but nothing matches, the list is simply empty, so a foreach loop runs zero times. A malformed expression or invalid context node produces false; check for that before iterating, as in the examples above.

Keep the query focused on the elements you need. //a[@href] searches the document for matching anchors. If you already have a context element and want only descendants beneath it, pass that element as the second argument to query() and use a relative expression beginning with a dot:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
$buttons = $xpath->query('.//button[@type="submit"]', $container);
if ($buttons === false) {
    throw new RuntimeException('Invalid XPath expression or context node');
}

The leading . matters: .//button expresses a descendant search relative to the context node. An expression beginning with // searches from the document root instead of expressing that relative path.

Find data attributes and combine conditions

HTML data attributes are ordinary attributes from XPath’s perspective. These examples show common patterns:

  • //*[@data-product] selects any element with a data-product attribute.
  • //*[@data-product="chair"] selects any element whose data-product value is exactly chair.
  • //article[@data-category="news"] restricts the match to article elements with that value.
  • //a[@href and @data-track] selects anchors that have both attributes.

Attribute predicates are useful when conditions belong in the selection itself—for example, when you want only tracked links, not every link followed by a PHP-side filter. If the conditions are simpler to express in PHP or only a fixed set of tags is relevant, you can traverse those tags and test their attributes there instead.

Choose XPath or tag traversal

Approach Best fit Trade-off
DOMXPath predicates Combined tag and attribute conditions, or searches spanning several tag names. Requires a valid XPath expression; check for false when the expression or context is invalid.
Tag-based traversal and attribute checks A small, fixed set of tags where PHP-side conditions are easy to read. You may need to inspect more elements and write the filtering logic yourself.

For instance, traversal can be straightforward if the only question is whether any button has a particular attribute:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
foreach ($doc->getElementsByTagName('button') as $button) {
    if ($button->hasAttribute('type') && $button->getAttribute('type') === 'submit') {
        echo $button->getAttribute('type'), PHP_EOL;
    }
}

Use XPath when its predicates make the requested selection clearer. Use traversal when the tag set is narrow and the condition is easier to understand as ordinary PHP.

Work with namespaced attributes

For an attribute in an XML namespace, use getAttributeNS($namespaceUri, $localName). The namespace URI identifies the namespace, and the local name identifies the attribute within it. This is different from looking up an attribute by its literal prefixed spelling.

$value = $element->getAttributeNS('https://example.com/ns', 'code');

To query namespace-qualified nodes or attributes with XPath, register a prefix on the DOMXPath instance using registerNamespace(), then use that prefix in the expression. The prefix you register is a query alias; it need not be the same prefix used in the source document. Use namespace-aware access when the namespace is part of the document’s meaning rather than treating a colon-containing name as an ordinary unqualified attribute.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

PHP versions, HTML input, and encoding

Traditional and PHP 8.4 DOMXPath APIs

The examples in this article use the traditional DOMDocument, DOMXPath, and DOMElement APIs. PHP 8.4 also provides the newer DomXPath class, described by the PHP manual as a modern, spec-compliant equivalent. Do not substitute that class into older-runtime code without confirming the PHP version: use the API available in the runtime where the script will run.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Parsing an HTML fragment or page

DOMDocument::loadHTML() is used above to parse an HTML string. For a full remote page, retrieve the response separately and pass its HTML to the parser; parsing does not itself fetch a URL. HTML may be imperfect or incomplete, so if a query returns no matches, first confirm that the intended markup was actually loaded and that the node names and attribute values in the parsed document are what you expect.

Encoding

PHP’s DOM extension uses UTF-8. Ordinary UTF-8 snippets generally need no special handling, but legacy documents in another encoding may require conversion before parsing so text and attribute values are interpreted correctly.

Troubleshoot common problems

  • The result list is empty. The XPath ran but found no matching nodes. Check that the source HTML contains the target element and attribute, that the spelling and value match, and that the parser received the expected document.
  • query() returned false. The expression may be malformed or the context node invalid. Check brackets and quotes in the XPath and verify the context belongs to the document being queried.
  • A value appears empty. The attribute may be absent, or it may exist with an empty value. Call hasAttribute() to distinguish those cases.
  • A scoped query finds elements elsewhere. Use a relative expression such as .//button with the context node. An expression starting with // searches from the document root.
  • Namespaced attributes are not found. Register the namespace URI with DOMXPath::registerNamespace() for XPath queries, or retrieve a known namespaced value with getAttributeNS().
  • Non-ASCII values look wrong. Check the input encoding. DOM expects UTF-8, so convert legacy-encoded input as needed.

Or skip the browser setup

PHP DOM is the right choice when you need to inspect attributes in markup. If instead you need a visual screenshot or PDF of a live page, ScreenshotNeo is a separate option: it captures a URL, rather than returning DOM elements or attribute values. Its API can remove cookie and consent banners, newsletter popups, and chat widgets before capture; those cleanup steps can be turned off. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, with verdict and billing information in response headers.

For a screenshot of a page such as Stripe’s homepage, one GET request is:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

See the ScreenshotNeo API documentation for request parameters. It also offers an MCP server with take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 shots. Sign up for ScreenshotNeo’s free plan.

FAQ

Does this parse JavaScript-rendered page content?

DOMDocument::loadHTML() parses the HTML string supplied to it; it does not run a browser or execute page JavaScript. If the markup is generated only after scripts run, obtain the rendered HTML through an appropriate browser workflow before parsing it.

Can I use this to fetch a web page?

No. The examples parse HTML already held in a PHP string. Retrieving a URL is a separate networking step; handle that response and its errors before passing the returned markup to the DOM parser.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.