Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

Parse the HTML into a DOM, find the boundary elements with DOMXPath, then walk from the start node’s nextSibling until you reach the end node. That loop makes the stopping rule explicit and is the safest fit when sections may repeat. Use XPath’s following-sibling axis only when both markers are unique siblings in a known structure.

Choose the selection method that matches your HTML

Approach Best for Trade-off
DOM sibling loop Repeated sections, stopping at the first end marker, or needing control over whitespace and comments. Requires a few more lines, but termination is explicit.
XPath following-sibling A stable section with unique boundary markers among the same parent’s children. Can select too much when markers repeat or the structure changes.
Container-scoped XPath plus a loop Several independent sections inside reliable containers. Depends on identifying the right container and using a relative path.

Both methods operate on parsed DOM nodes, not raw HTML strings. XPath locates the markers; DOM traversal or XPath then selects the nodes between them.

Parse the HTML and collect nodes up to the end marker

This runnable example finds the first matching start and end headings, then collects non-empty text from element and text nodes between them. The boundary headings themselves are excluded.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
<?php
$html = <<<'HTML'
<div class="content">
  <h2 id="start">Start</h2>
  <p>First value</p>
  <p>Second <strong>value</strong></p>
  <h2 id="end">End</h2>
  <p>Outside the range</p>
</div>
HTML;

$doc = new DOMDocument();
$previousSetting = libxml_use_internal_errors(true);
$loaded = $doc->loadHTML($html, LIBXML_NOERROR | LIBXML_NOWARNING);
libxml_clear_errors();
libxml_use_internal_errors($previousSetting);

if (!$loaded) {
    throw new RuntimeException('Could not parse HTML');
}

$xpath = new DOMXPath($doc);
$startResults = $xpath->query("//h2[@id='start']");
$endResults = $xpath->query("//h2[@id='end']");
if ($startResults === false || $endResults === false) {
    throw new RuntimeException('Invalid XPath expression');
}
$start = $startResults->item(0);
$end = $endResults->item(0);

$values = [];
if ($start !== null && $end !== null) {
    for ($node = $start->nextSibling; $node !== null; $node = $node->nextSibling) {
        if ($node->isSameNode($end)) {
            break;
        }
        if ($node->nodeType === XML_ELEMENT_NODE || $node->nodeType === XML_TEXT_NODE) {
            $text = trim($node->textContent ?? '');
            if ($text !== '') {
                $values[] = $text;
            }
        }
    }
}

print_r($values);

The expected values are “First value” and “Second value”; “Outside the range” follows the end marker and is not collected. The second paragraph’s nested <strong> text is included because textContent reads descendant text as well.

Check queries and missing markers

DOMXPath::query() returns a DOMNodeList on success or false for an invalid expression or context node. Check for false before using item(). item(0) returns null when there is no match, so handle missing markers before starting the loop. PHP’s DOMXPath::query() documentation describes the return value and context-node behavior.

Why use nextSibling?

The loop advances through the start node’s siblings and stops as soon as it encounters the exact end node, tested with isSameNode(). This implements “up to the first end marker” directly. It also means whitespace text nodes and comments can appear in the sequence, even if they do not contribute values. The node’s nodeType lets the example include elements and text while skipping comments.

textContent produces readable text, not the original markup. A paragraph containing nested emphasis contributes its combined text as one value. If you need individual text nodes rather than one value per element, traverse descendant text nodes separately.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use XPath when the sibling boundaries are unique

When the start and end headings are unique siblings under the same parent, XPath can select the siblings that have the end marker somewhere after them:

$nodes = $xpath->query(
    "//h2[@id='start']/following-sibling::node()[following-sibling::h2[@id='end']]"
);

if ($nodes === false) {
    throw new RuntimeException('Invalid XPath expression');
}

$values = [];
foreach ($nodes as $node) {
    $text = trim($node->textContent ?? $node->nodeValue ?? '');
    if ($text !== '') {
        $values[] = $text;
    }
}

The predicate finds following siblings that have an h2 with id="end" later in the same sibling sequence. It does not include the end heading itself. This is concise for a stable, unique range, but repeated end markers can make it include nodes up to the last matching marker rather than stopping at the first. Prefer the procedural loop when “first end marker” is the requirement.

Scope extraction to a container

If the page contains multiple sections with the same marker names, first identify the intended container and run relative XPath queries from it. A leading . keeps the search relative to that context node:

$containers = $xpath->query("//div[@class='content']");
if ($containers === false) {
    throw new RuntimeException('Invalid container XPath');
}
$container = $containers->item(0);

if ($container !== null) {
    $starts = $xpath->query(".//h2[@id='start']", $container);
    $ends = $xpath->query(".//h2[@id='end']", $container);
    if ($starts === false || $ends === false) {
        throw new RuntimeException('Invalid boundary XPath');
    }
    $start = $starts->item(0);
    $end = $ends->item(0);
    // Apply the sibling loop only after confirming both markers are in
    // the intended sibling sequence.
}

A context query limits which descendants are considered, but the selected boundary nodes still need to be siblings for a direct nextSibling walk to reach the end. If the start and end live at different nesting levels, define the desired structural range first—for example, select a common ancestor or extract each matching child container—instead of assuming they share a sibling list.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Return original markup instead of text

When the result must retain links, emphasis, or other element tags, serialize each selected element node with saveHTML() rather than reading textContent:

$fragments = [];
for ($node = $start->nextSibling; $node !== null; $node = $node->nextSibling) {
    if ($node->isSameNode($end)) {
        break;
    }
    if ($node->nodeType === XML_ELEMENT_NODE) {
        $fragment = $doc->saveHTML($node);
        if ($fragment !== false) {
            $fragments[] = $fragment;
        }
    }
}

This serializes element siblings as HTML fragments; it does not return the boundary headings. Decide whether whitespace-only text nodes or comments belong in your output and handle those node types explicitly if they matter. Serialization is not a sanitizer: do not treat returned markup as safe for insertion into an unrelated page without suitable security handling.

Account for PHP’s HTML parser version

DOMDocument::loadHTML() accepts imperfect HTML, but it uses an HTML 4 parser and can build a tree that differs from a browser’s HTML5 parser. That matters when modern markup relies on browser-specific error recovery or implicit elements. The PHP manual recommends using DomHTMLDocument to parse modern HTML instead of DOMDocument. PHP 8.4 adds DomHTMLDocument::createFromString() and createFromFile() for HTML5-conforming parsing. See the PHP loadHTML() manual for the parser caveat and security warning.

Parsing behavior may also depend on the installed libxml version. If your environment supports PHP 8.4’s DomHTMLDocument, choose it when browser-like HTML5 parsing is important. For older environments or controlled, simple HTML, DOMDocument may be sufficient; verify the resulting tree for markup whose structure is significant. Do not use loadHTML() as an HTML sanitizer, especially for untrusted input.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Troubleshoot empty or incorrect selections

  • No values returned: Confirm both XPath queries match nodes, and check for null from item(0). Inspect whether the marker IDs and tag names match the parsed DOM.
  • The loop never reaches the end: The boundary may not be a sibling of the start, or it may occur in a different parent. Confirm both markers are in the same sibling sequence; scope queries to the right container if the document repeats markers.
  • Too many nodes selected by XPath: The expression can match nodes whose later siblings contain an end marker. Repeated markers or an unexpected tree can broaden the range. Use the loop to stop at the first matching end node.
  • Whitespace or comments appear: They are DOM siblings. Filter by nodeType as needed; the example includes elements and text nodes and skips comments.
  • Text combines unexpectedly: textContent includes descendant text. Use saveHTML() for a fragment or traverse child nodes if values must be separated more finely.
  • Browser and PHP structures differ: DOMDocument uses HTML 4 parsing. For modern HTML, consider DomHTMLDocument on PHP 8.4 or later and validate the parsed structure in the target PHP/libxml environment.
  • Parsing warnings clutter output: The example temporarily enables libxml internal errors, clears the parser’s errors, and restores the prior setting. Avoid changing this setting globally without restoring it in reusable code.

Or skip the browser setup

If the real task is capturing a webpage rather than extracting nodes from an HTML string, ScreenshotNeo is a website screenshot API and MCP server: one GET request can return PNG, JPEG, WebP, or PDF. For example, this PHP call saves a screenshot of the target URL:

<?php
$url = 'https://stripe.com';
$query = http_build_query([
    'access_key' => 'YOUR_API_KEY',
    'url' => $url,
]);
$body = file_get_contents('https://api.screenshotneo.com/v1/shot?' . $query);
if ($body === false) {
    throw new RuntimeException('Screenshot request failed');
}
file_put_contents('shot.webp', $body);

See the ScreenshotNeo documentation for the API and available options. Cookie banners, newsletter popups, and chat widgets are removed before the shot; bot checks, blank pages, and failed loads are never billed. Its MCP server lets AI agents take screenshots. The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000. Sign up for ScreenshotNeo’s free plan.

Frequently Asked Questions

Does the sibling loop include either boundary heading?

No. It starts at the start node’s next sibling and breaks before processing the end node.

Can XPath select nodes across different parents?

The shown following-sibling expression cannot; its range is among siblings under the same parent.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.