Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

Use driver.findElements(By.xpath("your XPath")) to retrieve every match in the currently loaded page, then repeat that lookup after each page change. Selenium does not aggregate elements from pages that are not loaded. Copy the text or attributes you need before navigating, wait for the next page’s content to be ready, and stop when the site’s own pagination state says there are no more pages.

What “duplicate XPath matches” means

There are two separate problems that are often described with the same phrase:

  • Multiple matches on one page: several elements in the current DOM satisfy one XPath.
  • Matches repeated across pages: the same XPath identifies elements on page 1, page 2, and later pages of a paginated or infinite-scroll result.

findElements solves the first problem for the page currently loaded. A loop that advances through the application solves the second. A single call cannot see elements in a page that WebDriver has not loaded.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Call Result Use it when
findElement(By.xpath(...)) One WebElement, specifically the first match The element is required and you expect one usable match
findElements(By.xpath(...)) List<WebElement>; an empty list when there are no matches You need every match, a count, or optional content

The official Selenium documentation states: “If there are no matches, an empty list is returned.” See Finding web elements for the singular and plural APIs.

Minimal Java lookup on one page

These imports and statements collect the visible text from every matching element in the active browsing context:

import java.util.ArrayList;
import java.util.List;
import org.openqa.selenium.By;
import org.openqa.selenium.WebDriver;
import org.openqa.selenium.WebElement;

List<String> values = new ArrayList<>();
List<WebElement> matches = driver.findElements(
    By.xpath("//div[@class='result']")
);

for (WebElement match : matches) {
    values.add(match.getText());
}

If no result cards match, matches.size() is zero and the loop simply does nothing. Use findElement only when a missing element should be an error or when you deliberately want the first match.

Collect attributes instead of text

Read attributes while the element belongs to the current page. For example:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
for (WebElement match : matches) {
    String title = match.getAttribute("data-id");
    String href = match.getAttribute("href");
    // Persist or process the values here.
}

After navigation, the old WebElement references may no longer be valid. Store immutable values, not page-bound elements.

Scope XPath correctly

When you locate a container first, begin the descendant expression with .//:

WebElement resultsPanel = driver.findElement(By.id("results"));
List<WebElement> cards = resultsPanel.findElements(
    By.xpath(".//article[contains(@class, 'result-card')]")
);

In Selenium’s WebElement API, an XPath beginning with // is documented as searching the whole document even when called from a WebElement. The dot keeps the search under the container and prevents unrelated matches elsewhere on the page.

Keep the locator maintainable

Verify the expression against the actual markup on every page template. Selenium’s locator guidance favors unique, predictable IDs, followed by a well-written CSS selector where it fits. XPath is useful for relationships, text conditions, and structures CSS cannot express, but long absolute paths are fragile. Prefer a compact expression based on stable attributes:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
//article[@data-testid='result']
//a[contains(@href, '/products/')]
//div[@role='row' and .//span[@class='status']]

Iterate through ordinary pagination

The following pattern visits a first page, copies each match’s data, clicks the site’s next control, waits for a page-specific change, and then repeats. The selectors and stopping condition are intentionally site-specific: replace them with the controls used by your application.

import java.time.Duration;
import java.util.ArrayList;
import java.util.HashSet;
import java.util.List;
import java.util.Set;

import org.openqa.selenium.By;
import org.openqa.selenium.StaleElementReferenceException;
import org.openqa.selenium.WebDriver;
import org.openqa.selenium.WebElement;
import org.openqa.selenium.support.ui.ExpectedConditions;
import org.openqa.selenium.support.ui.WebDriverWait;

List<String> collected = new ArrayList<>();
Set<String> uniqueIds = new HashSet<>();
WebDriverWait wait = new WebDriverWait(driver, Duration.ofSeconds(20));

By resultXPath = By.xpath("//div[@class='result']");
By nextButton = By.cssSelector("a[rel='next']");
By pageMarker = By.cssSelector(".pagination .current");

while (true) {
    // Wait for the current page's content, then locate it afresh.
    wait.until(ExpectedConditions.presenceOfAllElementsLocatedBy(resultXPath));
    List<WebElement> matches = driver.findElements(resultXPath);

    for (WebElement match : matches) {
        String id = match.getAttribute("data-id");
        String text = match.getText();
        if (id == null || uniqueIds.add(id)) {
            collected.add(text);
        }
    }

    List<WebElement> next = driver.findElements(nextButton);
    if (next.isEmpty()) {
        break;
    }

    WebElement currentMarker = driver.findElement(pageMarker);
    String oldPage = currentMarker.getText();
    next.get(0).click();

    wait.until(ExpectedConditions.not(
        ExpectedConditions.textToBe(pageMarker, oldPage)
    ));
}

The uniqueIds set is optional. Use it when pages can overlap or when the same record may be rendered twice. If the site has no stable ID, construct a key from values that identify a record, such as a normalized URL plus title. Do not discard legitimate duplicates merely because their visible text is equal unless that is the intended rule.

When the next control changes the URL

Capture the current URL, activate the link, and wait for the URL to change:

String oldUrl = driver.getCurrentUrl();
driver.findElement(By.cssSelector("a.next")).click();
wait.until(ExpectedConditions.urlToBeNot(oldUrl));

If the application updates content without changing the URL, wait for a page marker, a changed result element, or a loading indicator to disappear instead. The correct condition is the one that proves the new results—not an arbitrary delay.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

When pagination uses a URL pattern

If the site documents a stable query parameter, loading each URL directly can be simpler:

for (int page = 1; page <= 20; page++) {
    driver.get("https://example.com/search?page=" + page);
    wait.until(ExpectedConditions.presenceOfAllElementsLocatedBy(resultXPath));
    for (WebElement match : driver.findElements(resultXPath)) {
        collected.add(match.getText());
    }
}

Set a real upper bound or stop when a page returns no matches. Never assume that page numbers are contiguous or that a guessed maximum exists.

Dynamic pages and waits

WebDriver’s lookup methods are evaluated against the current page. An implicit wait can make a lookup poll for a limited period; the WebDriver API documents that behavior. It does not know which application state means “results are complete.”

Choose a condition that matches the application

  • Results are inserted once: wait for presence of the result locator.
  • Results are replaced after a click: wait for the old page marker to become stale or for its text to change.
  • A spinner controls readiness: wait for the spinner to become invisible, then locate results again.
  • At least one result is optional: wait for a stable container, then use findElements and handle an empty list.

A fixed Thread.sleep can be shorter than a slow response or unnecessarily long on a fast run, so it is not a universal synchronization strategy. Keep the wait tied to a state change you can observe.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Infinite scroll

There may be no next button. Scroll, wait for the count or sentinel to change, collect newly rendered elements, and stop when the application reports the end:

int previousCount = 0;
while (true) {
    List<WebElement> items = driver.findElements(resultXPath);
    for (int i = previousCount; i < items.size(); i++) {
        collected.add(items.get(i).getText());
    }
    previousCount = items.size();

    if (!driver.findElements(By.cssSelector(".no-more-results")).isEmpty()) {
        break;
    }
    ((org.openqa.selenium.JavascriptExecutor) driver)
        .executeScript("window.scrollTo(0, document.body.scrollHeight);");
    wait.until(ExpectedConditions.or(
        ExpectedConditions.numberOfElementsMoreThan(resultXPath, previousCount),
        ExpectedConditions.visibilityOfElementLocated(By.cssSelector(".no-more-results"))
    ));
}

Some virtualized lists remove off-screen rows. In that case, process each batch immediately rather than relying on one ever-growing list.

Prevent stale-element and duplicate-data errors

StaleElementReferenceException

This occurs when a framework rerenders a node or navigation replaces its document. Do not retain a WebElement across a page transition. Locate it again after the wait, and copy its text or attributes before triggering the transition. If a click itself causes rerendering, reacquire the button in a short retry rather than reusing the old reference.

Duplicate records caused by overlap

Some APIs show the last item of one page again at the top of the next. Deduplicate using a stable record key, not the Java object identity. A LinkedHashSet preserves insertion order while removing exact key duplicates:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Set<String> seen = new java.util.LinkedHashSet<>();
String key = match.getAttribute("data-id");
if (key != null && seen.add(key)) {
    // Save this record once.
}

Duplicate matches caused by the XPath itself

Inspect the DOM and determine whether the same logical item is represented by nested nodes, hidden templates, or mobile and desktop copies. Narrow the locator to the rendered region, add a visibility check when appropriate, or select the item container rather than every descendant text node.

Navigation patterns and stopping rules

Application behavior Advance action Reliable stop signal
Numbered links Click the next number or build the documented URL No next link, disabled next link, or no results
Next button with URL navigation Click and wait for URL change Next control absent or disabled
AJAX replacement Click and wait for marker/content change End-of-results message or unchanged page token
Infinite scroll Scroll and wait for count/sentinel No-more sentinel or no additional batch after a bounded retry

Check disabled state as well as presence. A disabled button may still be returned by findElements, so inspect isEnabled() or the application’s disabled attribute/class before clicking.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Common failures and fixes

Only one element is returned

Cause: findElement selects the first match. Fix: use findElements and iterate the returned list.

The list is empty

Cause: the XPath is wrong, the lookup is scoped incorrectly, or the page has not rendered the content. Fix: verify the live DOM, use .// for a container search, and wait for the application’s readiness condition.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Results from the previous page are collected twice

Cause: the click happened before the old content was replaced, or the site overlaps pages. Fix: wait for a marker change or staleness, then deduplicate with a stable key.

Element is not clickable

Cause: an overlay, disabled state, or off-screen control. Fix: wait for clickability, close the site’s overlay through its normal UI, scroll into view, and verify that the control is enabled.

XPath works on one page but not another

Cause: different templates, localized attributes, or a changed DOM. Fix: compare both templates, use stable IDs or data attributes where available, and keep page-specific locators when one shared expression would be misleading.

Browser hangs or the loop never ends

Cause: no explicit stop condition, a next button that remains present, or a wait for a state that never occurs. Fix: add a maximum page or retry guard, log the URL and page marker, and treat an unchanged marker as a failure requiring investigation rather than silently looping.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Performance, reliability, and data handling

  • Collect only the fields you need; calling getText() on thousands of complex nodes can be more expensive than reading a specific attribute.
  • Persist each page’s records before advancing if a long run must be resumable. Store the page URL or token with the data.
  • Use a bounded explicit wait and log timeout context: URL, XPath, page number, and the last observed marker.
  • Respect authentication, rate limits, robots policies, and the application’s terms. Selenium does not make a restricted site accessible by itself.
  • For parallel jobs, give each WebDriver its own session and output sink; WebDriver instances and page-bound elements should not be shared casually between threads.

The APIs establish current-page lookup and navigation behavior, but the exact transition timing belongs to the target application. Design the loop around observable page state rather than assumptions about a particular site.

Or skip the browser setup

If your goal is simply to produce clean images or PDFs of each page after you have identified the URLs, ScreenshotNeo provides a single HTTP endpoint. It accepts cookie and consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be turned off. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and the response identifies the result with X-Page-Verdict and X-Billed headers.

cURL:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

See the ScreenshotNeo documentation for parameters. Its MCP server exposes take_screenshot, get_page_info, and capture_pdf to Claude, Cursor, and other MCP clients. The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account to try it.

Frequently Asked Questions

Can one XPath query search pages that are already closed or not loaded?

No. WebDriver evaluates a locator in the current browsing context. Load or navigate to each page, collect its values, and then continue.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Should I use CSS instead of XPath?

Use a unique ID or a compact CSS selector when it clearly identifies the element; keep XPath for relationships, text conditions, or structures CSS cannot express.

How can I prove that pagination finished?

Use the application’s own signal, such as an absent or disabled next control, an end-of-results message, a final page token, or a bounded URL range. Log that signal so an unexpected template change is visible.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.