The Tool Desk
Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
Use driver.findElements(By.xpath("your XPath")) to retrieve every match in the currently loaded page, then repeat that lookup after each page change. Selenium does not aggregate elements from pages that are not loaded. Copy the text or attributes you need before navigating, wait for the next page’s content to be ready, and stop when the site’s own pagination state says there are no more pages.
What “duplicate XPath matches” means
There are two separate problems that are often described with the same phrase:
- Multiple matches on one page: several elements in the current DOM satisfy one XPath.
- Matches repeated across pages: the same XPath identifies elements on page 1, page 2, and later pages of a paginated or infinite-scroll result.
findElements solves the first problem for the page currently loaded. A loop that advances through the application solves the second. A single call cannot see elements in a page that WebDriver has not loaded.
Quick wins for a faster PC:
Repair Windows errors before they cause bigger problemsFix Now →Scan for outdated or missing drivers - takes under a minuteDriver Scan →| Call | Result | Use it when |
|---|---|---|
findElement(By.xpath(...)) |
One WebElement, specifically the first match |
The element is required and you expect one usable match |
findElements(By.xpath(...)) |
List<WebElement>; an empty list when there are no matches |
You need every match, a count, or optional content |
The official Selenium documentation states: “If there are no matches, an empty list is returned.” See Finding web elements for the singular and plural APIs.
#1 Best Overall
Minimal Java lookup on one page
These imports and statements collect the visible text from every matching element in the active browsing context:
import java.util.ArrayList;
import java.util.List;
import org.openqa.selenium.By;
import org.openqa.selenium.WebDriver;
import org.openqa.selenium.WebElement;
List<String> values = new ArrayList<>();
List<WebElement> matches = driver.findElements(
By.xpath("//div[@class='result']")
);
for (WebElement match : matches) {
values.add(match.getText());
}
If no result cards match, matches.size() is zero and the loop simply does nothing. Use findElement only when a missing element should be an error or when you deliberately want the first match.
Collect attributes instead of text
Read attributes while the element belongs to the current page. For example:
for (WebElement match : matches) {
String title = match.getAttribute("data-id");
String href = match.getAttribute("href");
// Persist or process the values here.
}
After navigation, the old WebElement references may no longer be valid. Store immutable values, not page-bound elements.
Scope XPath correctly
When you locate a container first, begin the descendant expression with .//:
WebElement resultsPanel = driver.findElement(By.id("results"));
List<WebElement> cards = resultsPanel.findElements(
By.xpath(".//article[contains(@class, 'result-card')]")
);
In Selenium’s WebElement API, an XPath beginning with // is documented as searching the whole document even when called from a WebElement. The dot keeps the search under the container and prevents unrelated matches elsewhere on the page.
Rank #2
Keep the locator maintainable
Verify the expression against the actual markup on every page template. Selenium’s locator guidance favors unique, predictable IDs, followed by a well-written CSS selector where it fits. XPath is useful for relationships, text conditions, and structures CSS cannot express, but long absolute paths are fragile. Prefer a compact expression based on stable attributes:
//article[@data-testid='result']
//a[contains(@href, '/products/')]
//div[@role='row' and .//span[@class='status']]
Iterate through ordinary pagination
The following pattern visits a first page, copies each match’s data, clicks the site’s next control, waits for a page-specific change, and then repeats. The selectors and stopping condition are intentionally site-specific: replace them with the controls used by your application.
import java.time.Duration;
import java.util.ArrayList;
import java.util.HashSet;
import java.util.List;
import java.util.Set;
import org.openqa.selenium.By;
import org.openqa.selenium.StaleElementReferenceException;
import org.openqa.selenium.WebDriver;
import org.openqa.selenium.WebElement;
import org.openqa.selenium.support.ui.ExpectedConditions;
import org.openqa.selenium.support.ui.WebDriverWait;
List<String> collected = new ArrayList<>();
Set<String> uniqueIds = new HashSet<>();
WebDriverWait wait = new WebDriverWait(driver, Duration.ofSeconds(20));
By resultXPath = By.xpath("//div[@class='result']");
By nextButton = By.cssSelector("a[rel='next']");
By pageMarker = By.cssSelector(".pagination .current");
while (true) {
// Wait for the current page's content, then locate it afresh.
wait.until(ExpectedConditions.presenceOfAllElementsLocatedBy(resultXPath));
List<WebElement> matches = driver.findElements(resultXPath);
for (WebElement match : matches) {
String id = match.getAttribute("data-id");
String text = match.getText();
if (id == null || uniqueIds.add(id)) {
collected.add(text);
}
}
List<WebElement> next = driver.findElements(nextButton);
if (next.isEmpty()) {
break;
}
WebElement currentMarker = driver.findElement(pageMarker);
String oldPage = currentMarker.getText();
next.get(0).click();
wait.until(ExpectedConditions.not(
ExpectedConditions.textToBe(pageMarker, oldPage)
));
}
The uniqueIds set is optional. Use it when pages can overlap or when the same record may be rendered twice. If the site has no stable ID, construct a key from values that identify a record, such as a normalized URL plus title. Do not discard legitimate duplicates merely because their visible text is equal unless that is the intended rule.
When the next control changes the URL
Capture the current URL, activate the link, and wait for the URL to change:
String oldUrl = driver.getCurrentUrl();
driver.findElement(By.cssSelector("a.next")).click();
wait.until(ExpectedConditions.urlToBeNot(oldUrl));
If the application updates content without changing the URL, wait for a page marker, a changed result element, or a loading indicator to disappear instead. The correct condition is the one that proves the new results—not an arbitrary delay.
Windows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallCrashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minuteWhen pagination uses a URL pattern
If the site documents a stable query parameter, loading each URL directly can be simpler:
Rank #3
for (int page = 1; page <= 20; page++) {
driver.get("https://example.com/search?page=" + page);
wait.until(ExpectedConditions.presenceOfAllElementsLocatedBy(resultXPath));
for (WebElement match : driver.findElements(resultXPath)) {
collected.add(match.getText());
}
}
Set a real upper bound or stop when a page returns no matches. Never assume that page numbers are contiguous or that a guessed maximum exists.
Dynamic pages and waits
WebDriver’s lookup methods are evaluated against the current page. An implicit wait can make a lookup poll for a limited period; the WebDriver API documents that behavior. It does not know which application state means “results are complete.”
Choose a condition that matches the application
- Results are inserted once: wait for presence of the result locator.
- Results are replaced after a click: wait for the old page marker to become stale or for its text to change.
- A spinner controls readiness: wait for the spinner to become invisible, then locate results again.
- At least one result is optional: wait for a stable container, then use
findElementsand handle an empty list.
A fixed Thread.sleep can be shorter than a slow response or unnecessarily long on a fast run, so it is not a universal synchronization strategy. Keep the wait tied to a state change you can observe.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Infinite scroll
There may be no next button. Scroll, wait for the count or sentinel to change, collect newly rendered elements, and stop when the application reports the end:
int previousCount = 0;
while (true) {
List<WebElement> items = driver.findElements(resultXPath);
for (int i = previousCount; i < items.size(); i++) {
collected.add(items.get(i).getText());
}
previousCount = items.size();
if (!driver.findElements(By.cssSelector(".no-more-results")).isEmpty()) {
break;
}
((org.openqa.selenium.JavascriptExecutor) driver)
.executeScript("window.scrollTo(0, document.body.scrollHeight);");
wait.until(ExpectedConditions.or(
ExpectedConditions.numberOfElementsMoreThan(resultXPath, previousCount),
ExpectedConditions.visibilityOfElementLocated(By.cssSelector(".no-more-results"))
));
}
Some virtualized lists remove off-screen rows. In that case, process each batch immediately rather than relying on one ever-growing list.
Prevent stale-element and duplicate-data errors
StaleElementReferenceException
This occurs when a framework rerenders a node or navigation replaces its document. Do not retain a WebElement across a page transition. Locate it again after the wait, and copy its text or attributes before triggering the transition. If a click itself causes rerendering, reacquire the button in a short retry rather than reusing the old reference.
Rank #4
Duplicate records caused by overlap
Some APIs show the last item of one page again at the top of the next. Deduplicate using a stable record key, not the Java object identity. A LinkedHashSet preserves insertion order while removing exact key duplicates:
Do these 3 things before closing this tab:
1Fix the driver behind crashes, sound loss and screen glitches2Repair Windows errors before they cause bigger problems3Scan for outdated or missing drivers - takes under a minuteSet<String> seen = new java.util.LinkedHashSet<>();
String key = match.getAttribute("data-id");
if (key != null && seen.add(key)) {
// Save this record once.
}
Duplicate matches caused by the XPath itself
Inspect the DOM and determine whether the same logical item is represented by nested nodes, hidden templates, or mobile and desktop copies. Narrow the locator to the rendered region, add a visibility check when appropriate, or select the item container rather than every descendant text node.
Navigation patterns and stopping rules
| Application behavior | Advance action | Reliable stop signal |
|---|---|---|
| Numbered links | Click the next number or build the documented URL | No next link, disabled next link, or no results |
| Next button with URL navigation | Click and wait for URL change | Next control absent or disabled |
| AJAX replacement | Click and wait for marker/content change | End-of-results message or unchanged page token |
| Infinite scroll | Scroll and wait for count/sentinel | No-more sentinel or no additional batch after a bounded retry |
Check disabled state as well as presence. A disabled button may still be returned by findElements, so inspect isEnabled() or the application’s disabled attribute/class before clicking.
Common failures and fixes
Only one element is returned
Cause: findElement selects the first match. Fix: use findElements and iterate the returned list.
The list is empty
Cause: the XPath is wrong, the lookup is scoped incorrectly, or the page has not rendered the content. Fix: verify the live DOM, use .// for a container search, and wait for the application’s readiness condition.
Results from the previous page are collected twice
Cause: the click happened before the old content was replaced, or the site overlaps pages. Fix: wait for a marker change or staleness, then deduplicate with a stable key.
Best Value
Element is not clickable
Cause: an overlay, disabled state, or off-screen control. Fix: wait for clickability, close the site’s overlay through its normal UI, scroll into view, and verify that the control is enabled.
XPath works on one page but not another
Cause: different templates, localized attributes, or a changed DOM. Fix: compare both templates, use stable IDs or data attributes where available, and keep page-specific locators when one shared expression would be misleading.
Browser hangs or the loop never ends
Cause: no explicit stop condition, a next button that remains present, or a wait for a state that never occurs. Fix: add a maximum page or retry guard, log the URL and page marker, and treat an unchanged marker as a failure requiring investigation rather than silently looping.
Performance, reliability, and data handling
- Collect only the fields you need; calling
getText()on thousands of complex nodes can be more expensive than reading a specific attribute. - Persist each page’s records before advancing if a long run must be resumable. Store the page URL or token with the data.
- Use a bounded explicit wait and log timeout context: URL, XPath, page number, and the last observed marker.
- Respect authentication, rate limits, robots policies, and the application’s terms. Selenium does not make a restricted site accessible by itself.
- For parallel jobs, give each WebDriver its own session and output sink; WebDriver instances and page-bound elements should not be shared casually between threads.
The APIs establish current-page lookup and navigation behavior, but the exact transition timing belongs to the target application. Design the loop around observable page state rather than assumptions about a particular site.
Or skip the browser setup
If your goal is simply to produce clean images or PDFs of each page after you have identified the URLs, ScreenshotNeo provides a single HTTP endpoint. It accepts cookie and consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be turned off. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and the response identifies the result with X-Page-Verdict and X-Billed headers.
cURL:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python:
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
See the ScreenshotNeo documentation for parameters. Its MCP server exposes take_screenshot, get_page_info, and capture_pdf to Claude, Cursor, and other MCP clients. The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account to try it.
Frequently Asked Questions
Can one XPath query search pages that are already closed or not loaded?
No. WebDriver evaluates a locator in the current browsing context. Load or navigate to each page, collect its values, and then continue.
Should I use CSS instead of XPath?
Use a unique ID or a compact CSS selector when it clearly identifies the element; keep XPath for relationships, text conditions, or structures CSS cannot express.
How can I prove that pagination finished?
Use the application’s own signal, such as an absent or disabled next control, an end-of-results message, a final page token, or a bounded URL range. Log that signal so an unexpected template change is visible.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

