What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

Short answer: use Python’s xml.etree.ElementTree when a small XPath subset is enough for an XML document, lxml.etree when you need complete XPath expressions, namespaces, variables, or repeated evaluation, and Selenium’s By.XPATH when the selector must run against a live browser DOM. In every case, start with a short XPath anchored to a stable attribute, then add predicates or relationships only as needed.

The correct selector depends on the input. ElementTree and lxml query a parsed tree that already exists in Python; Selenium queries the current page in a browser, where JavaScript, timing, frames, and changing state matter.

Choose the Python XPath tool first

Tool Best use XPath support Where the query runs
xml.etree.ElementTree Small XML extraction tasks with no third-party dependency Limited XPath subset; a full XPath engine is outside the module’s scope In-memory XML tree
lxml.etree Complex XML or HTML queries, namespaces, variables, functions, and repeated evaluation XPath 1.0 plus EXSLT extensions through libxml2/libxslt; supports custom extension functions Parsed XML or HTML tree
Selenium Finding elements in a live, JavaScript-driven browser page XPath expressions passed through By.XPATH WebDriver’s current browser DOM

If you are unsure, begin with ElementTree for straightforward XML. Move to lxml when your expression needs functions, variables, richer axes, or repeated evaluation. Use Selenium only when you need browser behavior rather than merely parsing markup.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

ElementTree: the standard-library option

ElementTree’s findall(), find(), and iterfind() accept a deliberately limited XPath subset. This is enough for child paths, descendant searches, attributes, simple predicates, and positional tests, but not for every XPath function or axis.

Basic descendant and positional queries

import xml.etree.ElementTree as ET

xml_text = '''<catalog>
  <book id="b1"><title>XPath Basics</title></book>
  <book id="b2"><title>Browser Automation</title></book>
</catalog>'''

root = ET.fromstring(xml_text)

books = root.findall('.//book')
second_book = root.findall('.//book[2]')
book_titles = root.findall('.//book/title')
book_b2 = root.findall(".//book[@id='b2']")

for book in books:
    print(book.get('id'), book.findtext('title'))

The leading . makes the expression relative to root. A path such as .//book searches descendants at any depth. An attribute predicate selects only matching elements, while [2] selects the second matching sibling in the relevant step.

Namespaces in ElementTree

Unqualified names do not match namespaced XML elements automatically. Use the namespace URI in Clark notation, or pass a namespace map when the API you are using supports that form.

import xml.etree.ElementTree as ET

xml_text = '''<record xmlns:dc="http://purl.org/dc/elements/1.1/">
  <dc:title>A namespaced title</dc:title>
</record>'''

root = ET.fromstring(xml_text)
titles = root.findall('.//{http://purl.org/dc/elements/1.1/}title')
print(titles[0].text)

This explicit URI is dependable when the document vocabulary is known. A query for .//title would return nothing because the element is actually namespace-qualified.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Know ElementTree’s boundary

ElementTree is not a full XPath engine. If you need expressions such as contains(), starts-with(), arbitrary axes, variables, or scalar functions such as count(), use lxml or perform a small amount of Python filtering after selecting a broader set.

lxml: full XPath expressions for parsed XML or HTML

Install lxml in the environment that runs your script:

python -m pip install lxml

lxml.etree supports XPath 1.0, XSLT 1.0, and EXSLT extensions through libxml2/libxslt. Its xpath() method returns the natural XPath result type: element nodes, strings, booleans, or numbers. It also supports compiled XPath objects and evaluators for repeated queries.

Variables and scalar results

from lxml import etree

root = etree.fromstring(
    b'<catalog><book id="b1">XPath</book></catalog>'
)

book_id = 'b1'
books = root.xpath('//book[@id=$wanted_id]', wanted_id=book_id)
texts = root.xpath('//book/text()')
book_count = root.xpath('count(//book)')

print(books[0].text)
print(texts)
print(book_count)

The variable is bound separately from the XPath string, which avoids interpolating untrusted text into the expression. //book/text() returns strings, while //book returns element objects and count(//book) returns a number.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Relative versus absolute context

from lxml import etree

root = etree.fromstring(
    b'<catalog><section><book id="b1"/></section></catalog>'
)
section = root.xpath('/catalog/section')[0]

all_books = root.xpath('/catalog/section/book')
books_in_section = section.xpath('.//book')

print(len(all_books), len(books_in_section))

/catalog/section/book starts at the document root. Once the context is the section element, use .//book to search within that subtree. Omitting the dot changes the meaning: an absolute expression is evaluated from the document, not from the selected element.

Namespaces and HTML parsing

For namespace-heavy XML, pass a prefix map and use that prefix in the expression:

from lxml import etree

xml = b'''<feed xmlns="urn:example:feed">
  <entry><title>Item one</title></entry>
</feed>'''
root = etree.fromstring(xml)
ns = {'f': 'urn:example:feed'}
print(root.xpath('//f:entry/f:title/text()', namespaces=ns))

When the vocabulary is unknown, local-name() can match by local part, but explicit namespaces are safer when you know the document schema because they avoid accidental matches. For HTML, parse the markup into an lxml tree first, then evaluate the same XPath methods.

Selenium: XPath against a live browser DOM

Selenium exposes XPath through By.XPATH. Use it when the page is generated or modified by JavaScript, or when you need to interact with the rendered document rather than a downloaded file.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A complete locator pattern with an explicit wait

from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC

url = 'https://example.com'
driver = webdriver.Chrome()
wait = WebDriverWait(driver, 15)

try:
    driver.get(url)
    heading = wait.until(
        EC.visibility_of_element_located(
            (By.XPATH, "//h1[normalize-space()]")
        )
    )
    print(heading.text)
finally:
    driver.quit()

The wait checks both presence and state before reading the element. In a real application, replace the example expression with a selector for your page and include the XPath in any failure message so a broken locator can be diagnosed quickly.

Scope a child query with a relative XPath

from selenium.webdriver.common.by import By

login = driver.find_element(By.XPATH, "//form[@id='loginForm']")
username = login.find_element(By.XPATH, ".//input[@name='username']")
submit = login.find_element(
    By.XPATH,
    ".//input[@continue and @type='submit']",
)

The dot in .//input keeps the search inside the form. Without it, a document-wide search could match an input belonging to another form.

The final example intentionally illustrates the structure; replace @continue with the attribute that actually exists in your markup, such as @name='continue'. A corrected submit locator would be:

submit = login.find_element(
    By.XPATH,
    ".//input[@name='continue' and @type='submit']",
)

XPath patterns that remain readable

Anchor to stable attributes

  • Prefer a unique, predictable id when the application treats it as a stable contract.
  • Use meaningful name, aria-label, or other semantic attributes when IDs are unavailable.
  • Combine predicates for a specific control, such as //input[@name='email' and @type='email'].

Keep expressions short. Generated CSS classes and long positional paths are coupled to implementation details and tend to break during harmless markup changes.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use relationships when no single attribute is unique

# A label relationship is often clearer than a deep absolute path
email = driver.find_element(
    By.XPATH,
    "//label[normalize-space()='Email']/following::input[1]",
)

# Select a button inside the article that owns a heading
button = driver.find_element(
    By.XPATH,
    "//h2[normalize-space()='Settings']/ancestor::section[1]//button",
)

Relationship-based selectors express the page’s meaning, but keep the relationship as local as possible so a new section elsewhere does not create an accidental match.

Text predicates and quoting

normalize-space() removes surrounding whitespace and collapses runs of internal whitespace, making text tests less sensitive to formatting. For text containing an apostrophe, use double quotes around the XPath literal; for text containing both quote types, construct an XPath concat() expression or match a stable attribute instead.

Why an XPath returns nothing

  1. Wrong context: a query intended for a selected element must begin with ., for example .//input, not //input.
  2. Namespace mismatch: unprefixed names do not match namespaced XML. Use ElementTree’s qualified name or an explicit lxml namespace map.
  3. Page not ready: in Selenium, wait for the element and the state you need before querying or clicking.
  4. Overly specific path: replace an absolute path such as /html/body/form[1] with a stable ID, name, label, or nearby relationship.
  5. Wrong result type: element paths return nodes, text() returns strings, and functions such as count() return numbers. Do not call element methods on a scalar result.
  6. Selector tested against different markup: inspect the exact parsed XML or current browser DOM. A server response may differ from what JavaScript later renders.

Reduce a failing expression incrementally

from selenium.webdriver.common.by import By

xpath = "//form[@id='loginForm']//input[@name='username']"

# Test the stable anchor first
form = driver.find_elements(By.XPATH, "//*[@id='loginForm']")
if not form:
    raise RuntimeError("No element matched //*[@id='loginForm']")

# Then add the child condition
fields = form[0].find_elements(By.XPATH, ".//input[@name='username']")
if not fields:
    raise RuntimeError(f"No element matched {xpath}")

Testing the smallest stable predicate first tells you whether the failure is the anchor, the relationship, or the page state.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Performance, reliability, and maintenance

Choose the lowest-cost context

  • ElementTree and lxml run locally on an already parsed tree; there is no browser startup or WebDriver round trip.
  • lxml is the practical choice for repeated complex queries. Compile a frequently reused expression with etree.XPath() rather than rebuilding it in a tight loop.
  • Selenium is heavier because it drives a browser and must wait for live state. Reuse one driver for a related workflow, but always quit it in a finally block.

Make selectors a deliberate interface

Keep XPath strings in one place, give each locator a descriptive name, and include the expression in error logs. Prefer semantic attributes supplied by the application over generated classes or indexes. Use positional indexes only when the document contract guarantees the order.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Validate before changing the selector

When a query suddenly returns zero results, first verify the document or browser state, then the namespace and context. Changing a working selector to a longer absolute path can hide the real timing or markup problem and make future changes more expensive.

Or skip the browser setup

If your goal is a clean image or PDF of a page rather than extracting DOM nodes, ScreenshotNeo can capture the URL with one request. It is not an XPath engine, so keep ElementTree, lxml, or Selenium for data extraction and interaction; use ScreenshotNeo when the deliverable is a screenshot.

Its API accepts consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and each response identifies the page verdict and billing state with X-Page-Verdict and X-Billed headers. An MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients.

One-call examples

See the ScreenshotNeo documentation for request details.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot failed: ${res.status}`);
const image = Buffer.from(await res.arrayBuffer());

ScreenshotNeo includes full-page capture with lazy images loaded, CSS-selector element capture, device presets and custom viewports, retina scale, PDF paper and margin controls, custom CSS and JavaScript, click-before-capture actions, waits, request blocking, headers, cookies, user agents, authorization, timezone and geolocation controls, transparent backgrounds, resizing, selectable cache TTLs, signed links, asynchronous jobs with signed webhooks, bulk capture for up to 100 URLs per call, a usage API, and an OpenAPI specification. Its parameter names are compatible with those used by other screenshot APIs, which can simplify migration.

The Free plan includes 1,000 screenshots per month without a card. Paid plans are Starter ($5 for 3,000), Growth ($15 for 15,000), Pro ($39 for 60,000), Scale ($99 for 250,000), and Business ($249 for 1,000,000); yearly billing gives two months free, and every feature is available on every plan. Create a free ScreenshotNeo account to start with the 1,000 included screenshots.

Frequently Asked Questions

Can XPath select an attribute instead of an element?

Yes. In lxml, an expression such as //book/@id returns attribute strings. ElementTree’s limited XPath support is better suited to selecting the element and reading element.get('id') in Python.

How do I handle an XPath literal containing both quote characters?

Use XPath’s concat() function to assemble the literal, or avoid text matching by selecting a stable attribute. lxml supports concat(); ElementTree does not provide the full function set.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Can Selenium XPath cross into an iframe?

No. Switch the driver to the target frame before locating elements, then switch back to the default content when the frame work is complete. The XPath itself is evaluated within the current document context.

Should I use CSS selectors instead?

Use a readable CSS selector when a stable ID or class is sufficient. XPath is especially useful for text conditions and relationships between elements; whichever syntax you choose, keep the locator short and tied to a stable contract.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.