What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
Short answer: use Python’s xml.etree.ElementTree when a small XPath subset is enough for an XML document, lxml.etree when you need complete XPath expressions, namespaces, variables, or repeated evaluation, and Selenium’s By.XPATH when the selector must run against a live browser DOM. In every case, start with a short XPath anchored to a stable attribute, then add predicates or relationships only as needed.
The correct selector depends on the input. ElementTree and lxml query a parsed tree that already exists in Python; Selenium queries the current page in a browser, where JavaScript, timing, frames, and changing state matter.
Choose the Python XPath tool first
| Tool | Best use | XPath support | Where the query runs |
|---|---|---|---|
xml.etree.ElementTree |
Small XML extraction tasks with no third-party dependency | Limited XPath subset; a full XPath engine is outside the module’s scope | In-memory XML tree |
lxml.etree |
Complex XML or HTML queries, namespaces, variables, functions, and repeated evaluation | XPath 1.0 plus EXSLT extensions through libxml2/libxslt; supports custom extension functions | Parsed XML or HTML tree |
| Selenium | Finding elements in a live, JavaScript-driven browser page | XPath expressions passed through By.XPATH |
WebDriver’s current browser DOM |
If you are unsure, begin with ElementTree for straightforward XML. Move to lxml when your expression needs functions, variables, richer axes, or repeated evaluation. Use Selenium only when you need browser behavior rather than merely parsing markup.
ElementTree: the standard-library option
ElementTree’s findall(), find(), and iterfind() accept a deliberately limited XPath subset. This is enough for child paths, descendant searches, attributes, simple predicates, and positional tests, but not for every XPath function or axis.
#1 Best Overall
Basic descendant and positional queries
import xml.etree.ElementTree as ET
xml_text = '''<catalog>
<book id="b1"><title>XPath Basics</title></book>
<book id="b2"><title>Browser Automation</title></book>
</catalog>'''
root = ET.fromstring(xml_text)
books = root.findall('.//book')
second_book = root.findall('.//book[2]')
book_titles = root.findall('.//book/title')
book_b2 = root.findall(".//book[@id='b2']")
for book in books:
print(book.get('id'), book.findtext('title'))
The leading . makes the expression relative to root. A path such as .//book searches descendants at any depth. An attribute predicate selects only matching elements, while [2] selects the second matching sibling in the relevant step.
Namespaces in ElementTree
Unqualified names do not match namespaced XML elements automatically. Use the namespace URI in Clark notation, or pass a namespace map when the API you are using supports that form.
import xml.etree.ElementTree as ET
xml_text = '''<record xmlns:dc="http://purl.org/dc/elements/1.1/">
<dc:title>A namespaced title</dc:title>
</record>'''
root = ET.fromstring(xml_text)
titles = root.findall('.//{http://purl.org/dc/elements/1.1/}title')
print(titles[0].text)
This explicit URI is dependable when the document vocabulary is known. A query for .//title would return nothing because the element is actually namespace-qualified.
Quick wins for a faster PC:
Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Clear out junk files and repair common Windows errorsFree Scan →Know ElementTree’s boundary
ElementTree is not a full XPath engine. If you need expressions such as contains(), starts-with(), arbitrary axes, variables, or scalar functions such as count(), use lxml or perform a small amount of Python filtering after selecting a broader set.
lxml: full XPath expressions for parsed XML or HTML
Install lxml in the environment that runs your script:
Rank #2
python -m pip install lxml
lxml.etree supports XPath 1.0, XSLT 1.0, and EXSLT extensions through libxml2/libxslt. Its xpath() method returns the natural XPath result type: element nodes, strings, booleans, or numbers. It also supports compiled XPath objects and evaluators for repeated queries.
Variables and scalar results
from lxml import etree
root = etree.fromstring(
b'<catalog><book id="b1">XPath</book></catalog>'
)
book_id = 'b1'
books = root.xpath('//book[@id=$wanted_id]', wanted_id=book_id)
texts = root.xpath('//book/text()')
book_count = root.xpath('count(//book)')
print(books[0].text)
print(texts)
print(book_count)
The variable is bound separately from the XPath string, which avoids interpolating untrusted text into the expression. //book/text() returns strings, while //book returns element objects and count(//book) returns a number.
Relative versus absolute context
from lxml import etree
root = etree.fromstring(
b'<catalog><section><book id="b1"/></section></catalog>'
)
section = root.xpath('/catalog/section')[0]
all_books = root.xpath('/catalog/section/book')
books_in_section = section.xpath('.//book')
print(len(all_books), len(books_in_section))
/catalog/section/book starts at the document root. Once the context is the section element, use .//book to search within that subtree. Omitting the dot changes the meaning: an absolute expression is evaluated from the document, not from the selected element.
Namespaces and HTML parsing
For namespace-heavy XML, pass a prefix map and use that prefix in the expression:
from lxml import etree
xml = b'''<feed xmlns="urn:example:feed">
<entry><title>Item one</title></entry>
</feed>'''
root = etree.fromstring(xml)
ns = {'f': 'urn:example:feed'}
print(root.xpath('//f:entry/f:title/text()', namespaces=ns))
When the vocabulary is unknown, local-name() can match by local part, but explicit namespaces are safer when you know the document schema because they avoid accidental matches. For HTML, parse the markup into an lxml tree first, then evaluate the same XPath methods.
Selenium: XPath against a live browser DOM
Selenium exposes XPath through By.XPATH. Use it when the page is generated or modified by JavaScript, or when you need to interact with the rendered document rather than a downloaded file.
A complete locator pattern with an explicit wait
from selenium import webdriver
from selenium.webdriver.common.by import By
from selenium.webdriver.support.ui import WebDriverWait
from selenium.webdriver.support import expected_conditions as EC
url = 'https://example.com'
driver = webdriver.Chrome()
wait = WebDriverWait(driver, 15)
try:
driver.get(url)
heading = wait.until(
EC.visibility_of_element_located(
(By.XPATH, "//h1[normalize-space()]")
)
)
print(heading.text)
finally:
driver.quit()
The wait checks both presence and state before reading the element. In a real application, replace the example expression with a selector for your page and include the XPath in any failure message so a broken locator can be diagnosed quickly.
Scope a child query with a relative XPath
from selenium.webdriver.common.by import By
login = driver.find_element(By.XPATH, "//form[@id='loginForm']")
username = login.find_element(By.XPATH, ".//input[@name='username']")
submit = login.find_element(
By.XPATH,
".//input[@continue and @type='submit']",
)
The dot in .//input keeps the search inside the form. Without it, a document-wide search could match an input belonging to another form.
The final example intentionally illustrates the structure; replace @continue with the attribute that actually exists in your markup, such as @name='continue'. A corrected submit locator would be:
submit = login.find_element(
By.XPATH,
".//input[@name='continue' and @type='submit']",
)
XPath patterns that remain readable
Anchor to stable attributes
- Prefer a unique, predictable
idwhen the application treats it as a stable contract. - Use meaningful
name,aria-label, or other semantic attributes when IDs are unavailable. - Combine predicates for a specific control, such as
//input[@name='email' and @type='email'].
Keep expressions short. Generated CSS classes and long positional paths are coupled to implementation details and tend to break during harmless markup changes.
Use relationships when no single attribute is unique
# A label relationship is often clearer than a deep absolute path
email = driver.find_element(
By.XPATH,
"//label[normalize-space()='Email']/following::input[1]",
)
# Select a button inside the article that owns a heading
button = driver.find_element(
By.XPATH,
"//h2[normalize-space()='Settings']/ancestor::section[1]//button",
)
Relationship-based selectors express the page’s meaning, but keep the relationship as local as possible so a new section elsewhere does not create an accidental match.
Text predicates and quoting
normalize-space() removes surrounding whitespace and collapses runs of internal whitespace, making text tests less sensitive to formatting. For text containing an apostrophe, use double quotes around the XPath literal; for text containing both quote types, construct an XPath concat() expression or match a stable attribute instead.
Why an XPath returns nothing
- Wrong context: a query intended for a selected element must begin with
., for example.//input, not//input. - Namespace mismatch: unprefixed names do not match namespaced XML. Use ElementTree’s qualified name or an explicit lxml namespace map.
- Page not ready: in Selenium, wait for the element and the state you need before querying or clicking.
- Overly specific path: replace an absolute path such as
/html/body/form[1]with a stable ID, name, label, or nearby relationship. - Wrong result type: element paths return nodes,
text()returns strings, and functions such ascount()return numbers. Do not call element methods on a scalar result. - Selector tested against different markup: inspect the exact parsed XML or current browser DOM. A server response may differ from what JavaScript later renders.
Reduce a failing expression incrementally
from selenium.webdriver.common.by import By
xpath = "//form[@id='loginForm']//input[@name='username']"
# Test the stable anchor first
form = driver.find_elements(By.XPATH, "//*[@id='loginForm']")
if not form:
raise RuntimeError("No element matched //*[@id='loginForm']")
# Then add the child condition
fields = form[0].find_elements(By.XPATH, ".//input[@name='username']")
if not fields:
raise RuntimeError(f"No element matched {xpath}")
Testing the smallest stable predicate first tells you whether the failure is the anchor, the relationship, or the page state.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Performance, reliability, and maintenance
Choose the lowest-cost context
- ElementTree and lxml run locally on an already parsed tree; there is no browser startup or WebDriver round trip.
- lxml is the practical choice for repeated complex queries. Compile a frequently reused expression with
etree.XPath()rather than rebuilding it in a tight loop. - Selenium is heavier because it drives a browser and must wait for live state. Reuse one driver for a related workflow, but always quit it in a
finallyblock.
Make selectors a deliberate interface
Keep XPath strings in one place, give each locator a descriptive name, and include the expression in error logs. Prefer semantic attributes supplied by the application over generated classes or indexes. Use positional indexes only when the document contract guarantees the order.
Validate before changing the selector
When a query suddenly returns zero results, first verify the document or browser state, then the namespace and context. Changing a working selector to a longer absolute path can hide the real timing or markup problem and make future changes more expensive.
Best Value
Or skip the browser setup
If your goal is a clean image or PDF of a page rather than extracting DOM nodes, ScreenshotNeo can capture the URL with one request. It is not an XPath engine, so keep ElementTree, lxml, or Selenium for data extraction and interaction; use ScreenshotNeo when the deliverable is a screenshot.
Its API accepts consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and each response identifies the page verdict and billing state with X-Page-Verdict and X-Billed headers. An MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients.
One-call examples
See the ScreenshotNeo documentation for request details.
Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchWindows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallcurl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`Screenshot failed: ${res.status}`);
const image = Buffer.from(await res.arrayBuffer());
ScreenshotNeo includes full-page capture with lazy images loaded, CSS-selector element capture, device presets and custom viewports, retina scale, PDF paper and margin controls, custom CSS and JavaScript, click-before-capture actions, waits, request blocking, headers, cookies, user agents, authorization, timezone and geolocation controls, transparent backgrounds, resizing, selectable cache TTLs, signed links, asynchronous jobs with signed webhooks, bulk capture for up to 100 URLs per call, a usage API, and an OpenAPI specification. Its parameter names are compatible with those used by other screenshot APIs, which can simplify migration.
The Free plan includes 1,000 screenshots per month without a card. Paid plans are Starter ($5 for 3,000), Growth ($15 for 15,000), Pro ($39 for 60,000), Scale ($99 for 250,000), and Business ($249 for 1,000,000); yearly billing gives two months free, and every feature is available on every plan. Create a free ScreenshotNeo account to start with the 1,000 included screenshots.
Frequently Asked Questions
Can XPath select an attribute instead of an element?
Yes. In lxml, an expression such as //book/@id returns attribute strings. ElementTree’s limited XPath support is better suited to selecting the element and reading element.get('id') in Python.
How do I handle an XPath literal containing both quote characters?
Use XPath’s concat() function to assemble the literal, or avoid text matching by selecting a stable attribute. lxml supports concat(); ElementTree does not provide the full function set.
Free tools Windows power users keep installed
One-click scans. No signup required.
Can Selenium XPath cross into an iframe?
No. Switch the driver to the target frame before locating elements, then switch back to the default content when the frame work is complete. The XPath itself is evaluated within the current document context.
Should I use CSS selectors instead?
Use a readable CSS selector when a stable ID or class is sufficient. XPath is especially useful for text conditions and relationships between elements; whichever syntax you choose, keep the locator short and tied to a stable contract.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

