Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
Use XPath when you need to select nodes in an HTML or XML tree by structure, attributes, text, or relationships such as parent and ancestor. In Scrapy, a typical extraction is response.xpath("//span/text()").get() for one value or response.xpath("//a/@href").getall() for every link. The details that prevent most scraping bugs are context-relative paths, correctly scoped position predicates, and stable selectors.
What XPath does in a scraper
XPath is an expression language for addressing nodes in XML-derived data models. The W3C XPath 1.0 Recommendation was published on 16 November 1999, and browser-oriented DOM guidance still describes XPath 1.0 as the simple mechanism for accessing a DOM tree. Scrapy’s documentation summarizes the practical use: XPath is a language for selecting nodes in XML documents that can also be used with HTML.
An XPath expression can return elements, text nodes, attributes, or a calculated value. A path is made from steps separated by /; predicates in square brackets filter the result. The expression is evaluated against a document or against a selected subtree, depending on the API call that starts it.
Free tools Windows power users keep installed
One-click scans. No signup required.
The pieces you use most
//articlefindsarticleelements anywhere below the current context./html/body/mainfollows a specific child path from the document root.@hrefselects an attribute, such as a link destination.text()selects direct text-node children;string(.)converts all descendant text in the current node to one string.[condition]applies a predicate, for example[@data-id="42"]...moves to a parent, whileancestor::sectionselects matching ancestors.
Basic extraction with Scrapy and Parsel
Scrapy selectors and the stand-alone Parsel library use lxml underneath. lxml parses both HTML and XML, so the same XPath concepts work in a spider, a shell session, or a small script.
#1 Best Overall
Get one text value
title = response.xpath("//h1/text()").get()
# Example result: "Product name"
.get() returns the first serialized result or None when there is no match. It does not raise an error merely because a page changed.
Get all values
links = response.xpath("//a/@href").getall()
labels = response.xpath("//nav//a//text()").getall()
.getall() returns every match as a list. Text may be split across several descendant nodes, so use ::text with CSS only for direct text or use XPath’s string(.) when you need the complete visible string from an element.
Build a structured item
def parse(self, response):
for card in response.xpath("//article[contains(@class, 'card')]"):
yield {
"title": card.xpath(".//h2//text()").get(default="").strip(),
"url": card.xpath(".//a[1]/@href").get(),
"price": card.xpath(".//*[contains(@class, 'price')]//text()").get(),
}
The dot before each nested path is intentional. It keeps the query inside the current card selector.
Relative versus absolute XPath in nested selections
This is the most common reason a nested XPath returns unrelated elements. Suppose you select cards first:
cards = response.xpath("//div[contains(@class, 'card')]")
Then card.xpath("//p") starts at the document root and finds every paragraph on the page for every card. Use a relative path beginning with .:
for card in cards:
summary = card.xpath(".//p[contains(@class, 'summary')]//text()").getall()
published = card.xpath("./time/@datetime").get()
.//p searches descendants of the current card. ./time selects an immediate child named time. A leading slash without a dot resets the search to the document root, so reserve absolute paths for cases where you deliberately want document-wide context.
Predicates, text tests, and position
Filter by attributes
response.xpath("//a[@rel='next']/@href").get()
response.xpath("//input[@name='q']/@value").get()
response.xpath("//*[@data-testid='result']").getall()
For a class token, avoid exact equality when an element can have multiple classes. This test matches a standalone token:
contains(concat(' ', normalize-space(@class), ' '), ' product-card ')
Used in context, it becomes //*[contains(concat(' ', normalize-space(@class), ' '), ' product-card ')].
Match text carefully
response.xpath("//button[normalize-space(.)='Load more']").get()
response.xpath("//*[contains(normalize-space(.), 'Shipping')]").getall()
normalize-space(.) collapses runs of whitespace and includes descendant text. That is usually more reliable than text()='...', which checks only direct text nodes and fails when a label contains a nested span.
Understand the two meanings of “first”
//li[1] means the first li under each parent selected by //. It can therefore return one item from many lists. To select the first li in document order, group the whole expression:
first_per_list = response.xpath("//ul/li[1]").getall()
first_in_document = response.xpath("(//li)[1]").get()
Apply the same distinction to links, rows, cards, and any repeated component. If you need the third item within each card, use a relative expression such as .//li[3]; if you need the third item overall, use (//li)[3].
Recommended Free Tools
Extracting links, attributes, and complete text
Links and URLs
hrefs = response.xpath("//a[@href]/@href").getall()
anchor = response.xpath("//a[@href][1]")
url = anchor.xpath("./@href").get()
label = " ".join(anchor.xpath(".//text()").getall()).strip()
Scrapy’s response object can resolve relative links when you use its URL utilities, but XPath itself only reads the attribute value. Treat empty, fragment-only, JavaScript, and malformed URLs as validation cases before following them.
Rank #3
Direct text versus descendant text
//h2/text() returns only direct text nodes. For markup such as <h2>Buy <em>now</em></h2>, use //h2//text() and join the pieces, or select the heading and call string(.):
heading = response.xpath("string(//h2[1])").get()
When whitespace and hidden decorative nodes matter, inspect the HTML and normalize the result in Python rather than assuming every text node is user-visible.
Namespaces and XML feeds
HTML scraping often has no namespace, but XML documents commonly use one. If an element has a prefix, pass a prefix-to-URI mapping and use that prefix in the XPath:
Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchPC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11ns = {
"atom": "http://www.w3.org/2005/Atom",
"media": "http://search.yahoo.com/mrss/",
}
entries = response.xpath("//atom:entry", namespaces=ns)
titles = response.xpath("//atom:entry/atom:title/text()", namespaces=ns).getall()
The prefix you choose in the mapping is local to your query; it does not have to match the prefix used in the source document. If the source uses a default namespace, it still needs an explicit prefix in your selector. For mixed or unknown namespaces, inspect the parsed tree and map each URI deliberately.
Regex and implementation extensions
XPath 1.0 itself does not provide the regular-expression functions many scrapers expect. Scrapy pre-registers EXSLT namespaces, including re:test(), through lxml. For example:
response.xpath(
"//*[re:test(@class, '^price-', 'i')]",
namespaces={"re": "http://exslt.org/regular-expressions"},
).getall()
This is an implementation extension, not a portable XPath 1.0 feature. The Scrapy documentation notes that lxml’s Python regular-expression hook can add a small performance penalty. For simple matching, prefer ordinary predicates; for complex parsing, extract the attribute and use Python’s regular-expression engine after selection.
XPath or CSS selectors?
Neither syntax wins every task. Scrapy exposes both response.xpath() and response.css(), so a maintainable spider can use the clearest selector for each field.
Quick wins for a faster PC:
Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Clear out junk files and repair common Windows errorsFree Scan →| Need | Usually the clearer choice | Reason |
|---|---|---|
| Stable tag, class, or ID | CSS | Short syntax is easy for a team to read. |
| Parent, ancestor, or sibling relationship | XPath | Axes express structural relationships directly. |
| Match visible text | XPath | Predicates such as normalize-space(.) are built in. |
| XML with namespaces | XPath | Namespace-aware paths are part of the model. |
| Browser automation locator | Either | Selenium says XPath works as well as CSS, but its syntax is often more complicated and harder to debug. |
Do not choose on an assumed speed advantage: the cited documentation establishes API and debugging trade-offs, not a general benchmark. Prefer stable IDs, data attributes, or semantic structure over long chains of incidental div elements. Keep selectors short enough that a reviewer can explain what makes a match correct.
Debugging and failure recovery
Zero results
- Wrong response: log
response.urland status, then save the body to verify a redirect, block page, or login form. - JavaScript-only content: the server response may not contain the nodes you see in a browser. Use a rendering-capable workflow, wait for the content, or find the underlying data request.
- Namespace mismatch: add the correct URI mapping for XML.
- Text split across elements: replace
text()='...'withnormalize-space(.). - Nested context error: add
.to paths evaluated on a selected element.
Too many results
- Scope the path to a card, row, or section before selecting descendants.
- Use an attribute predicate rather than a broad tag name.
- Check whether
//li[1]is selecting the first item per parent; use(//li)[1]for one global result.
Wrong or unstable values
- Inspect the raw HTML, not just the rendered inspector, and identify whether classes change between requests.
- Normalize whitespace and strip results at the Python boundary.
- Validate missing attributes, relative URLs, and duplicate records before storing them.
- Write a small fixture test containing the DOM shapes you support. A selector should fail visibly when its required structure disappears, rather than silently producing plausible but incorrect data.
Performance, reliability, and maintainability
Start with one broad selection and perform relative queries on each selected subtree. This makes intent explicit and avoids repeatedly scanning the entire document. Select only the node type you need—an attribute instead of a whole element when collecting URLs—and use .get() when one value is expected.
Use retries and rate limits at the HTTP layer, not inside XPath. XPath cannot solve authentication, robots policies, consent flows, bot challenges, pagination, or content that is absent from the response. Record the URL, status, selector name, and match count for each extraction so a layout change can be diagnosed quickly. When an expression becomes a long chain of positional steps, replace it with stable attributes or a nearby semantic relationship.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Or skip the browser setup
If your goal is a clean image or PDF of a page rather than DOM data, ScreenshotNeo provides a single HTTP endpoint. Before capture it accepts the cookie or consent banner like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each cleanup step can be disabled. Bot checks and CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing status. Its MCP server exposes take_screenshot, get_page_info, and capture_pdf to Claude, Cursor, and other MCP clients.
Use the ScreenshotNeo API documentation for all options. A minimal request is:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
The service supports full-page captures with lazy images loaded, CSS-selector element capture, dark mode, 12 device presets or any viewport, retina scale, PDF paper settings and page ranges, custom CSS and JavaScript, clicks, waits, request blocking, headers, cookies, user agents, authorization, timezone, geolocation, transparent backgrounds, resizing, chosen cache TTLs, signed image links, asynchronous jobs with signed webhooks, bulk capture of up to 100 URLs per call, usage reporting, and an OpenAPI specification. Parameter names used by other screenshot APIs also work.
Best Value
There is a free allowance of 1,000 screenshots per month with no card. Paid plans start at $5 for 3,000 screenshots; yearly billing gives two months free, and every feature is available on every plan. Create a free ScreenshotNeo account to try it.
FAQ
Can XPath select an element’s parent?
Yes. From a selected node, use .. for its parent or an axis such as ancestor::article[1] for the nearest matching ancestor.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Why does my selector work in a browser but not in Scrapy?
The browser may have executed JavaScript, while Scrapy is inspecting the original HTTP response. Compare the response body with the rendered DOM and locate the data request or use a rendering workflow when necessary.
Is XPath limited to HTML?
No. It was designed for XML-derived data models and is also used with HTML parsers. XML namespaces must be mapped when the document uses them.
Should every scraper standardize on XPath?
No. Use CSS for simple, stable selectors and XPath where text predicates or structural relationships make the intent clearer.
Frequently Asked Questions
Can XPath select an element’s parent?
Yes. From a selected node, use .. for its parent or an axis such as ancestor::article[1] for the nearest matching ancestor.
The Tool Desk
Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Why does my selector work in a browser but not in Scrapy?
The browser may have executed JavaScript, while Scrapy is inspecting the original HTTP response. Compare the response body with the rendered DOM and locate the data request or use a rendering workflow when necessary.
Is XPath limited to HTML?
No. It was designed for XML-derived data models and is also used with HTML parsers. XML namespaces must be mapped when the document uses them.
Should every scraper standardize on XPath?
No. Use CSS for simple, stable selectors and XPath where text predicates or structural relationships make the intent clearer.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

