Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

Use XPath when you need to select nodes in an HTML or XML tree by structure, attributes, text, or relationships such as parent and ancestor. In Scrapy, a typical extraction is response.xpath("//span/text()").get() for one value or response.xpath("//a/@href").getall() for every link. The details that prevent most scraping bugs are context-relative paths, correctly scoped position predicates, and stable selectors.

What XPath does in a scraper

XPath is an expression language for addressing nodes in XML-derived data models. The W3C XPath 1.0 Recommendation was published on 16 November 1999, and browser-oriented DOM guidance still describes XPath 1.0 as the simple mechanism for accessing a DOM tree. Scrapy’s documentation summarizes the practical use: XPath is a language for selecting nodes in XML documents that can also be used with HTML.

An XPath expression can return elements, text nodes, attributes, or a calculated value. A path is made from steps separated by /; predicates in square brackets filter the result. The expression is evaluated against a document or against a selected subtree, depending on the API call that starts it.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The pieces you use most

  • //article finds article elements anywhere below the current context.
  • /html/body/main follows a specific child path from the document root.
  • @href selects an attribute, such as a link destination.
  • text() selects direct text-node children; string(.) converts all descendant text in the current node to one string.
  • [condition] applies a predicate, for example [@data-id="42"].
  • .. moves to a parent, while ancestor::section selects matching ancestors.

Basic extraction with Scrapy and Parsel

Scrapy selectors and the stand-alone Parsel library use lxml underneath. lxml parses both HTML and XML, so the same XPath concepts work in a spider, a shell session, or a small script.

Get one text value

title = response.xpath("//h1/text()").get()
# Example result: "Product name"

.get() returns the first serialized result or None when there is no match. It does not raise an error merely because a page changed.

Get all values

links = response.xpath("//a/@href").getall()
labels = response.xpath("//nav//a//text()").getall()

.getall() returns every match as a list. Text may be split across several descendant nodes, so use ::text with CSS only for direct text or use XPath’s string(.) when you need the complete visible string from an element.

Build a structured item

def parse(self, response):
    for card in response.xpath("//article[contains(@class, 'card')]"):
        yield {
            "title": card.xpath(".//h2//text()").get(default="").strip(),
            "url": card.xpath(".//a[1]/@href").get(),
            "price": card.xpath(".//*[contains(@class, 'price')]//text()").get(),
        }

The dot before each nested path is intentional. It keeps the query inside the current card selector.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Relative versus absolute XPath in nested selections

This is the most common reason a nested XPath returns unrelated elements. Suppose you select cards first:

cards = response.xpath("//div[contains(@class, 'card')]")

Then card.xpath("//p") starts at the document root and finds every paragraph on the page for every card. Use a relative path beginning with .:

for card in cards:
    summary = card.xpath(".//p[contains(@class, 'summary')]//text()").getall()
    published = card.xpath("./time/@datetime").get()

.//p searches descendants of the current card. ./time selects an immediate child named time. A leading slash without a dot resets the search to the document root, so reserve absolute paths for cases where you deliberately want document-wide context.

Predicates, text tests, and position

Filter by attributes

response.xpath("//a[@rel='next']/@href").get()
response.xpath("//input[@name='q']/@value").get()
response.xpath("//*[@data-testid='result']").getall()

For a class token, avoid exact equality when an element can have multiple classes. This test matches a standalone token:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
contains(concat(' ', normalize-space(@class), ' '), ' product-card ')

Used in context, it becomes //*[contains(concat(' ', normalize-space(@class), ' '), ' product-card ')].

Match text carefully

response.xpath("//button[normalize-space(.)='Load more']").get()
response.xpath("//*[contains(normalize-space(.), 'Shipping')]").getall()

normalize-space(.) collapses runs of whitespace and includes descendant text. That is usually more reliable than text()='...', which checks only direct text nodes and fails when a label contains a nested span.

Understand the two meanings of “first”

//li[1] means the first li under each parent selected by //. It can therefore return one item from many lists. To select the first li in document order, group the whole expression:

first_per_list = response.xpath("//ul/li[1]").getall()
first_in_document = response.xpath("(//li)[1]").get()

Apply the same distinction to links, rows, cards, and any repeated component. If you need the third item within each card, use a relative expression such as .//li[3]; if you need the third item overall, use (//li)[3].

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Extracting links, attributes, and complete text

Links and URLs

hrefs = response.xpath("//a[@href]/@href").getall()
anchor = response.xpath("//a[@href][1]")
url = anchor.xpath("./@href").get()
label = " ".join(anchor.xpath(".//text()").getall()).strip()

Scrapy’s response object can resolve relative links when you use its URL utilities, but XPath itself only reads the attribute value. Treat empty, fragment-only, JavaScript, and malformed URLs as validation cases before following them.

Direct text versus descendant text

//h2/text() returns only direct text nodes. For markup such as <h2>Buy <em>now</em></h2>, use //h2//text() and join the pieces, or select the heading and call string(.):

heading = response.xpath("string(//h2[1])").get()

When whitespace and hidden decorative nodes matter, inspect the HTML and normalize the result in Python rather than assuming every text node is user-visible.

Namespaces and XML feeds

HTML scraping often has no namespace, but XML documents commonly use one. If an element has a prefix, pass a prefix-to-URI mapping and use that prefix in the XPath:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
ns = {
    "atom": "http://www.w3.org/2005/Atom",
    "media": "http://search.yahoo.com/mrss/",
}
entries = response.xpath("//atom:entry", namespaces=ns)
titles = response.xpath("//atom:entry/atom:title/text()", namespaces=ns).getall()

The prefix you choose in the mapping is local to your query; it does not have to match the prefix used in the source document. If the source uses a default namespace, it still needs an explicit prefix in your selector. For mixed or unknown namespaces, inspect the parsed tree and map each URI deliberately.

Regex and implementation extensions

XPath 1.0 itself does not provide the regular-expression functions many scrapers expect. Scrapy pre-registers EXSLT namespaces, including re:test(), through lxml. For example:

response.xpath(
    "//*[re:test(@class, '^price-', 'i')]",
    namespaces={"re": "http://exslt.org/regular-expressions"},
).getall()

This is an implementation extension, not a portable XPath 1.0 feature. The Scrapy documentation notes that lxml’s Python regular-expression hook can add a small performance penalty. For simple matching, prefer ordinary predicates; for complex parsing, extract the attribute and use Python’s regular-expression engine after selection.

XPath or CSS selectors?

Neither syntax wins every task. Scrapy exposes both response.xpath() and response.css(), so a maintainable spider can use the clearest selector for each field.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Need Usually the clearer choice Reason
Stable tag, class, or ID CSS Short syntax is easy for a team to read.
Parent, ancestor, or sibling relationship XPath Axes express structural relationships directly.
Match visible text XPath Predicates such as normalize-space(.) are built in.
XML with namespaces XPath Namespace-aware paths are part of the model.
Browser automation locator Either Selenium says XPath works as well as CSS, but its syntax is often more complicated and harder to debug.

Do not choose on an assumed speed advantage: the cited documentation establishes API and debugging trade-offs, not a general benchmark. Prefer stable IDs, data attributes, or semantic structure over long chains of incidental div elements. Keep selectors short enough that a reviewer can explain what makes a match correct.

Debugging and failure recovery

Zero results

  • Wrong response: log response.url and status, then save the body to verify a redirect, block page, or login form.
  • JavaScript-only content: the server response may not contain the nodes you see in a browser. Use a rendering-capable workflow, wait for the content, or find the underlying data request.
  • Namespace mismatch: add the correct URI mapping for XML.
  • Text split across elements: replace text()='...' with normalize-space(.).
  • Nested context error: add . to paths evaluated on a selected element.

Too many results

  • Scope the path to a card, row, or section before selecting descendants.
  • Use an attribute predicate rather than a broad tag name.
  • Check whether //li[1] is selecting the first item per parent; use (//li)[1] for one global result.

Wrong or unstable values

  • Inspect the raw HTML, not just the rendered inspector, and identify whether classes change between requests.
  • Normalize whitespace and strip results at the Python boundary.
  • Validate missing attributes, relative URLs, and duplicate records before storing them.
  • Write a small fixture test containing the DOM shapes you support. A selector should fail visibly when its required structure disappears, rather than silently producing plausible but incorrect data.

Performance, reliability, and maintainability

Start with one broad selection and perform relative queries on each selected subtree. This makes intent explicit and avoids repeatedly scanning the entire document. Select only the node type you need—an attribute instead of a whole element when collecting URLs—and use .get() when one value is expected.

Use retries and rate limits at the HTTP layer, not inside XPath. XPath cannot solve authentication, robots policies, consent flows, bot challenges, pagination, or content that is absent from the response. Record the URL, status, selector name, and match count for each extraction so a layout change can be diagnosed quickly. When an expression becomes a long chain of positional steps, replace it with stable attributes or a nearby semantic relationship.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup

If your goal is a clean image or PDF of a page rather than DOM data, ScreenshotNeo provides a single HTTP endpoint. Before capture it accepts the cookie or consent banner like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each cleanup step can be disabled. Bot checks and CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing status. Its MCP server exposes take_screenshot, get_page_info, and capture_pdf to Claude, Cursor, and other MCP clients.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use the ScreenshotNeo API documentation for all options. A minimal request is:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

The service supports full-page captures with lazy images loaded, CSS-selector element capture, dark mode, 12 device presets or any viewport, retina scale, PDF paper settings and page ranges, custom CSS and JavaScript, clicks, waits, request blocking, headers, cookies, user agents, authorization, timezone, geolocation, transparent backgrounds, resizing, chosen cache TTLs, signed image links, asynchronous jobs with signed webhooks, bulk capture of up to 100 URLs per call, usage reporting, and an OpenAPI specification. Parameter names used by other screenshot APIs also work.

There is a free allowance of 1,000 screenshots per month with no card. Paid plans start at $5 for 3,000 screenshots; yearly billing gives two months free, and every feature is available on every plan. Create a free ScreenshotNeo account to try it.

FAQ

Can XPath select an element’s parent?

Yes. From a selected node, use .. for its parent or an axis such as ancestor::article[1] for the nearest matching ancestor.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Why does my selector work in a browser but not in Scrapy?

The browser may have executed JavaScript, while Scrapy is inspecting the original HTTP response. Compare the response body with the rendered DOM and locate the data request or use a rendering workflow when necessary.

Is XPath limited to HTML?

No. It was designed for XML-derived data models and is also used with HTML parsers. XML namespaces must be mapped when the document uses them.

Should every scraper standardize on XPath?

No. Use CSS for simple, stable selectors and XPath where text predicates or structural relationships make the intent clearer.

Frequently Asked Questions

Can XPath select an element’s parent?

Yes. From a selected node, use .. for its parent or an axis such as ancestor::article[1] for the nearest matching ancestor.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Why does my selector work in a browser but not in Scrapy?

The browser may have executed JavaScript, while Scrapy is inspecting the original HTTP response. Compare the response body with the rendered DOM and locate the data request or use a rendering workflow when necessary.

Is XPath limited to HTML?

No. It was designed for XML-derived data models and is also used with HTML parsers. XML namespaces must be mapped when the document uses them.

Should every scraper standardize on XPath?

No. Use CSS for simple, stable selectors and XPath where text predicates or structural relationships make the intent clearer.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.