XPath lets you select elements, text nodes, attributes, and structural relationships in a parsed HTML document. In Scrapy, start with response.xpath(...), then use .get() for one result or .getall() for all results. The most important scope rule is simple: // searches from the document, while .// searches beneath the current element.
What XPath does in an HTML scraper
XPath is a language for addressing parts of a document tree. The W3C XPath 1.0 Recommendation, published on 16 November 1999, describes it as “a language for addressing parts of an XML document, designed to be used by both XSLT and XPointer.” An HTML parser builds a tree from the response markup; XPath expressions select nodes from that tree. XPath does not fetch a page or execute its JavaScript by itself.
The examples below use Scrapy’s selector API and assume response is an HTML response. Scrapy selectors are a thin wrapper over Parsel, which uses lxml underneath. Parsel can also be used independently of Scrapy; lxml parses HTML and XML but is not included in Python’s standard library.
Core XPath expressions for scraping
| Goal | XPath | Result and use |
|---|---|---|
| Select all H1 elements | //h1 |
Element selectors for each matching heading. |
| Extract text nodes directly inside H1 | //h1/text() |
Text nodes that are direct children of each H1; nested child text is not included by this step. |
| Get link destinations | //a/@href |
The href attribute values of matching anchors. |
| Find links whose href contains a value | //a[contains(@href, "image")]/@href |
Anchor hrefs containing the substring image. |
| Select a div by ID | //div[@id="images"] |
Div elements whose ID exactly matches. |
| Get the title text | //title/text() |
Use .get() for one string, or .getall() for all matches. |
| Get image source attributes | //img/@src |
All matching src values when retrieved with .getall(). |
| Find paragraphs below the current selector | .//p |
Paragraph descendants of the current selected element. |
Extract one value or a list in Scrapy
response.xpath() returns selector objects, not plain strings. Call .get() to serialize the first match, or .getall() to serialize all matches into a list. If there is no match, .get() returns None unless you pass a default value.
#1 Best Overall
- HTML CSS Design and Build Web Sites
- Comes with secure packaging
- It can be a gift option
title = response.xpath("//title/text()").get()
image_urls = response.xpath("//img/@src").getall()
first_title_or_empty = response.xpath("//title/text()").get(default="")
Choose based on what the next stage needs: a scalar for a field expected once, a list for repeated values, and an explicit default when absence should become a particular value. Do not assume .get() validates uniqueness; if the page has multiple matches, it silently gives you the first.
Understand //, .//, and relative scope
In a nested selector, // begins a document-level search. The dot in .// anchors the search under the current node. This distinction matters when iterating through repeated containers.
Rank #2
for card in response.xpath("//article"):
# Searches the document, not just this card:
all_paragraphs = card.xpath("//p").getall()
# Searches descendants of this card:
card_paragraphs = card.xpath(".//p").getall()
# Searches only direct child paragraphs:
direct_paragraphs = card.xpath("p").getall()
Use .//p when a paragraph may appear at any depth inside the selected container. Use p when only immediate child paragraphs are wanted. A document-level query inside a loop can repeat the same site-wide results for every container, producing plausible-looking but incorrect output.
Position predicates: first per parent or first overall?
Predicates apply to the node context they are written in. Consequently, //li[1] selects each li that is the first matching li child under its relevant parent. A page with several lists can therefore produce several results. To select the first li in the document-wide result set, write (//li)[1].
Quick wins for a faster PC:
Repair Windows errors before they cause bigger problemsFix Now →Scan for outdated or missing drivers - takes under a minuteDriver Scan →Clear out junk files and repair common Windows errorsFree Scan →Rank #3
first_item_in_each_list = response.xpath("//li[1]").getall()
first_item_in_document = response.xpath("(//li)[1]").get()
Parentheses change the order of operations: first evaluate the full set of li matches, then take its first node. Whenever a positional result looks unexpectedly numerous, check which parent context the predicate applies to.
Text nodes, nested markup, and element string values
text() selects text-node children; .//text() selects text nodes below the current element at any depth. This is useful when you need separate text fragments, but those fragments may be split by nested markup. When checking an element’s combined text, use its string value, written as ..
Rank #4
- Brand: Wiley
- Set of 2 Volumes
- A handy two-book set that uniquely combines related technologies Highly visual format and accessible language makes these books highly effective learning tools Perfect for beginning web designers and front-end developers
# Individual descendant text nodes:
response.xpath("//a//text()").getall()
# Test combined descendant text, even if nested markup splits it:
response.xpath("//a[contains(., 'Next Page')]").getall()
A string function applied to a node set such as .//text() can convert only the first text node to a string. Thus contains(.//text(), 'Next Page') may fail when the words are split across nested elements, while contains(., 'Next Page') tests the element’s combined string value.
Match class tokens safely
HTML class attributes can contain several whitespace-separated tokens. An exact comparison such as //*[@class='product'] misses an element whose class is product featured. A raw substring check such as contains(@class, 'prod') can match an unintended token like product-card.
Best Value
For a token-safe match, normalize whitespace and pad both sides so the token is checked as a whole:
//*[contains(concat(' ', normalize-space(@class), ' '), ' product ')]
For straightforward class selection, Scrapy’s CSS selector can be more readable; CSS queries are translated into XPath internally. You can select by CSS first, then chain XPath when you need text nodes, attributes, or more complex structural predicates.
Namespaces, parser choice, and JavaScript-rendered pages
These are separate problems from XPath syntax. A correct XPath expression can still return nothing if it does not match the parsed tree or the response type.
- Namespaces: An XML feed may put elements in a namespace, so a namespace-free expression such as
//linkmay not match. Use namespace-aware queries with mappings, or deliberately remove namespaces with Scrapy’sremove_namespaces(). Removing namespaces changes the tree and has a processing cost. - Parser and response type: Scrapy documents response type selection and namespace behavior. Check that the response is parsed as the type you expect and inspect the parsed content rather than assuming the server returned ordinary HTML.
- JavaScript content: XPath addresses the tree the parser receives. It does not make a client-rendered page populate missing elements; when needed content is absent from the parsed response, you need a workflow that supplies rendered page content before extraction.
Choose XPath or CSS for the job
| Need | Often a good choice | Why |
|---|---|---|
| Simple class-based selection | CSS | Often easier to read for common class selectors. |
| Text nodes or attributes | XPath | Expressions such as //a/@href and //p/text() address these directly. |
| Structural relationships or predicates | XPath | Supports context, positional predicates, and other query logic. |
| Repeated nested containers | Either, with explicit scope | Choose based on readability and ensure the selector is relative to the current container when appropriate. |
There is no basis here for a blanket speed claim for one selector style. Choose by readability, needed node type, parser behavior, scraper integration, and scope. Scrapy offers both response.css() and response.xpath(); Parsel is available without Scrapy if that better fits your workflow.
PC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minuteA practical extraction workflow
- Inspect the response content. Confirm that the response contains the element you intend to extract and that the parser is handling it as HTML or XML as appropriate.
- Anchor on a stable structure. Use a meaningful ID or container where available, rather than relying on a fragile page-wide positional match.
- Decide the scope. Use
//for a document search,.//under a selected node, and a bare child name such aspfor direct children. - Select the right node type. Use an element expression for elements,
/@attributefor an attribute, and/text()for direct text nodes. - Choose cardinality deliberately. Use
.get()for one serialized match or.getall()for a list; provide a default if a missing single value needs one. - Test awkward cases. Check pages with absent fields, multiple containers, nested markup, and classes containing more than one token.
Troubleshooting common XPath misses
- Nested loop returns the same paragraphs for every container: replace
//pwith.//pfor descendants of the current selector, orpfor direct children. //li[1]returns several results: that expression selects the first matching child in each parent context. Use(//li)[1]for the first item in the overall document result.- Exact class match finds nothing: the element may have multiple class tokens. Use the token-safe normalized expression or select the class with CSS.
- Class substring match finds the wrong element: substring matching does not respect token boundaries. Use the padded
normalize-space(@class)form. - Text test fails around bold or span markup: use
.to test combined descendant text; use.//text()only when individual text nodes are what you need. .get()givesNone: no node matched in the parsed response. Check the response content, selector scope, HTML/XML type, and namespace before changing the expression blindly.- Only one result appears though several were expected: use
.getall()and confirm the XPath matches all intended nodes;.get()returns only the first result. - Namespace-free query misses XML elements: query with the namespace mapping or deliberately remove namespaces before querying.
- Rendered content is absent: XPath cannot select nodes that are not in the parsed tree. Arrange for the page content to be rendered or otherwise present in the response before extracting it.
Or skip the browser setup
When your workflow needs a screenshot or PDF rather than HTML nodes to query, ScreenshotNeo provides a one-request capture API. It is not an XPath parser: use Scrapy and XPath for structured extraction, and use this endpoint when the output you need is a rendered image.
Quick Recap
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
See the ScreenshotNeo API documentation for request options. Before capture, it accepts cookie or consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing, with response headers indicating the page verdict and billing status. Its MCP server offers take_screenshot, get_page_info, and capture_pdf for AI agents. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. Sign up for the free plan.
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




