Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
XPath is a language for selecting nodes in a document, including parsed HTML. In Scrapy, you can use it through response.xpath() to extract text, attributes, and elements; Scrapy also supports CSS selectors through response.css(). Use XPath when your match depends on text or document structure, and CSS when a straightforward tag or class selector says what you mean more clearly.
The examples below explain the distinctions that most often trip up people moving from tutorial sites such as book.toscrape.com and quotes.toscrape.com to unfamiliar pages. They focus on selectors and extraction, not on whether a particular site may lawfully be scraped.
What XPath does in a scraper
XPath stands for XML Path Language. It addresses nodes in a structured document; HTML is one of the document types it can be used with. A scraper first parses a page into a document tree, then evaluates a selector against that tree. The selector can identify elements, text nodes, or attribute values.
Windows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallOutdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchIn Scrapy, the response object exposes both XPath and CSS selector methods. The expression is passed as a string, and the result can be reduced to one value or collected as a list:
#1 Best Overall
response.xpath("//title/text()").get()selects title text and returns the first result.response.xpath("//a/@href").getall()selects every link’shrefattribute and returns all results.response.css("title::text").get()selects title text with CSS syntax.
Scrapy’s .get() returns a single result (the first match), while .getall() returns all matching results. If a query finds nothing, .get() has no selected value to return; account for absent fields rather than assuming every page has the same markup.
To learn a selector, inspect the actual HTML around the value you need. Identify the element containing it, determine whether the value is text or an attribute, and then test a narrow selector before extracting a whole page of fields. A tutorial page is useful practice, but production pages often have extra wrappers, repeated components, or markup that differs from the example you started with.
How to choose between XPath and CSS
Scrapy supports both, so the choice is about expressing the match accurately and keeping it understandable. CSS is often concise for a tag, class, or straightforward descendant match. XPath is useful when the query needs to inspect text, navigate structural relationships, or test attributes as part of the selection.
| Need | Reasonable starting point | Example |
|---|---|---|
| Text inside a known element | Either, depending on the markup | response.xpath("//title/text()").get() |
| Attribute values such as link destinations | XPath makes the attribute target explicit | response.xpath("//a/@href").getall() |
| Simple tag or class matching | CSS is often direct | response.css(".product-title::text").get() |
| Match an element using its text or a structural relationship | XPath can express the condition in the selector | Use a text predicate or a path through the relevant parent and child elements. |
This is a readability decision, not a speed guarantee: the cited Scrapy and XPath documentation establishes that both selector styles are available, but does not establish that one is universally faster. Prefer the shortest expression that clearly describes the intended match and is easy for the next maintainer to verify.
Extract text and attributes without mixing them up
HTML elements can contain text, nested elements, and attributes. Decide which of those the field actually represents. In XPath, /text() selects text-node children, while /@href selects the named attribute.
# First title text node
page_title = response.xpath("//title/text()").get()
# All href attributes from links
links = response.xpath("//a/@href").getall()
# All text-node children of paragraphs
paragraph_parts = response.xpath("//p/text()").getall()
A common surprise is that //p/text() does not mean “all visible text anywhere inside each paragraph.” It selects text nodes that are direct children of the matched paragraph elements. If a paragraph includes nested markup such as <strong>, text inside that child is a separate text node. When the field is the element’s combined text, select or test the element itself as described below rather than assuming a direct-child text query covers its descendants.
Why nested selectors need a relative path
A selector can be run against a response or against a selected element. In a nested query, a path beginning with / is still absolute to the whole document. That can make a selector appear to ignore the element you already chose.
# Find each product card
cards = response.xpath("//article")
for card in cards:
# Relative: search inside this card
date = card.xpath("./time/@datetime").get()
paragraphs = card.xpath(".//p").getall()
The leading dot marks a relative expression. ./time/@datetime looks for a time element in the selected context, while .//p searches for paragraph descendants within that context. Use a relative path whenever each nested result must stay attached to the current card, row, article, or other selected container.
Rank #3
Without the dot, a nested expression such as //time/@datetime can search the document rather than just the current card. In a loop, that can return a date from somewhere else on the page and associate it with the wrong record. When results look identical for every card, check whether the inner XPath is absolute.
What does [1] mean in XPath?
The position predicate applies to the location step it follows. Consequently, //li[1] and (//li)[1] do not mean the same thing.
//li[1]selects anlithat is first among the matchinglisiblings under its parent. A document with several lists can therefore yield a first item from each list.(//li)[1]first forms the document-wide set of matchinglielements, then selects the first one in that set.
If you mean “the first item in this particular list,” select the list first and use a relative query, for example ul.xpath("./li[1]") when ul is already a selected element. If you mean “the first matching list item on the page,” grouping the full expression as (//li)[1] makes that scope explicit.
How to match text that includes nested elements
Suppose an anchor’s label is split across text and a nested element:
<a href="/next">Next <strong>Page</strong></a>
Its text is not necessarily one direct text node. Scrapy cautions against using contains(.//text(), 'Next Page') for this case: when a node-set is converted to a string for the function, that conversion can inspect only its first text node, not concatenate all descendant text nodes as the query’s author might expect.
Use contains(., 'Next Page') to test the anchor’s aggregate descendant text:
next_link = response.xpath("//a[contains(., 'Next Page')]/@href").get()
The dot refers to the context element, and the string value of that element includes descendant text. This is useful for selecting an element based on a label that crosses nested markup. Keep the surrounding element constraint specific enough that a matching phrase elsewhere on the page does not select an unintended link.
Free tools Windows power users keep installed
One-click scans. No signup required.
A practical selector-debugging workflow
- Confirm the page contains the target markup. Inspect the HTML available to the scraper and locate the precise element, text, or attribute. A selector cannot match content that is absent from the parsed document.
- Start with the element. Test a broad, recognizable tag or class in
response.xpath()orresponse.css(), then narrow it to the record container you need. - Choose text or an attribute deliberately. Use XPath’s
/text()for direct text nodes and/@attributefor attributes. Remember that nested child text is separate from direct text. - Check selector scope. In a query on a selected element, use
.for a relative path. Confirm whether a positional predicate applies per parent or to a grouped document-wide selection. - Test both one and many matches. Use
.get()to inspect a single expected value and.getall()to check whether repeated items are selected as intended. - Validate against more than one record or page. A selector that works for the first item may be accidentally global or may rely on markup not shared by other records.
Common XPath and Scrapy selector failures
| Symptom | Likely cause | What to check |
|---|---|---|
| The same value appears for every item in a loop | The nested path is absolute and searches the whole document. | Use a relative path beginning with ., such as ./time/@datetime or .//p. |
| A text-based selector misses an element with nested markup | The text is split among multiple descendant text nodes. | For a combined-text test, use contains(., 'phrase') rather than treating .//text() as concatenated text. |
| A “first item” selector returns several elements | //li[1] selects first matching siblings under multiple parents. |
Use (//li)[1] for the first document-wide match, or scope to the intended parent first. |
| A field is missing or empty | The target may be absent, represented as an attribute rather than text, nested below another element, or not present in the parsed HTML. | Inspect the markup and try the relevant text or attribute selector; handle a missing result in the scraper. |
| A broad selector returns unrelated matches | The selector does not identify a specific record or component. | Anchor the query to a distinctive parent, then use a relative selector for the field inside it. |
Scraping etiquette: what robots.txt does and does not mean
RFC 9309, the IETF’s September 2022 Robots Exclusion Protocol specification, describes crawler rules published in robots.txt. It says crawlers that successfully retrieve the file must follow parseable rules. It also states: “These rules are not a form of access authorization.” A robots.txt file is a crawler protocol, not a grant of permission or a complete legal decision.
Best Value
Whether a particular scraping project is permitted can depend on the site’s terms, the data, purpose, authentication, jurisdiction, and other circumstances. A robots.txt rule alone does not settle those questions. For consequential projects, assess the specific situation and applicable law with qualified counsel.
Or skip the browser setup
If your goal is to capture a page as an image or PDF rather than extract structured fields, a screenshot API is a different tool from XPath. ScreenshotNeo accepts a URL in one GET request and returns an image or PDF. For selector learning and data extraction, continue using Scrapy; a screenshot does not replace parsing HTML into fields.
Here is a cURL example saving a WebP capture of the same kind of page:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
See the ScreenshotNeo documentation for request options. ScreenshotNeo removes cookie banners, newsletter popups, and chat widgets before the capture; bot checks, blank pages, and failed loads are never billed. Its MCP server lets AI agents take screenshots. The Free plan includes 1,000 screenshots per month with no card, and paid plans start at $5 for 3,000.
Sign up for 1,000 free screenshots a month, with no card required.
Frequently Asked Questions
Does XPath work only with XML?
No. XPath is used to address nodes in XML and other XML-like documents, including HTML and SVG.
Is robots.txt permission to scrape a site?
No. RFC 9309 explicitly says robots.txt rules are not access authorization; they do not, by themselves, determine whether a project is lawful.
Recommended Free Tools
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

