Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

XPath is a language for selecting nodes in a document, including parsed HTML. In Scrapy, you can use it through response.xpath() to extract text, attributes, and elements; Scrapy also supports CSS selectors through response.css(). Use XPath when your match depends on text or document structure, and CSS when a straightforward tag or class selector says what you mean more clearly.

The examples below explain the distinctions that most often trip up people moving from tutorial sites such as book.toscrape.com and quotes.toscrape.com to unfamiliar pages. They focus on selectors and extraction, not on whether a particular site may lawfully be scraped.

What XPath does in a scraper

XPath stands for XML Path Language. It addresses nodes in a structured document; HTML is one of the document types it can be used with. A scraper first parses a page into a document tree, then evaluates a selector against that tree. The selector can identify elements, text nodes, or attribute values.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

In Scrapy, the response object exposes both XPath and CSS selector methods. The expression is passed as a string, and the result can be reduced to one value or collected as a list:

  • response.xpath("//title/text()").get() selects title text and returns the first result.
  • response.xpath("//a/@href").getall() selects every link’s href attribute and returns all results.
  • response.css("title::text").get() selects title text with CSS syntax.

Scrapy’s .get() returns a single result (the first match), while .getall() returns all matching results. If a query finds nothing, .get() has no selected value to return; account for absent fields rather than assuming every page has the same markup.

To learn a selector, inspect the actual HTML around the value you need. Identify the element containing it, determine whether the value is text or an attribute, and then test a narrow selector before extracting a whole page of fields. A tutorial page is useful practice, but production pages often have extra wrappers, repeated components, or markup that differs from the example you started with.

How to choose between XPath and CSS

Scrapy supports both, so the choice is about expressing the match accurately and keeping it understandable. CSS is often concise for a tag, class, or straightforward descendant match. XPath is useful when the query needs to inspect text, navigate structural relationships, or test attributes as part of the selection.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Need Reasonable starting point Example
Text inside a known element Either, depending on the markup response.xpath("//title/text()").get()
Attribute values such as link destinations XPath makes the attribute target explicit response.xpath("//a/@href").getall()
Simple tag or class matching CSS is often direct response.css(".product-title::text").get()
Match an element using its text or a structural relationship XPath can express the condition in the selector Use a text predicate or a path through the relevant parent and child elements.

This is a readability decision, not a speed guarantee: the cited Scrapy and XPath documentation establishes that both selector styles are available, but does not establish that one is universally faster. Prefer the shortest expression that clearly describes the intended match and is easy for the next maintainer to verify.

Extract text and attributes without mixing them up

HTML elements can contain text, nested elements, and attributes. Decide which of those the field actually represents. In XPath, /text() selects text-node children, while /@href selects the named attribute.

# First title text node
page_title = response.xpath("//title/text()").get()

# All href attributes from links
links = response.xpath("//a/@href").getall()

# All text-node children of paragraphs
paragraph_parts = response.xpath("//p/text()").getall()

A common surprise is that //p/text() does not mean “all visible text anywhere inside each paragraph.” It selects text nodes that are direct children of the matched paragraph elements. If a paragraph includes nested markup such as <strong>, text inside that child is a separate text node. When the field is the element’s combined text, select or test the element itself as described below rather than assuming a direct-child text query covers its descendants.

Why nested selectors need a relative path

A selector can be run against a response or against a selected element. In a nested query, a path beginning with / is still absolute to the whole document. That can make a selector appear to ignore the element you already chose.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
# Find each product card
cards = response.xpath("//article")

for card in cards:
    # Relative: search inside this card
    date = card.xpath("./time/@datetime").get()
    paragraphs = card.xpath(".//p").getall()

The leading dot marks a relative expression. ./time/@datetime looks for a time element in the selected context, while .//p searches for paragraph descendants within that context. Use a relative path whenever each nested result must stay attached to the current card, row, article, or other selected container.

Without the dot, a nested expression such as //time/@datetime can search the document rather than just the current card. In a loop, that can return a date from somewhere else on the page and associate it with the wrong record. When results look identical for every card, check whether the inner XPath is absolute.

What does [1] mean in XPath?

The position predicate applies to the location step it follows. Consequently, //li[1] and (//li)[1] do not mean the same thing.

  • //li[1] selects an li that is first among the matching li siblings under its parent. A document with several lists can therefore yield a first item from each list.
  • (//li)[1] first forms the document-wide set of matching li elements, then selects the first one in that set.

If you mean “the first item in this particular list,” select the list first and use a relative query, for example ul.xpath("./li[1]") when ul is already a selected element. If you mean “the first matching list item on the page,” grouping the full expression as (//li)[1] makes that scope explicit.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How to match text that includes nested elements

Suppose an anchor’s label is split across text and a nested element:

<a href="/next">Next <strong>Page</strong></a>

Its text is not necessarily one direct text node. Scrapy cautions against using contains(.//text(), 'Next Page') for this case: when a node-set is converted to a string for the function, that conversion can inspect only its first text node, not concatenate all descendant text nodes as the query’s author might expect.

Use contains(., 'Next Page') to test the anchor’s aggregate descendant text:

next_link = response.xpath("//a[contains(., 'Next Page')]/@href").get()

The dot refers to the context element, and the string value of that element includes descendant text. This is useful for selecting an element based on a label that crosses nested markup. Keep the surrounding element constraint specific enough that a matching phrase elsewhere on the page does not select an unintended link.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A practical selector-debugging workflow

  1. Confirm the page contains the target markup. Inspect the HTML available to the scraper and locate the precise element, text, or attribute. A selector cannot match content that is absent from the parsed document.
  2. Start with the element. Test a broad, recognizable tag or class in response.xpath() or response.css(), then narrow it to the record container you need.
  3. Choose text or an attribute deliberately. Use XPath’s /text() for direct text nodes and /@attribute for attributes. Remember that nested child text is separate from direct text.
  4. Check selector scope. In a query on a selected element, use . for a relative path. Confirm whether a positional predicate applies per parent or to a grouped document-wide selection.
  5. Test both one and many matches. Use .get() to inspect a single expected value and .getall() to check whether repeated items are selected as intended.
  6. Validate against more than one record or page. A selector that works for the first item may be accidentally global or may rely on markup not shared by other records.

Common XPath and Scrapy selector failures

Symptom Likely cause What to check
The same value appears for every item in a loop The nested path is absolute and searches the whole document. Use a relative path beginning with ., such as ./time/@datetime or .//p.
A text-based selector misses an element with nested markup The text is split among multiple descendant text nodes. For a combined-text test, use contains(., 'phrase') rather than treating .//text() as concatenated text.
A “first item” selector returns several elements //li[1] selects first matching siblings under multiple parents. Use (//li)[1] for the first document-wide match, or scope to the intended parent first.
A field is missing or empty The target may be absent, represented as an attribute rather than text, nested below another element, or not present in the parsed HTML. Inspect the markup and try the relevant text or attribute selector; handle a missing result in the scraper.
A broad selector returns unrelated matches The selector does not identify a specific record or component. Anchor the query to a distinctive parent, then use a relative selector for the field inside it.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Scraping etiquette: what robots.txt does and does not mean

RFC 9309, the IETF’s September 2022 Robots Exclusion Protocol specification, describes crawler rules published in robots.txt. It says crawlers that successfully retrieve the file must follow parseable rules. It also states: “These rules are not a form of access authorization.” A robots.txt file is a crawler protocol, not a grant of permission or a complete legal decision.

Whether a particular scraping project is permitted can depend on the site’s terms, the data, purpose, authentication, jurisdiction, and other circumstances. A robots.txt rule alone does not settle those questions. For consequential projects, assess the specific situation and applicable law with qualified counsel.

Or skip the browser setup

If your goal is to capture a page as an image or PDF rather than extract structured fields, a screenshot API is a different tool from XPath. ScreenshotNeo accepts a URL in one GET request and returns an image or PDF. For selector learning and data extraction, continue using Scrapy; a screenshot does not replace parsing HTML into fields.

Here is a cURL example saving a WebP capture of the same kind of page:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

See the ScreenshotNeo documentation for request options. ScreenshotNeo removes cookie banners, newsletter popups, and chat widgets before the capture; bot checks, blank pages, and failed loads are never billed. Its MCP server lets AI agents take screenshots. The Free plan includes 1,000 screenshots per month with no card, and paid plans start at $5 for 3,000.

Sign up for 1,000 free screenshots a month, with no card required.

Frequently Asked Questions

Does XPath work only with XML?

No. XPath is used to address nodes in XML and other XML-like documents, including HTML and SVG.

Is robots.txt permission to scrape a site?

No. RFC 9309 explicitly says robots.txt rules are not access authorization; they do not, by themselves, determine whether a project is lawful.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.