Driver FixRecommendedSound, Wi-Fi or graphics acting up? Check drivers firstFind missing or outdated drivers fast.Check DriversOctober DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsClean PCRecommendedOne scan can reveal what keeps slowing WindowsLook for cleanup and repair opportunities.Run Scan×
Skip to content
HTML parsing

Ultimate XPath Cheatsheet for HTML Parsing in Web Scraping

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

XPath lets you select elements, text nodes, attributes, and structural relationships in a parsed HTML document. In Scrapy, start with response.xpath(...), then use .get() for one result or .getall() for all results. The most important scope rule is simple: // searches from the document, while .// searches beneath the current element.

What XPath does in an HTML scraper

XPath is a language for addressing parts of a document tree. The W3C XPath 1.0 Recommendation, published on 16 November 1999, describes it as “a language for addressing parts of an XML document, designed to be used by both XSLT and XPointer.” An HTML parser builds a tree from the response markup; XPath expressions select nodes from that tree. XPath does not fetch a page or execute its JavaScript by itself.

The examples below use Scrapy’s selector API and assume response is an HTML response. Scrapy selectors are a thin wrapper over Parsel, which uses lxml underneath. Parsel can also be used independently of Scrapy; lxml parses HTML and XML but is not included in Python’s standard library.

Core XPath expressions for scraping

Goal XPath Result and use
Select all H1 elements //h1 Element selectors for each matching heading.
Extract text nodes directly inside H1 //h1/text() Text nodes that are direct children of each H1; nested child text is not included by this step.
Get link destinations //a/@href The href attribute values of matching anchors.
Find links whose href contains a value //a[contains(@href, "image")]/@href Anchor hrefs containing the substring image.
Select a div by ID //div[@id="images"] Div elements whose ID exactly matches.
Get the title text //title/text() Use .get() for one string, or .getall() for all matches.
Get image source attributes //img/@src All matching src values when retrieved with .getall().
Find paragraphs below the current selector .//p Paragraph descendants of the current selected element.

Extract one value or a list in Scrapy

response.xpath() returns selector objects, not plain strings. Call .get() to serialize the first match, or .getall() to serialize all matches into a list. If there is no match, .get() returns None unless you pass a default value.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
#1 Best Overall
Sale
HTML and CSS: Design and Build Websites
  • HTML CSS Design and Build Web Sites
  • Comes with secure packaging
  • It can be a gift option
title = response.xpath("//title/text()").get()
image_urls = response.xpath("//img/@src").getall()
first_title_or_empty = response.xpath("//title/text()").get(default="")

Choose based on what the next stage needs: a scalar for a field expected once, a list for repeated values, and an explicit default when absence should become a particular value. Do not assume .get() validates uniqueness; if the page has multiple matches, it silently gives you the first.

Understand //, .//, and relative scope

In a nested selector, // begins a document-level search. The dot in .// anchors the search under the current node. This distinction matters when iterating through repeated containers.

for card in response.xpath("//article"):
    # Searches the document, not just this card:
    all_paragraphs = card.xpath("//p").getall()

    # Searches descendants of this card:
    card_paragraphs = card.xpath(".//p").getall()

    # Searches only direct child paragraphs:
    direct_paragraphs = card.xpath("p").getall()

Use .//p when a paragraph may appear at any depth inside the selected container. Use p when only immediate child paragraphs are wanted. A document-level query inside a loop can repeat the same site-wide results for every container, producing plausible-looking but incorrect output.

Position predicates: first per parent or first overall?

Predicates apply to the node context they are written in. Consequently, //li[1] selects each li that is the first matching li child under its relevant parent. A page with several lists can therefore produce several results. To select the first li in the document-wide result set, write (//li)[1].

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
first_item_in_each_list = response.xpath("//li[1]").getall()
first_item_in_document = response.xpath("(//li)[1]").get()

Parentheses change the order of operations: first evaluate the full set of li matches, then take its first node. Whenever a positional result looks unexpectedly numerous, check which parent context the predicate applies to.

Text nodes, nested markup, and element string values

text() selects text-node children; .//text() selects text nodes below the current element at any depth. This is useful when you need separate text fragments, but those fragments may be split by nested markup. When checking an element’s combined text, use its string value, written as ..

Rank #4
Sale
Web Design with HTML, CSS, JavaScript and jQuery Set
  • Brand: Wiley
  • Set of 2 Volumes
  • A handy two-book set that uniquely combines related technologies Highly visual format and accessible language makes these books highly effective learning tools Perfect for beginning web designers and front-end developers
# Individual descendant text nodes:
response.xpath("//a//text()").getall()

# Test combined descendant text, even if nested markup splits it:
response.xpath("//a[contains(., 'Next Page')]").getall()

A string function applied to a node set such as .//text() can convert only the first text node to a string. Thus contains(.//text(), 'Next Page') may fail when the words are split across nested elements, while contains(., 'Next Page') tests the element’s combined string value.

Match class tokens safely

HTML class attributes can contain several whitespace-separated tokens. An exact comparison such as //*[@class='product'] misses an element whose class is product featured. A raw substring check such as contains(@class, 'prod') can match an unintended token like product-card.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

For a token-safe match, normalize whitespace and pad both sides so the token is checked as a whole:

//*[contains(concat(' ', normalize-space(@class), ' '), ' product ')]

For straightforward class selection, Scrapy’s CSS selector can be more readable; CSS queries are translated into XPath internally. You can select by CSS first, then chain XPath when you need text nodes, attributes, or more complex structural predicates.

Namespaces, parser choice, and JavaScript-rendered pages

These are separate problems from XPath syntax. A correct XPath expression can still return nothing if it does not match the parsed tree or the response type.

  • Namespaces: An XML feed may put elements in a namespace, so a namespace-free expression such as //link may not match. Use namespace-aware queries with mappings, or deliberately remove namespaces with Scrapy’s remove_namespaces(). Removing namespaces changes the tree and has a processing cost.
  • Parser and response type: Scrapy documents response type selection and namespace behavior. Check that the response is parsed as the type you expect and inspect the parsed content rather than assuming the server returned ordinary HTML.
  • JavaScript content: XPath addresses the tree the parser receives. It does not make a client-rendered page populate missing elements; when needed content is absent from the parsed response, you need a workflow that supplies rendered page content before extraction.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Choose XPath or CSS for the job

Need Often a good choice Why
Simple class-based selection CSS Often easier to read for common class selectors.
Text nodes or attributes XPath Expressions such as //a/@href and //p/text() address these directly.
Structural relationships or predicates XPath Supports context, positional predicates, and other query logic.
Repeated nested containers Either, with explicit scope Choose based on readability and ensure the selector is relative to the current container when appropriate.

There is no basis here for a blanket speed claim for one selector style. Choose by readability, needed node type, parser behavior, scraper integration, and scope. Scrapy offers both response.css() and response.xpath(); Parsel is available without Scrapy if that better fits your workflow.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A practical extraction workflow

  1. Inspect the response content. Confirm that the response contains the element you intend to extract and that the parser is handling it as HTML or XML as appropriate.
  2. Anchor on a stable structure. Use a meaningful ID or container where available, rather than relying on a fragile page-wide positional match.
  3. Decide the scope. Use // for a document search, .// under a selected node, and a bare child name such as p for direct children.
  4. Select the right node type. Use an element expression for elements, /@attribute for an attribute, and /text() for direct text nodes.
  5. Choose cardinality deliberately. Use .get() for one serialized match or .getall() for a list; provide a default if a missing single value needs one.
  6. Test awkward cases. Check pages with absent fields, multiple containers, nested markup, and classes containing more than one token.

Troubleshooting common XPath misses

  • Nested loop returns the same paragraphs for every container: replace //p with .//p for descendants of the current selector, or p for direct children.
  • //li[1] returns several results: that expression selects the first matching child in each parent context. Use (//li)[1] for the first item in the overall document result.
  • Exact class match finds nothing: the element may have multiple class tokens. Use the token-safe normalized expression or select the class with CSS.
  • Class substring match finds the wrong element: substring matching does not respect token boundaries. Use the padded normalize-space(@class) form.
  • Text test fails around bold or span markup: use . to test combined descendant text; use .//text() only when individual text nodes are what you need.
  • .get() gives None: no node matched in the parsed response. Check the response content, selector scope, HTML/XML type, and namespace before changing the expression blindly.
  • Only one result appears though several were expected: use .getall() and confirm the XPath matches all intended nodes; .get() returns only the first result.
  • Namespace-free query misses XML elements: query with the namespace mapping or deliberately remove namespaces before querying.
  • Rendered content is absent: XPath cannot select nodes that are not in the parsed tree. Arrange for the page content to be rendered or otherwise present in the response before extracting it.

Or skip the browser setup

When your workflow needs a screenshot or PDF rather than HTML nodes to query, ScreenshotNeo provides a one-request capture API. It is not an XPath parser: use Scrapy and XPath for structured extraction, and use this endpoint when the output you need is a rendered image.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

See the ScreenshotNeo API documentation for request options. Before capture, it accepts cookie or consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing, with response headers indicating the page verdict and billing status. Its MCP server offers take_screenshot, get_page_info, and capture_pdf for AI agents. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. Sign up for the free plan.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Read next

Recommended PC Tool
Recommended PC Tool
Windows Errors? Fix Them Before They SpreadFree repair scan
Crashes, No Sound, or Screen Glitches?Free driver scan

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.