Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

CSS selectors let a scraper locate matching elements in the HTML tree it has parsed. You then use your scraping library to read the selected element’s text or attributes—for example, an article heading or a link’s href. The selector finds the node; extraction code gets the value.

What a CSS selector does in a scraper

A CSS selector is a pattern for matching elements in a document tree. It can identify elements by tag name, ID, class, attributes, and relationships to other elements. The W3C Selectors Level 4 specification describes simple, compound, and complex selectors, as well as selector lists: W3C Selectors Level 4.

Suppose a page has several product cards, each with a heading and a link. The selector article.product identifies article elements that also have the product class. Your code can then query each matching card for its heading and link. This is useful because it scopes the smaller queries to one card at a time instead of accidentally combining a heading from one product with a link from another.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A selector does not fetch a page, run extraction code, or guarantee that the page contains the data you expect. It queries the parsed tree supplied to the selector engine. The scraper is responsible for obtaining the response and deciding what to do with the matched nodes and their values.

Build a selector from the page structure

Start with the element that contains the data you want, then narrow the match using attributes or relationships. These building blocks are supported by the CSS selector model; the exact selector features available in a scraper also depend on its selector engine.

Pattern What it matches Example use
article Elements with that tag name Find article elements.
#main The element with the main ID Target a known page section.
.product Elements with the product class Find elements marked as product content.
article.featured An article element that also has the featured class Combine conditions on the same element; there is no space between them.
article h2 An h2 anywhere inside an article descendant Find a heading nested somewhere within an article.
article > h2 An h2 that is a direct child of an article Use when the nesting relationship matters.
a[href^="https"] An anchor whose href begins with https Filter links by an attribute value prefix.

Commas can form a selector list when you want to match any of several patterns. Before making a selector more complicated, inspect the HTML and confirm which element actually holds the data. A selector can match the right kind of node but still be aimed at the wrong part of the page.

Use CSS selectors in Scrapy

Scrapy exposes response.css() as a shortcut for querying a response. Its selector stack uses Parsel with lxml underneath, so test selectors against the Scrapy selector implementation you will actually run. The current Scrapy selector documentation, accessed September 29, 2026, showed version 2.17.0; APIs and documentation versions can change. See Scrapy: Selectors.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Inside a spider callback, a basic pattern for extracting a name and link from each product card is:

for card in response.css("article.product"):
    name = card.css("h2::text").get()
    link = card.css("a::attr(href)").get()

The outer query produces each matching card. The nested queries are scoped to that card. In Scrapy, ::text selects text content and ::attr(href) selects an attribute value; these are Scrapy selector extensions used with CSS queries. .get() returns one result or no result, while .getall() returns all results as a list. If the page may omit a heading or link, handle a missing result rather than assuming every card has one.

For example, calling card.css("a::attr(href)").get() returns the value of the selected anchor’s href attribute. It does not return the anchor’s visible text. To collect multiple matching values, use .getall() on the relevant query. The right choice depends on whether the page structure is expected to contain one match or several.

Use CSS selectors in Beautiful Soup

Beautiful Soup provides select() to return matching elements and select_one() to return the first match. Its current CSS selector support is implemented by Soup Sieve. The documentation accessed September 29, 2026, showed Beautiful Soup 4.14.3; check the installed versions in your own environment because selector support can vary with the library and parser. See Beautiful Soup documentation.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

This standalone example parses a short HTML string, selects each product card, and extracts the heading text and link attribute:

from bs4 import BeautifulSoup

html = """
<article class="product">
  <h2>Notebook</h2>
  <a href="https://example.com/notebook">View product</a>
</article>
"""
soup = BeautifulSoup(html, "html.parser")

for card in soup.select("article.product"):
    heading = card.select_one("h2")
    name = heading.get_text(strip=True) if heading else None

    link = card.select_one("a")
    href = link.get("href") if link else None

    print(name, href)

Here, select() returns all matching cards, and select_one() looks for the first matching heading or link inside each card. get_text(strip=True) reads the heading text, while get("href") reads the link attribute. The if checks keep the example from trying to extract a value from a missing element.

Scrapy and Beautiful Soup express the same basic selection idea, but their extraction APIs differ: Scrapy’s examples use ::text, ::attr(...), and .get(); Beautiful Soup returns tag objects whose text and attributes you read with methods such as get_text() and get(). Do not copy an extraction expression from one library into the other and assume it works unchanged.

Test with the same parser and selector engine

A selector that works in a browser is not guaranteed to work identically in a scraping library. The parsed input, parser, installed library versions, and selector engine all matter. In particular, current Beautiful Soup CSS selection is implemented by Soup Sieve, while Scrapy uses Parsel with lxml underneath. Check the installed packages and test the exact query against the same parsing path used by your scraper. The libraries’ documentation describes their respective implementations: Scrapy selectors and Beautiful Soup.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  1. Inspect the HTML your scraper actually parsed. Do not rely only on what you see in a browser; the response content may not contain the same elements as a fully rendered page.
  2. Find the target element and check its tag, attributes, and nesting. Confirm whether the desired value is text inside an element or an attribute such as href.
  3. Try a small selector first. Test the tag, ID, class, or relationship that should identify the target before combining several conditions.
  4. Run the selector using your scraper’s library and parser. If it returns no results, confirm that the selector features you used are supported by the installed implementation.
  5. Check the extraction separately. A selector may match an element even when the attribute you request is absent or the text is empty.

CSS selection only queries the tree provided to the selector engine. It does not establish that a fetched response contains every element a visitor sees after a browser renders the page. If the data is absent from the parsed response, changing a selector cannot make that missing node appear.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

When CSS is enough—and when to consider XPath

For straightforward element, class, ID, relationship, and attribute selection, CSS is readable and supported by both Scrapy and Beautiful Soup. In Scrapy, XPath is also available through response.xpath(). XPath may be a better fit when a query is more naturally expressed as a path or needs XPath-specific capabilities. Choose based on readability for the query, feature support in the engine you use, and how clearly the result can be extracted. Scrapy documents both shortcuts in its selector guide.

Beautiful Soup’s documentation notes that if you need CSS selectors only, parsing with lxml directly may be faster. Treat that as the library documentation’s guidance, not a guarantee for every page, selector, or workload. If speed matters, measure the actual task and input you plan to process rather than assuming a general performance advantage.

Troubleshoot selectors that return no results

  • The selector is aimed at the wrong structure. Inspect the parsed HTML and verify the element’s tag, class, ID, attributes, and nesting. A descendant query and a direct-child query do not match the same relationships.
  • The response does not contain the element. Check the content your scraper parsed. A browser view is not proof that the same node exists in the response tree queried by the selector.
  • The selector feature is unsupported in the installed engine. Confirm the package versions and selector implementation, then test the query with that same runtime combination.
  • The element matches, but the extracted value is missing. Check whether you requested text or an attribute, and confirm the target element actually contains that text or attribute. In Scrapy, use the appropriate ::text or ::attr(name) query; in Beautiful Soup, read text or the attribute from the selected tag.
  • A query returns more or fewer matches than expected. Check whether you used a descendant relationship or direct-child relationship and whether multiple elements share the selected class or tag. Scope nested queries to a specific parent when you need to keep values associated with the correct record.

Or skip the browser setup

For a screenshot of a rendered page rather than structured text or attributes, ScreenshotNeo offers a one-request screenshot API and an MCP server for AI agents. It is not a replacement for CSS selectors when your goal is to extract fields from HTML; it serves the different task of capturing a page image or PDF.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

For example, this cURL request saves a WebP screenshot:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

See the ScreenshotNeo API documentation for request options. ScreenshotNeo accepts cookie or consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers report the page verdict and billing status. Its MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 screenshots.

Sign up for 1,000 free screenshots a month—no card required.

Further reading

For a broader treatment of scraping, O’Reilly lists Web Scraping with Python, 3rd Edition by Ryan Mitchell, published in February 2024 and 352 pages long. It covers HTML, CSS, JavaScript, and scraping mechanics beyond selectors: O’Reilly book listing.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.