Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
Parsel is the extraction layer, not the downloader. Give it an HTML, XML, or JSON document, create a Selector, and use CSS, XPath, JMESPath, or regular expressions to turn matching nodes into Python values. A separate HTTP client (such as requests) or Scrapy must fetch pages; Parsel does not run JavaScript, schedule requests, or provide a complete crawler.
This guide builds a practical scraper from supplied markup, explains when to choose CSS, XPath, or JMESPath, covers common selector mistakes, and shows how the same selectors fit into Scrapy.
Install Parsel and verify your Python environment
Parsel is published as the parsel package. The current PyPI project page lists Parsel 1.12.1, released September 28, 2026, and Python >=3.10. Package requirements can change, so check the project page when you create or upgrade an environment.
python -m pip install parsel
python -c "import parsel; print(parsel.__version__)"
The project is licensed under BSD-3-Clause. Parsel’s release history shows why old tutorials can mislead: version 1.11.0 removed Python 3.9 and PyPy 3.10 support while adding Python 3.14 and PyPy 3.11 support. Use the interpreter and dependency versions installed in your own virtual environment rather than copying an outdated requirement.
#1 Best Overall
How do I use Parsel in Python to scrape a webpage?
First obtain the response body with an HTTP client, then pass the text to Selector. The selector API returns selector objects; call .get() for one string or .getall() for a list.
from parsel import Selector
html = """<html><body>
<h1>Example</h1>
<a href="/guide">Read the guide</a>
</body></html>"""
sel = Selector(text=html)
title = sel.css("h1::text").get()
link = sel.css("a::attr(href)").get()
all_links = sel.css("a::attr(href)").getall()
print(title) # Example
print(link) # /guide
print(all_links) # ['/guide']
According to the Parsel usage documentation, .get() always returns one result: the first match, or None when there is no match. Use .get(default="not-found") when a missing value should become a known fallback. .getall() consistently returns a list, including an empty list when nothing matches.
Fetch the page separately
import requests
from parsel import Selector
url = "https://example.com/"
response = requests.get(url, timeout=30)
response.raise_for_status()
sel = Selector(text=response.text)
heading = sel.css("h1::text").get(default="")
print(heading.strip())
This example downloads ordinary server-rendered HTML. Respect the site’s terms, robots policy, rate limits, authentication rules, and applicable law. If the useful content is inserted only after JavaScript runs, a plain HTTP request will not contain it; Parsel cannot render that JavaScript.
How do I select elements with CSS or XPath in Parsel?
CSS for straightforward element and class selection
CSS is concise when the relationship is visual and structural: article h2::text, .product a::attr(href), or ul.items li. Parsel adds scraping-oriented pseudo-elements:
::textselects text nodes.::attr(name)selects an attribute value.
These extensions are Parsel/Scrapy features, not portable standard CSS. The documentation warns that they may not work in libraries such as lxml or PyQuery.
names = sel.css(".product .name::text").getall()
urls = sel.css(".product a::attr(href)").getall()
XPath for traversal, XML, and precise text handling
XPath is useful for XML, document-relative navigation, conditions, and cases where CSS cannot express the relationship clearly.
dates = sel.css(".shout").xpath("./time/@datetime").getall()
# The leading dot keeps this XPath relative to each .shout node.
all_text = sel.xpath("normalize-space(string(//article[1]))").get()
Inside a nested selector, begin with . when the path should be relative. A leading slash refers to the document root, so /time can unexpectedly escape the node you just selected.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
JMESPath for JSON
Use JMESPath when the input is JSON rather than an HTML tree. Parsel can select JSON text embedded in a script and then apply a JMESPath expression:
json_values = sel.css("script::text").jmespath("a").getall()
For a JSON response body, construct a selector with JSON input according to the installed Parsel version, then use JMESPath expressions for fields and arrays. Keep HTML/XML selection and JSON selection conceptually separate: CSS and XPath describe nodes, while JMESPath describes JSON data.
Regular expressions after structural selection
Parsel also supports regular-expression extraction. Prefer selecting the relevant node first and applying a regex to its text; using regex as a replacement for parsing nested HTML makes escaped markup, whitespace, and repeated elements harder to handle.
Rank #3
How do I extract text, links, and attributes with Parsel?
One value versus many values
first_price = sel.css(".price::text").get()
prices = sel.css(".price::text").getall()
missing = sel.css(".does-not-exist::text").get(default="unknown")
Normalize values at the boundary of your program:
clean_prices = [value.strip() for value in prices]
clean_title = (first_price or "").strip()
Attributes and links
hrefs = sel.css("a::attr(href)").getall()
images = sel.xpath("//img/@src").getall()
labels = sel.css("button::attr(aria-label)").getall()
These are usually relative URLs. Resolve them against the page URL with Python’s standard library before storing or requesting them:
Quick wins for a faster PC:
Scan for outdated or missing drivers - takes under a minuteDriver Scan →Repair Windows errors before they cause bigger problemsFix Now →from urllib.parse import urljoin
absolute = [urljoin(url, href) for href in hrefs]
Nested text is a frequent surprise
::text and XPath text() select direct text nodes. They can omit words inside child elements:
html = "<p>Hello <strong>world</strong>!</p>"
node = Selector(text=html)
print(node.css("p::text").getall()) # ['Hello ', '!']
print(node.xpath("string(//p)").get()) # Hello world!
Use string(.) for all descendant text and normalize-space(.) when trimming and collapsing whitespace is appropriate.
Classes, scripts, and malformed documents
- Prefer
.css('.someclass'). Exact@class='someclass'misses elements with multiple classes, while an unboundedcontains(@class, 'someclass')can match unrelated names. - Script and style contents are parsed as plain text. Tag-like strings inside them do not become child nodes.
- On a malformed document with multiple roots, CSS starts from the first root. If every root matters, select all roots with XPath first, then apply CSS to each nested selector.
Can I use Parsel without Scrapy?
Yes. Standalone Parsel is appropriate when another component already supplied the HTML, XML, or JSON body. Scrapy's selectors are a thin wrapper around Parsel designed to integrate with Response objects; in a spider callback, response.css() and response.xpath() are convenient shortcuts.
| Need | Use | Why |
|---|---|---|
| Already have markup or JSON | Standalone Parsel | Small dependency focused on selection and extraction |
| Requests, scheduling, retries, concurrency and item pipelines | Scrapy with Parsel selectors | Scrapy supplies the crawler and response workflow |
| JavaScript-rendered content | A browser or rendering service, then Parsel if desired | Parsel itself does not execute JavaScript |
The relationship and response shortcuts are documented by the Scrapy selector documentation. This is an architectural distinction, not a performance claim.
Free tools Windows power users keep installed
One-click scans. No signup required.
A maintainable extraction pattern
- Fetch once and check the HTTP status and content type.
- Create one
Selectorfor the response body. - Use stable landmarks such as semantic elements, IDs, or dedicated classes rather than styling-only selectors.
- Keep cardinality explicit: use
getall()for collections and validate required fields returned byget(). - Normalize whitespace and resolve URLs before writing records.
- Log missing selectors and sample pages so a template change is visible instead of silently producing empty data.
from dataclasses import dataclass
from urllib.parse import urljoin
from parsel import Selector
@dataclass
class Article:
title: str
url: str
tags: list[str]
def parse_article(body: str, base_url: str) -> Article:
sel = Selector(text=body)
title = sel.css("h1::text").get(default="").strip()
href = sel.css("article a.permalink::attr(href)").get(default="")
tags = [t.strip() for t in sel.css(".tags a::text").getall() if t.strip()]
if not title or not href:
raise ValueError("required article fields were not found")
return Article(title, urljoin(base_url, href), tags)
Or skip the browser setup
If your goal is a clean image or PDF of a page rather than structured fields, ScreenshotNeo can handle the capture in one request. It accepts consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be disabled. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and the response identifies the page verdict and billing status in X-Page-Verdict and X-Billed headers. Its MCP server exposes take_screenshot, get_page_info, and capture_pdf to Claude, Cursor, and other MCP clients.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
See the ScreenshotNeo API documentation for options such as PNG/JPEG/WebP output, full-page and element capture, device presets, custom CSS or JavaScript, waits, cookies and headers, blocking rules, PDF settings, caching, signed links, asynchronous webhooks, bulk capture, and the usage API.
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
The Free plan includes 1,000 shots per month with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Troubleshooting Parsel scrapers
The selector returns None or an empty list
Print a short response sample, confirm the HTTP status and content type, and inspect the actual server response. Check spelling, nesting, class names, and whether the content is JavaScript-rendered. Use get(default=...) deliberately, but do not hide a required-field failure.
Recommended Free Tools
Only the first item is extracted
Replace .get() with .getall() for repeated nodes, then iterate over the returned list.
Nested text is missing
Replace direct-text selection with XPath string(.) or normalize-space(.).
Best Value
Nested XPath selects the wrong place
Add the relative dot: .xpath('.//a/@href'). A path beginning with / starts at the document root.
JSON data is empty
Verify that the selected script actually contains JSON, that it is valid for the expression you use, and that you are applying JMESPath to JSON rather than trying to use an HTML CSS path.
The Tool Desk
Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →The page is blocked or incomplete
Parsel cannot solve bot challenges, consent overlays, network failures, or browser-only rendering. Handle those concerns in the HTTP client, crawler, browser, or capture service before passing the resulting body to Parsel.
FAQ
Does Parsel crawl websites?
No. It extracts from a document supplied by another component; Scrapy adds crawling and request integration.
Which selector should I learn first?
Start with CSS for simple HTML, then add XPath for relative traversal, XML, and descendant-text cases; use JMESPath for JSON.
Is Parsel's ::text valid CSS everywhere?
No. It is a Parsel/Scrapy extension and is not portable to every CSS selector library.
PC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchFrequently Asked Questions
Does Parsel execute JavaScript?
No. Fetch or render the page with another tool, then give Parsel the resulting HTML.
What does get() return when there is no match?
It returns None unless you provide a default value.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

