How do you scrape a website with Scrapy? Install Scrapy 2.19.0 in an isolated Python 3.10-or-newer environment, create a project and spider, extract fields with CSS or XPath selectors, yield structured items, and export them with feed exports. When browser content is missing from the downloaded response, investigate the page’s underlying data requests before adding a headless browser.
What Scrapy does
Scrapy is a Python framework for crawling websites and extracting structured data. A spider defines requests and parses responses, yielding items and additional requests. The scheduler queues requests, the downloader fetches them, and the engine coordinates the flow. Items are key-value records; downloader middleware changes request/response behavior such as headers, retries, authentication, redirects and proxies; spider middleware processes responses, requests and items around callbacks; item pipelines clean, validate, filter or persist items; extensions provide cross-cutting functions such as statistics and crawl-progress logging.
This guide follows the Scrapy 2.19.0 documentation, which is the documented release at the time of writing. Release notes and compatibility can change, so check the official installation page before upgrading. Scrapy 2.19.0 lists Python 3.10 or newer as its minimum.
Install Scrapy in a virtual environment
An isolated environment prevents Scrapy’s dependencies from conflicting with system packages.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
#1 Best Overall
- Install Python 3.10 or newer for your operating system.
- Create and activate an environment:
python -m venv .venv
# macOS/Linux
source .venv/bin/activate
# Windows PowerShell
.venvScriptsActivate.ps1
- Install Scrapy with pip:
python -m pip install --upgrade pip
python -m pip install Scrapy
The official guide also documents conda-forge. On Windows, pip dependencies can require Microsoft C++ Build Tools; conda-forge can avoid many of those setup problems. Optional extras add integrations such as HTTPX, S3, Google Cloud Storage, image pipelines or shell interfaces, but none is required for a basic crawl.
scrapy version
If installation fails, read the current platform-specific installation notes rather than pinning arbitrary dependency versions.
Create your first spider
Use quotes.toscrape.com as a contained learning site. For any real target, establish that your intended access and use are appropriate under the site’s terms, access rules and applicable law.
- Create a project:
scrapy startproject quotes_project
cd quotes_project
- Create a spider:
scrapy genspider quotes quotes.toscrape.com
Replace quotes_project/spiders/quotes.py with:
import scrapy
class QuotesSpider(scrapy.Spider):
name = "quotes"
allowed_domains = ["quotes.toscrape.com"]
start_urls = ["https://quotes.toscrape.com/"]
def parse(self, response):
for quote in response.css("div.quote"):
yield {
"text": quote.css("span.text::text").get(),
"author": quote.css("small.author::text").get(),
"tags": quote.css("div.tags a.tag::text").getall(),
}
next_page = response.css("li.next a::attr(href)").get()
if next_page:
yield response.follow(next_page, callback=self.parse)
The callback extracts every quote, then follows the site’s next-page link until no link remains. response.follow() resolves a relative URL safely against the current response.
Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minutePC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Use CSS selectors or XPath
Scrapy exposes Parsel selectors through response.css() and response.xpath(). Choose the expression that matches the response structure and that your team can maintain; neither style is universally superior.
CSS examples
response.css("h1::text").get()
response.css("article.product").getall()
response.css("a::attr(href)").getall()
XPath examples
response.xpath("//h1/text()").get()
response.xpath("//article[contains(@class, 'product')]")
response.xpath("//a/@href").getall()
get() returns the first match or None; getall() returns a list. Test selectors against the actual response before embedding them in a spider:
Rank #2
scrapy shell https://quotes.toscrape.com/
In the shell, inspect response.text, then try response.css(...).getall() or response.xpath(...).getall(). A selector that worked yesterday can fail after a site changes its HTML, so keep expressions tied to stable structure where possible.
Yield items, then export JSON, CSV or XML
Yield dictionaries for simple records, or define explicit item classes when a project benefits from declared fields. Feed exports are the direct route for serializing items; a custom pipeline is not needed merely to write a file.
The Tool Desk
Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →scrapy crawl quotes -O quotes.json
scrapy crawl quotes -O quotes.csv
scrapy crawl quotes -O quotes.xml
The -O option overwrites the output file. Use -o to append to an existing feed where supported by your workflow. Scrapy feed exports support formats including JSON, CSV and XML, plus storage backends configured in project settings.
When to use an item pipeline
Put item-level cleanup, validation, duplicate filtering and persistence in a pipeline rather than turning a spider callback into a database layer. Enable pipelines in settings.py:
ITEM_PIPELINES = {
"quotes_project.pipelines.ValidateQuotePipeline": 300,
}
Priorities run from lower numbers to higher numbers. A validation pipeline might look like this:
from itemadapter import ItemAdapter
class ValidateQuotePipeline:
def process_item(self, item, spider):
adapter = ItemAdapter(item)
text = adapter.get("text")
author = adapter.get("author")
if not text or not author:
raise ValueError("quote requires text and author")
adapter["text"] = text.strip()
adapter["author"] = author.strip()
return item
Use multiple pipelines when ordering matters—for example, normalize first, validate second, then write to a database. Keep request/response behavior in downloader middleware and response/item flow behavior in spider middleware.
Quick wins for a faster PC:
Clear out junk files and repair common Windows errorsFree Scan →Scan for outdated or missing drivers - takes under a minuteDriver Scan →Repair Windows errors before they cause bigger problemsFix Now →Follow links without losing crawl control
Pagination can be followed recursively as in the example, but unrestricted link crawling can expand rapidly. Restrict callbacks to links that represent the records you need, use allowed_domains, and monitor item counts and errors. For recurring crawls, configure download delay, per-domain concurrency and AutoThrottle. AutoThrottle adapts concurrency and delay to observed server load; appropriate values depend on the target and workload, not on a universal benchmark.
Review the target’s terms, access controls and applicable law before crawling. A robots.txt file is not, by itself, a legal authorization or a complete policy decision.
Why Scrapy cannot see content visible in your browser
A normal Scrapy request receives the server response; it does not automatically execute the page’s JavaScript like a browser. If a product list or article appears in a browser but is absent from response.text, diagnose the source in this order:
- Open browser developer tools and inspect the Network panel while the content loads.
- Identify the JSON, GraphQL or other request that returns the data.
- Reproduce that request in a Scrapy callback with the required URL, method, parameters, headers or body, while respecting access rules.
- Check whether the data is embedded in a script tag or loaded from an external resource.
- Only if the needed content is available exclusively in the rendered DOM, consider a headless browser integration.
Direct source extraction is usually simpler and cheaper to operate than rendering an entire browser. A headless browser is an escalation for pages whose required state cannot be obtained from an accessible underlying request, not the default fix for every missing element.
Do these 3 things before closing this tab:
1Repair Windows errors before they cause bigger problems2Fix the driver behind crashes, sound loss and screen glitches3Clear out junk files and repair common Windows errorsExample data request
def parse(self, response):
api_url = response.css("script[data-api-url]::attr(data-api-url)").get()
if api_url:
yield response.follow(api_url, callback=self.parse_api)
def parse_api(self, response):
data = response.json()
for record in data.get("results", []):
yield record
The attribute and JSON shape above are illustrative; inspect the target response and adapt them to its current structure.
Or skip the browser setup
If your goal is a clean image or PDF of a rendered page rather than structured records, ScreenshotNeo provides a single screenshot API call. It accepts cookie and consent banners as a visitor and removes more than 60 known consent platforms, newsletter popups and chat widgets before capture; each step can be disabled. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads and cache hits are not billed, and response headers identify the page verdict and billing status. Its MCP server exposes take_screenshot, get_page_info and capture_pdf to Claude, Cursor and other MCP clients.
Read the complete parameters in the ScreenshotNeo documentation. This cURL request saves a WebP image:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python:
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
Every plan includes the feature set: full-page and element capture, device presets, custom viewports, retina scale, PDF controls, custom CSS and JavaScript, waits, request blocking, headers, cookies, user agents, timezone and geolocation, transparent backgrounds, resizing, configurable caching, signed links, asynchronous webhooks, bulk capture of up to 100 URLs per call, a usage API and an OpenAPI specification. Pricing is Free for 1,000 shots per month with no card, Starter $5 for 3,000, Growth $15 for 15,000, Pro $39 for 60,000, Scale $99 for 250,000 and Business $249 for 1,000,000; yearly billing gives two months free.
Sign up free for 1,000 screenshots a month with no card.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Troubleshoot common failures
“No module named scrapy”
The virtual environment is not active or Scrapy was installed into another interpreter. Activate .venv, then run python -m pip install Scrapy and python -m scrapy version.
Empty fields
Inspect response.text in scrapy shell. The selector may target changed markup, a different page template or content loaded by JavaScript. Test a narrower selector and verify whether ::text or an attribute selector is appropriate.
Relative links fail
Use response.follow() or response.urljoin() instead of concatenating strings. Confirm that the callback receives the expected response URL.
403, 429 or repeated retries
Check that your request is appropriate, reduce concurrency, set a respectful download delay or enable AutoThrottle, and investigate required authentication or headers through documented access methods. Do not assume rotating proxies or bypassing controls is acceptable.
Best Value
Pipeline changes do not appear
Confirm the fully qualified class path in ITEM_PIPELINES, check its priority, and watch the crawl log for exceptions. A pipeline receives yielded items, not arbitrary response objects.
The crawl is too slow
Measure where time is spent: DNS and downloads, retries, parsing, rendering or persistence. Avoid browser rendering when a data request provides the needed records, and tune concurrency only within limits suitable for the target.
When to move beyond a tutorial spider
Use project settings for shared behavior and a spider’s custom_settings for per-spider overrides. Add explicit item definitions and pipelines when validation or persistence becomes important. Add extensions for statistics and progress reporting. For recurring jobs, plan storage, retries, deduplication, scheduling and deployment separately from extraction logic. Scrapy Cloud appears in the project’s deployment workflow, but service details and availability should be checked currently before choosing it.
Frequently Asked Questions
Which Python version does Scrapy 2.19.0 require?
The documented minimum is Python 3.10. Use a dedicated virtual environment and verify compatibility in the current installation notes.
Should I choose CSS or XPath selectors?
Both are supported through response.css() and response.xpath(). Choose the expression that best matches the current HTML structure and your team’s maintenance skills.
Do I need a custom pipeline to create a JSON file?
No. Feed exports handle ordinary JSON, CSV and XML output. Add a pipeline for cleanup, validation, filtering or custom persistence.
When should I use a headless browser?
First inspect network requests, embedded data and external resources. Use a headless browser when the required content is available only after rendering and cannot be obtained from an underlying source request.
Recommended Free Tools
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

