Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

To crawl a page with Scrapy, install it in a Python virtual environment, create a spider that yields requests and extracted items, then run the spider with the command-line tool and export its results. This walkthrough uses Scrapy’s official quotes.toscrape.com tutorial site and the current documented asynchronous spider interface. The installation documentation presented Scrapy 2.19.0 on September 30, 2026, and requires Python 3.10 or newer.

What a Scrapy crawl does

Scrapy is a Python framework for crawling websites and extracting structured data. A spider defines the pages to request and the logic for parsing their responses. As the crawler downloads a page, the spider can yield data items, additional requests, or both. Scrapy schedules those requests, passes downloaded responses to callbacks, and can export yielded items to a file.

This is different from manually opening a page and copying its content: the spider describes a repeatable workflow. The example below extracts quotes and authors from a site designed for practice. Selectors and page structure vary across websites, so treat the example as a pattern—not as a selector set that will work unchanged on arbitrary pages.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Install Scrapy and create a project

Scrapy’s current installation guide requires Python 3.10 or newer. It recommends installing into a dedicated virtual environment, which helps keep project dependencies separate from system Python packages. The commands below use the standard venv module and pip.

  1. Check that Python is available. On systems where the command is python3, use that command instead of python in the following steps.

    python --version
  2. Create and activate an environment from the directory where you want the project. Activation differs by shell:

    python -m venv .venv
    # macOS/Linux, bash or zsh:
    source .venv/bin/activate
    # Windows PowerShell:
    .venvScriptsActivate.ps1
    # Windows Command Prompt:
    .venvScriptsactivate.bat
  3. Install Scrapy and generate its project structure:

    python -m pip install Scrapy
    scrapy startproject tutorial
    cd tutorial

    Scrapy can also be installed through conda-forge. Its dependencies include packages such as lxml, parsel, w3lib, Twisted, cryptography, and pyOpenSSL; some dependencies can require platform-specific setup. If installation fails, use the official installation guide for the current platform-specific advice.

    Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  4. Identify your crawler in the project settings. Open tutorial/settings.py and set a descriptive user agent, for example:

    USER_AGENT = "practice-crawler (contact: [email protected])"

    Use contact details you control. An identifying user agent helps site owners identify and contact the crawler operator.

The generated project contains settings, item and pipeline modules, and a spiders directory. For this walkthrough, the spider itself is the main file you need to edit.

Write a spider that extracts data and follows pages

Create tutorial/spiders/quotes.py and add the following code. It starts at the practice site, extracts each quote and author, then follows the next-page link when one exists.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
import scrapy


class QuotesSpider(scrapy.Spider):
    name = "quotes"

    async def start(self):
        yield scrapy.Request("https://quotes.toscrape.com/")

    def parse(self, response):
        for quote in response.css("div.quote"):
            yield {
                "text": quote.css("span.text::text").get(),
                "author": quote.css("small.author::text").get(),
            }

        next_page = response.css("li.next a::attr(href)").get()
        if next_page:
            yield response.follow(next_page, callback=self.parse)

What each part means

  • QuotesSpider subclasses scrapy.Spider. Its name is the unique identifier used to select it from the project.

  • async def start(self) is the asynchronous start interface used in the current official tutorial. It yields a Scrapy request; Scrapy downloads the response and sends it to parse. Older tutorials may show a different interface, so check that examples match the Scrapy version you are using.

  • response.css(...) selects matching elements in the downloaded response. .get() returns the first matching value, or None if there is no match. The yielded dictionary is one result item.

  • The next-page selector retrieves the link’s href. response.follow resolves relative links against the current response URL and schedules a request whose response is handled by parse again.

    Free tools Windows power users keep installed

    One-click scans. No signup required.

    Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Choose CSS or XPath based on the markup

Scrapy supports both response.css() and response.xpath(); its documentation explains that CSS selectors are converted to XPath internally. CSS is often readable for matching familiar classes and elements, as in div.quote. XPath can be useful when the selection depends on document structure or text content—for example, finding an anchor by the text it displays. Neither is universally better: choose the expression that makes the intended match clear and is easiest to maintain as the page changes.

Do not assume a selector is right because it looks plausible. Inspect the actual response and refine the selector against the HTML your spider receives.

Inspect a response before relying on selectors

The Scrapy shell lets you interact with a response and test CSS or XPath expressions before embedding them in a spider. From the project directory, open a shell for the practice page:

scrapy shell https://quotes.toscrape.com/

At the shell prompt, try expressions such as:

response.css("div.quote span.text::text").getall()
response.css("li.next a::attr(href)").get()
response.xpath("//a[contains(., 'Next')]/@href").get()

Inspect the returned values and compare them with the page’s markup. An empty list or None usually means the selector did not match the response content; it does not by itself prove that Scrapy failed to download the page. The response may differ from the browser view, or the site’s markup may have changed.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Run the spider and save its results

Run the spider from the project root—the directory containing scrapy.cfg—and ask Scrapy to export the yielded dictionaries as JSON:

scrapy crawl quotes -O quotes.json

The spider name is quotes, matching the class’s name attribute. The capital -O option writes the feed and overwrites an existing output file. If you want to append to an existing feed instead, Scrapy also provides lowercase -o; consult the current command-line documentation for format-specific behavior. Feed export supports choosing a format by the output filename or an explicit feed URI.

After the crawl completes, inspect quotes.json. Each yielded dictionary becomes an exported item. If the file is empty, check that the spider ran, that the selectors matched, and that the response reached the expected callback.

Pass values into a spider

Spider arguments are useful when the same crawl needs a configurable starting point or other input. Scrapy passes command-line arguments supplied with -a to the spider constructor. For example, add a configurable URL:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
class QuotesSpider(scrapy.Spider):
    name = "quotes"

    def __init__(self, start_url="https://quotes.toscrape.com/", **kwargs):
        super().__init__(**kwargs)
        self.start_url = start_url

    async def start(self):
        yield scrapy.Request(self.start_url)

    def parse(self, response):
        for quote in response.css("div.quote"):
            yield {
                "text": quote.css("span.text::text").get(),
                "author": quote.css("small.author::text").get(),
            }

        next_page = response.css("li.next a::attr(href)").get()
        if next_page:
            yield response.follow(next_page, callback=self.parse)

Run it with a different starting URL by passing the argument after the spider name:

scrapy crawl quotes -a start_url=https://quotes.toscrape.com/ -O quotes.json

The example URL is still the tutorial site. If you point a spider at a different website, revise its extraction and link-following logic to match that site’s pages and crawling rules.

When to add an item pipeline

For a first crawl, exporting yielded dictionaries directly is usually the simplest route. A pipeline is an optional next step when items need additional processing before they are stored or exported. Scrapy documents pipelines as a place to clean, validate, deduplicate, or store items.

A pipeline is a class in a Python module with processing methods such as process_item. To enable one, add its dotted Python class path to ITEM_PIPELINES in tutorial/settings.py. Each configured pipeline has a numeric priority; lower numbers run before higher ones. Keep processing out of the pipeline until there is a concrete need—for example, validating required fields or removing duplicate records—so a beginner crawl remains easy to inspect.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Common problems and fixes

  • scrapy is not found. The virtual environment may not be active, or Scrapy may have been installed into a different Python environment. Activate .venv and run python -m pip show Scrapy; if it is absent, install it with python -m pip install Scrapy.

  • Python is too old. The current installation guidance requires Python 3.10 or newer. Check python --version, then create the environment with a compatible interpreter.

  • Dependency installation fails. Scrapy depends on several libraries, and setup can vary by operating system and Python environment. Read the official installation guide for platform-specific prerequisites rather than guessing at system-library changes.

  • The spider is not listed or cannot be selected. Confirm that the file is under the project’s spiders directory, that it imports successfully, and that its name is set. Run commands from the project root.

    Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • The crawl returns no items. Test the selectors in scrapy shell, verify that the response contains the expected markup, and confirm the callback is running. The page may have changed, or its response content may differ from what you see in a browser.

  • Only the first page is crawled. Check that the next-page selector matches a real link, that its href is present, and that the callback yields the follow-up request. response.follow handles relative links when given the response and link value.

  • The export file does not contain the latest run. Use -O to overwrite an existing feed. Lowercase -o has append-oriented behavior; choose intentionally to avoid confusing old and new records.

Reliability, performance, and crawl boundaries

Scrapy handles request scheduling and response callbacks, but successful results still depend on the response and your parsing assumptions. Selectors can break when a site changes its HTML, and missing fields should be anticipated when building a crawl for real data. Inspect a sample response and exported items before treating the output as complete.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Follow the target website’s rules and applicable requirements before crawling. The official tutorial demonstrates the software workflow; it does not grant permission to crawl every site. What is allowed can depend on the site, its terms, the data, the jurisdiction, and the intended use. Identify your crawler with a meaningful user agent and crawl only where your access and use are appropriate.

Or skip the browser setup

Scrapy is the right fit when you need a Python crawl with custom parsing and control over how links are followed. If your goal is simply to obtain a screenshot or PDF of a URL, ScreenshotNeo offers a one-request alternative. It is a screenshot API and MCP server for developers.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://quotes.toscrape.com/ -o shot.webp

See the ScreenshotNeo API documentation for request options. Its clean-shot flow accepts cookie or consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be turned off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers report the page verdict and billing status. An MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients. The free plan includes 1,000 shots per month with no card; paid plans start at $5 for 3,000 shots. It captures rendered images or PDFs rather than replacing Scrapy’s custom structured-data crawling.

Sign up for ScreenshotNeo’s free plan: 1,000 screenshots a month, no card required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Frequently Asked Questions

Can Scrapy follow links on other sites?

Technically, a spider can be configured to request pages on different sites, but what you may crawl depends on the particular site, its terms, the data, jurisdiction, and intended use. The practice-site example does not establish permission for other targets.

Do I need an item pipeline to export data?

No. A spider can yield dictionaries and Scrapy can export them directly. Pipelines are optional processing stages for needs such as validation, cleaning, deduplication, or storage.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.