Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Pass a run-specific value with scrapy crawl myspider -a name=value. Scrapy’s default spider initializer exposes that value as an attribute, so your spider can read self.name. When you start a crawl in Python, pass the same values as keyword arguments to CrawlerProcess.crawl() or CrawlerRunner.crawl(). All spider arguments arrive as strings; parse and validate lists, numbers, booleans, and JSON yourself.

The shortest working examples

From a terminal, provide one -a option for each argument:

scrapy crawl myspider -a category=electronics -a region=west

With a Python launcher, pass keyword arguments when scheduling the spider:

from scrapy.crawler import CrawlerProcess

process = CrawlerProcess()
process.crawl(MySpider, category='electronics', region='west')
process.start()

In either case, the spider receives category and region as attributes. The command-line form is best for an operator starting one crawl; the Python APIs are for applications, scheduled jobs, and pipelines that already control execution.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Design the spider’s parameter contract first

Decide which inputs vary from run to run, what format each value uses, and what should happen when an input is missing. A useful contract might look like this:

Argument Example Type at the boundary Spider policy
category electronics String Optional; use a general listing when omitted
region west String Optional; validate against supported regions
start_urls ["https://example.test/a", "https://example.test/b"] String containing JSON Required for a targeted run; decode and validate as a list
limit 100 String containing an integer Convert with explicit range checks

Keeping this contract explicit prevents a common mistake: assuming that a value that looks like a list, number, or Boolean has already been converted. Scrapy passes spider arguments as strings, including values supplied through Python keyword arguments.

Pass arguments from the command line

1. Start with a named spider

Run the command from the Scrapy project directory and use the spider’s name value:

scrapy crawl quotes -a tag=python

For multiple values, repeat -a; do not combine unrelated values into one option:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
scrapy crawl quotes -a tag=python -a region=west -a limit=25

2. Read the attribute in the spider

The default initializer copies supplied arguments onto the spider. An optional argument can therefore be read with getattr and a default:

import scrapy

class QuotesSpider(scrapy.Spider):
    name = 'quotes'

    async def start(self):
        tag = getattr(self, 'tag', None)
        url = 'https://quotes.toscrape.com/'
        if tag is not None:
            url += f'tag/{tag}'
        yield scrapy.Request(url, self.parse)

    def parse(self, response):
        for quote in response.css('div.quote'):
            yield {
                'text': quote.css('span.text::text').get(),
                'author': quote.css('small.author::text').get(),
            }

With this spider, scrapy crawl quotes -a tag=python requests the tag page, while omitting -a tag=... requests the site root. The current documentation also demonstrates an async def start() method; use the start-method style supported by the Scrapy version installed in your project.

3. Accept values in __init__ when you need custom setup

For simple access, no custom initializer is needed. Define one when you want to normalize or validate arguments before requests are generated, and call the base initializer so Scrapy can perform its normal setup:

import scrapy

class CatalogSpider(scrapy.Spider):
    name = 'catalog'

    def __init__(self, category=None, region=None, *args, **kwargs):
        super().__init__(*args, **kwargs)
        self.category = category
        self.region = region

    def start_requests(self):
        if not self.category:
            raise ValueError('category is required')
        url = f'https://example.test/catalog/{self.category}'
        yield scrapy.Request(url, callback=self.parse)

    def parse(self, response):
        yield {'url': response.url, 'region': self.region}

Use a default such as None for optional values. For required values, fail early with a clear message rather than issuing a broad or malformed request.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

4. Quote shell values deliberately

Shell syntax is applied before Scrapy receives the argument. Quote values containing spaces, ampersands, question marks, or shell metacharacters:

scrapy crawl catalog -a query='noise cancelling headphones' -a region='north america'

For a URL, quote the complete value so the shell does not interpret its query string:

scrapy crawl catalog -a start_url='https://example.test/search?q=phone&sort=price'

If a value contains credentials or other sensitive material, avoid putting it directly in shell history. Prefer environment variables or a protected secret mechanism, then read the variable in your launcher and pass the resulting value to the crawl.

Start a crawl from Python

CrawlerProcess: the self-contained launcher

CrawlerProcess configures and starts Scrapy when your script owns the reactor. Pass the spider class (or a spider name) followed by keyword arguments:

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
from scrapy.crawler import CrawlerProcess
from myproject.spiders.catalog import CatalogSpider

process = CrawlerProcess()
process.crawl(
    CatalogSpider,
    category='electronics',
    region='west',
    limit='25',
)
process.start()

The keyword values become spider arguments. They are still boundary values, so the spider should convert limit to an integer and validate it before using it in pagination or filtering.

CrawlerRunner: when another application owns the reactor

Use CrawlerRunner when your surrounding application already manages the reactor. Its crawl method accepts the spider class or name plus positional and keyword initialization arguments:

from twisted.internet import reactor, defer
from scrapy.crawler import CrawlerRunner
from myproject.spiders.catalog import CatalogSpider

runner = CrawlerRunner()

def run():
    return runner.crawl(
        CatalogSpider,
        category='electronics',
        region='west',
    )

d = run()
d.addBoth(lambda _: reactor.stop())
reactor.run()

Do not call reactor.run() a second time inside an application that already has a running reactor. That is the practical distinction between the process helper and the runner: choose the helper that matches who owns the event loop.

Async process and runner APIs

Current Scrapy API documentation also lists AsyncCrawlerProcess and AsyncCrawlerRunner for coroutine-based control flow. Their reactor and event-loop requirements depend on how the application is configured, so check the current Scrapy Core API documentation before embedding one in an existing asyncio service. The argument shape remains the same: pass run-specific values as keyword arguments to crawl.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Parse and validate structured values

Scrapy does not turn a command-line string into a Python collection automatically. If you pass start_urls=https://a.example,https://b.example and iterate over it as though it were a list, Python iterates over individual characters. Choose an unambiguous representation and decode it explicitly.

JSON for lists and nested data

import json
import scrapy

class MultiStartSpider(scrapy.Spider):
    name = 'multi_start'

    def __init__(self, start_urls_json=None, *args, **kwargs):
        super().__init__(*args, **kwargs)
        if not start_urls_json:
            raise ValueError('start_urls_json is required')
        try:
            value = json.loads(start_urls_json)
        except json.JSONDecodeError as exc:
            raise ValueError('start_urls_json must be valid JSON') from exc
        if not isinstance(value, list) or not all(isinstance(url, str) for url in value):
            raise ValueError('start_urls_json must be a JSON list of strings')
        self.start_urls = value

    def parse(self, response):
        yield {'url': response.url}

Invoke it with a shell-quoted JSON value:

scrapy crawl multi_start -a start_urls_json='["https://example.test/a", "https://example.test/b"]'

The official spider documentation mentions json.loads() and ast.literal_eval() as possible parsing approaches. JSON is usually easier to specify consistently across shells and languages; whichever format you choose, validate its shape before requests are scheduled.

Numbers and booleans

def parse_limit(raw):
    try:
        limit = int(raw)
    except (TypeError, ValueError) as exc:
        raise ValueError('limit must be an integer') from exc
    if limit < 1 or limit > 1000:
        raise ValueError('limit must be between 1 and 1000')
    return limit

def parse_bool(raw):
    normalized = str(raw).strip().lower()
    if normalized in {'1', 'true', 'yes'}:
        return True
    if normalized in {'0', 'false', 'no'}:
        return False
    raise ValueError('follow_next must be true or false')

Do not rely on Python’s bool('false'); every non-empty string is truthy. Define the accepted spellings and reject everything else.

URLs and lists with commas

A comma-separated convention is only safe when commas cannot occur in an item. URLs can contain commas, so JSON is safer for arbitrary URL lists. If you do use a delimiter, document escaping rules and test values containing the delimiter before deploying the spider.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Arguments versus settings

Scrapy’s FAQ says there is no rigid rule. Use spider arguments for inputs that change often between runs or apply only to one crawl, such as a target URL, customer partition, category, or temporary limit. Use settings for behavior that changes infrequently across runs, such as downloader middleware, concurrency policy, or a project-wide feed configuration. See the Scrapy FAQ for the documented distinction.

Putting stable configuration into -a makes every deployment command longer and easier to misconfigure. Putting a run-specific target into a settings file hides the input that actually changed. Keep the boundary visible: arguments describe this crawl; settings describe the project.

Choose the invocation that matches your environment

Situation Use How arguments are supplied Important constraint
Developer or operator starts one crawl CLI scrapy crawl name -a key=value Quote shell-sensitive values
Standalone Python script CrawlerProcess process.crawl(Spider, key=value) Let the process helper start the reactor
Application already owns Twisted’s reactor CrawlerRunner runner.crawl(Spider, key=value) Do not start a second reactor
Coroutine-based application AsyncCrawlerProcess or AsyncCrawlerRunner Keyword arguments to the async crawl API Match the configured reactor and event loop

The spider-arguments guide is available at Scrapy’s spider documentation. The documentation index currently identifies Scrapy 2.19.0; the linked spider page is version 2.12, while the API and FAQ links above provide current-version context. Check the version installed in your environment when an API detail differs.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Troubleshoot the common failures

“Unknown spider” or the command ignores your arguments

  • Confirm that the first positional value after scrapy crawl matches the spider’s name, not its Python class name.
  • Run scrapy list from the project directory to verify that Scrapy can discover the spider.
  • Put every custom value after its own -a option. A misspelled option or a value placed before the spider name is not a spider argument.

The attribute is missing

  • Use getattr(self, 'key', default) for optional values.
  • If you override __init__, call super().__init__(*args, **kwargs) and accept the keyword arguments you intend to use.
  • Check spelling and capitalization; region and Region are different attributes.

A list behaves like a string

That is expected until you parse it. Decode JSON (or another documented format), verify the resulting type, and only then iterate. Log the parsed type during development rather than assuming the input was converted.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A number or Boolean has the wrong behavior

Convert numeric strings with int() or float() and reject conversion errors. For Booleans, accept a small explicit set such as true/false and reject ambiguous spellings. Never use truthiness alone to interpret the string false.

The Python launcher fails with a reactor error

If your script owns execution, use CrawlerProcess. If a framework or service already runs the reactor, use CrawlerRunner and attach completion handling instead of calling reactor.run() again. For asyncio integrations, verify the reactor and event-loop combination required by the current API documentation.

The request URL is malformed

Print or log the final URL after interpolation. Shell quoting problems, unescaped query characters, and unvalidated parameters can all alter the value before Scrapy builds a request. Pass a complete URL as one quoted argument and validate its scheme and host in the spider.

Operational practices that keep parameterized crawls reliable

  • Document the contract: list every argument, whether it is required, its accepted format, and an example invocation beside the spider.
  • Validate before scheduling requests: reject bad values in initialization or the first start method so a typo does not create a large unintended crawl.
  • Use deterministic defaults: an omitted optional argument should select a deliberate, testable behavior rather than depend on a mutable global.
  • Keep values serializable: simple strings and explicitly encoded JSON travel cleanly from shells, Python launchers, schedulers, and job APIs.
  • Separate run inputs from project policy: put the former in arguments and the latter in settings, following the distinction documented in the FAQ.
  • Test each entry point: run one CLI invocation and one Python invocation with the same values, then compare the first requested URL and parsed types.

Or skip the browser setup:

If your crawl workflow also needs a rendered webpage image or PDF, ScreenshotNeo provides a single HTTP request instead of maintaining browser-capture code. It accepts cookie and consent banners as a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each cleanup step can be disabled. Only clean shots are billed: bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not charged, and the response reports the result in X-Page-Verdict and X-Billed headers. It also includes an MCP server with take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Example request (see the ScreenshotNeo API documentation):

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://quotes.toscrape.com -o shot.webp

The Free plan includes 1,000 screenshots per month with no card. Paid plans start at $5 for 3,000 shots; every feature is available on every plan, and yearly billing gives two months free. Create a free ScreenshotNeo account to get the 1,000 monthly screenshots without a card.

Frequently asked questions

Frequently Asked Questions

Which Scrapy documentation should I use for spider arguments?

The detailed spider-argument examples are on the Scrapy 2.12 spider page, while the current API and FAQ pages identify the present 2.19.0 documentation context. Match examples to the Scrapy version installed in your project.

Can I pass a spider name instead of a spider class from Python?

Yes. The crawler process and runner APIs accept either a spider class or a registered spider name, followed by the initialization arguments.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Does Scrapy convert command-line values to Python types?

No. Treat every supplied value as a string at the boundary and perform the conversion and validation your spider requires.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.