Quick wins for a faster PC:
Clear out junk files and repair common Windows errorsFree Scan →Scan for outdated or missing drivers - takes under a minuteDriver Scan →Pass a run-specific value with scrapy crawl myspider -a name=value. Scrapy’s default spider initializer exposes that value as an attribute, so your spider can read self.name. When you start a crawl in Python, pass the same values as keyword arguments to CrawlerProcess.crawl() or CrawlerRunner.crawl(). All spider arguments arrive as strings; parse and validate lists, numbers, booleans, and JSON yourself.
The shortest working examples
From a terminal, provide one -a option for each argument:
scrapy crawl myspider -a category=electronics -a region=west
With a Python launcher, pass keyword arguments when scheduling the spider:
from scrapy.crawler import CrawlerProcess
process = CrawlerProcess()
process.crawl(MySpider, category='electronics', region='west')
process.start()
In either case, the spider receives category and region as attributes. The command-line form is best for an operator starting one crawl; the Python APIs are for applications, scheduled jobs, and pipelines that already control execution.
The Tool Desk
Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →#1 Best Overall
Design the spider’s parameter contract first
Decide which inputs vary from run to run, what format each value uses, and what should happen when an input is missing. A useful contract might look like this:
| Argument | Example | Type at the boundary | Spider policy |
|---|---|---|---|
category |
electronics |
String | Optional; use a general listing when omitted |
region |
west |
String | Optional; validate against supported regions |
start_urls |
["https://example.test/a", "https://example.test/b"] |
String containing JSON | Required for a targeted run; decode and validate as a list |
limit |
100 |
String containing an integer | Convert with explicit range checks |
Keeping this contract explicit prevents a common mistake: assuming that a value that looks like a list, number, or Boolean has already been converted. Scrapy passes spider arguments as strings, including values supplied through Python keyword arguments.
Pass arguments from the command line
1. Start with a named spider
Run the command from the Scrapy project directory and use the spider’s name value:
scrapy crawl quotes -a tag=python
For multiple values, repeat -a; do not combine unrelated values into one option:
Do these 3 things before closing this tab:
1Repair Windows errors before they cause bigger problems2Fix the driver behind crashes, sound loss and screen glitches3Clear out junk files and repair common Windows errorsscrapy crawl quotes -a tag=python -a region=west -a limit=25
2. Read the attribute in the spider
The default initializer copies supplied arguments onto the spider. An optional argument can therefore be read with getattr and a default:
import scrapy
class QuotesSpider(scrapy.Spider):
name = 'quotes'
async def start(self):
tag = getattr(self, 'tag', None)
url = 'https://quotes.toscrape.com/'
if tag is not None:
url += f'tag/{tag}'
yield scrapy.Request(url, self.parse)
def parse(self, response):
for quote in response.css('div.quote'):
yield {
'text': quote.css('span.text::text').get(),
'author': quote.css('small.author::text').get(),
}
With this spider, scrapy crawl quotes -a tag=python requests the tag page, while omitting -a tag=... requests the site root. The current documentation also demonstrates an async def start() method; use the start-method style supported by the Scrapy version installed in your project.
Rank #2
3. Accept values in __init__ when you need custom setup
For simple access, no custom initializer is needed. Define one when you want to normalize or validate arguments before requests are generated, and call the base initializer so Scrapy can perform its normal setup:
import scrapy
class CatalogSpider(scrapy.Spider):
name = 'catalog'
def __init__(self, category=None, region=None, *args, **kwargs):
super().__init__(*args, **kwargs)
self.category = category
self.region = region
def start_requests(self):
if not self.category:
raise ValueError('category is required')
url = f'https://example.test/catalog/{self.category}'
yield scrapy.Request(url, callback=self.parse)
def parse(self, response):
yield {'url': response.url, 'region': self.region}
Use a default such as None for optional values. For required values, fail early with a clear message rather than issuing a broad or malformed request.
4. Quote shell values deliberately
Shell syntax is applied before Scrapy receives the argument. Quote values containing spaces, ampersands, question marks, or shell metacharacters:
scrapy crawl catalog -a query='noise cancelling headphones' -a region='north america'
For a URL, quote the complete value so the shell does not interpret its query string:
scrapy crawl catalog -a start_url='https://example.test/search?q=phone&sort=price'
If a value contains credentials or other sensitive material, avoid putting it directly in shell history. Prefer environment variables or a protected secret mechanism, then read the variable in your launcher and pass the resulting value to the crawl.
Start a crawl from Python
CrawlerProcess: the self-contained launcher
CrawlerProcess configures and starts Scrapy when your script owns the reactor. Pass the spider class (or a spider name) followed by keyword arguments:
Free tools Windows power users keep installed
One-click scans. No signup required.
from scrapy.crawler import CrawlerProcess
from myproject.spiders.catalog import CatalogSpider
process = CrawlerProcess()
process.crawl(
CatalogSpider,
category='electronics',
region='west',
limit='25',
)
process.start()
The keyword values become spider arguments. They are still boundary values, so the spider should convert limit to an integer and validate it before using it in pagination or filtering.
CrawlerRunner: when another application owns the reactor
Use CrawlerRunner when your surrounding application already manages the reactor. Its crawl method accepts the spider class or name plus positional and keyword initialization arguments:
from twisted.internet import reactor, defer
from scrapy.crawler import CrawlerRunner
from myproject.spiders.catalog import CatalogSpider
runner = CrawlerRunner()
def run():
return runner.crawl(
CatalogSpider,
category='electronics',
region='west',
)
d = run()
d.addBoth(lambda _: reactor.stop())
reactor.run()
Do not call reactor.run() a second time inside an application that already has a running reactor. That is the practical distinction between the process helper and the runner: choose the helper that matches who owns the event loop.
Async process and runner APIs
Current Scrapy API documentation also lists AsyncCrawlerProcess and AsyncCrawlerRunner for coroutine-based control flow. Their reactor and event-loop requirements depend on how the application is configured, so check the current Scrapy Core API documentation before embedding one in an existing asyncio service. The argument shape remains the same: pass run-specific values as keyword arguments to crawl.
Windows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallOutdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchParse and validate structured values
Scrapy does not turn a command-line string into a Python collection automatically. If you pass start_urls=https://a.example,https://b.example and iterate over it as though it were a list, Python iterates over individual characters. Choose an unambiguous representation and decode it explicitly.
JSON for lists and nested data
import json
import scrapy
class MultiStartSpider(scrapy.Spider):
name = 'multi_start'
def __init__(self, start_urls_json=None, *args, **kwargs):
super().__init__(*args, **kwargs)
if not start_urls_json:
raise ValueError('start_urls_json is required')
try:
value = json.loads(start_urls_json)
except json.JSONDecodeError as exc:
raise ValueError('start_urls_json must be valid JSON') from exc
if not isinstance(value, list) or not all(isinstance(url, str) for url in value):
raise ValueError('start_urls_json must be a JSON list of strings')
self.start_urls = value
def parse(self, response):
yield {'url': response.url}
Invoke it with a shell-quoted JSON value:
scrapy crawl multi_start -a start_urls_json='["https://example.test/a", "https://example.test/b"]'
The official spider documentation mentions json.loads() and ast.literal_eval() as possible parsing approaches. JSON is usually easier to specify consistently across shells and languages; whichever format you choose, validate its shape before requests are scheduled.
Numbers and booleans
def parse_limit(raw):
try:
limit = int(raw)
except (TypeError, ValueError) as exc:
raise ValueError('limit must be an integer') from exc
if limit < 1 or limit > 1000:
raise ValueError('limit must be between 1 and 1000')
return limit
def parse_bool(raw):
normalized = str(raw).strip().lower()
if normalized in {'1', 'true', 'yes'}:
return True
if normalized in {'0', 'false', 'no'}:
return False
raise ValueError('follow_next must be true or false')
Do not rely on Python’s bool('false'); every non-empty string is truthy. Define the accepted spellings and reject everything else.
URLs and lists with commas
A comma-separated convention is only safe when commas cannot occur in an item. URLs can contain commas, so JSON is safer for arbitrary URL lists. If you do use a delimiter, document escaping rules and test values containing the delimiter before deploying the spider.
Arguments versus settings
Scrapy’s FAQ says there is no rigid rule. Use spider arguments for inputs that change often between runs or apply only to one crawl, such as a target URL, customer partition, category, or temporary limit. Use settings for behavior that changes infrequently across runs, such as downloader middleware, concurrency policy, or a project-wide feed configuration. See the Scrapy FAQ for the documented distinction.
Putting stable configuration into -a makes every deployment command longer and easier to misconfigure. Putting a run-specific target into a settings file hides the input that actually changed. Keep the boundary visible: arguments describe this crawl; settings describe the project.
Choose the invocation that matches your environment
| Situation | Use | How arguments are supplied | Important constraint |
|---|---|---|---|
| Developer or operator starts one crawl | CLI | scrapy crawl name -a key=value |
Quote shell-sensitive values |
| Standalone Python script | CrawlerProcess |
process.crawl(Spider, key=value) |
Let the process helper start the reactor |
| Application already owns Twisted’s reactor | CrawlerRunner |
runner.crawl(Spider, key=value) |
Do not start a second reactor |
| Coroutine-based application | AsyncCrawlerProcess or AsyncCrawlerRunner |
Keyword arguments to the async crawl API |
Match the configured reactor and event loop |
The spider-arguments guide is available at Scrapy’s spider documentation. The documentation index currently identifies Scrapy 2.19.0; the linked spider page is version 2.12, while the API and FAQ links above provide current-version context. Check the version installed in your environment when an API detail differs.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Troubleshoot the common failures
“Unknown spider” or the command ignores your arguments
- Confirm that the first positional value after
scrapy crawlmatches the spider’sname, not its Python class name. - Run
scrapy listfrom the project directory to verify that Scrapy can discover the spider. - Put every custom value after its own
-aoption. A misspelled option or a value placed before the spider name is not a spider argument.
The attribute is missing
- Use
getattr(self, 'key', default)for optional values. - If you override
__init__, callsuper().__init__(*args, **kwargs)and accept the keyword arguments you intend to use. - Check spelling and capitalization;
regionandRegionare different attributes.
A list behaves like a string
That is expected until you parse it. Decode JSON (or another documented format), verify the resulting type, and only then iterate. Log the parsed type during development rather than assuming the input was converted.
Recommended Free Tools
Best Value
A number or Boolean has the wrong behavior
Convert numeric strings with int() or float() and reject conversion errors. For Booleans, accept a small explicit set such as true/false and reject ambiguous spellings. Never use truthiness alone to interpret the string false.
The Python launcher fails with a reactor error
If your script owns execution, use CrawlerProcess. If a framework or service already runs the reactor, use CrawlerRunner and attach completion handling instead of calling reactor.run() again. For asyncio integrations, verify the reactor and event-loop combination required by the current API documentation.
The request URL is malformed
Print or log the final URL after interpolation. Shell quoting problems, unescaped query characters, and unvalidated parameters can all alter the value before Scrapy builds a request. Pass a complete URL as one quoted argument and validate its scheme and host in the spider.
Operational practices that keep parameterized crawls reliable
- Document the contract: list every argument, whether it is required, its accepted format, and an example invocation beside the spider.
- Validate before scheduling requests: reject bad values in initialization or the first start method so a typo does not create a large unintended crawl.
- Use deterministic defaults: an omitted optional argument should select a deliberate, testable behavior rather than depend on a mutable global.
- Keep values serializable: simple strings and explicitly encoded JSON travel cleanly from shells, Python launchers, schedulers, and job APIs.
- Separate run inputs from project policy: put the former in arguments and the latter in settings, following the distinction documented in the FAQ.
- Test each entry point: run one CLI invocation and one Python invocation with the same values, then compare the first requested URL and parsed types.
Or skip the browser setup:
If your crawl workflow also needs a rendered webpage image or PDF, ScreenshotNeo provides a single HTTP request instead of maintaining browser-capture code. It accepts cookie and consent banners as a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each cleanup step can be disabled. Only clean shots are billed: bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not charged, and the response reports the result in X-Page-Verdict and X-Billed headers. It also includes an MCP server with take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients.
Example request (see the ScreenshotNeo API documentation):
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://quotes.toscrape.com -o shot.webp
The Free plan includes 1,000 screenshots per month with no card. Paid plans start at $5 for 3,000 shots; every feature is available on every plan, and yearly billing gives two months free. Create a free ScreenshotNeo account to get the 1,000 monthly screenshots without a card.
Frequently asked questions
Frequently Asked Questions
Which Scrapy documentation should I use for spider arguments?
The detailed spider-argument examples are on the Scrapy 2.12 spider page, while the current API and FAQ pages identify the present 2.19.0 documentation context. Match examples to the Scrapy version installed in your project.
Can I pass a spider name instead of a spider class from Python?
Yes. The crawler process and runner APIs accept either a spider class or a registered spider name, followed by the initialization arguments.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Does Scrapy convert command-line values to Python types?
No. Treat every supplied value as a string at the boundary and perform the conversion and validation your spider requires.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

