October DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsPC HealthRecommendedCrashes, freezes, slowdowns? Check your PC nowSpot repairable issues before they interrupt work.Check PCOctober DealsAmazon USDeal season is back - check today's better picksAmazon US: current deals, useful picks and tech finds.See Picks×
Skip to content
Callbacks

How to Pass Data Between Scrapy Callbacks: cb_kwargs, meta, and spider.state

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use cb_kwargs when a value belongs to your spider and should become an argument of the next callback. Create the follow-up Request with keys such as category or item, then give the callback parameters with the same names. Use meta mainly for data that Scrapy middleware, extensions, or other framework components must read. Use spider.state for spider-wide values that must survive a cleanly paused and resumed job.

The standard pattern: pass callback arguments with cb_kwargs

A callback is just a Python method invoked for a response. Scrapy does not infer arbitrary arguments from the URL, so put your application data on the next Request. The request’s cb_kwargs mapping is unpacked into keyword arguments when Scrapy calls the callback.

import scrapy

class ProductSpider(scrapy.Spider):
    name = "products"

    def parse(self, response):
        for product_url in response.css("a.product::attr(href)").getall():
            yield scrapy.Request(
                response.urljoin(product_url),
                callback=self.parse_product,
                cb_kwargs={
                    "category": "books",
                    "listing_url": response.url,
                },
            )

    def parse_product(self, response, category, listing_url):
        yield {
            "category": category,
            "listing_url": listing_url,
            "product_url": response.url,
            "title": response.css("h1::text").get(),
        }

The keys and callback parameter names must match. If you pass category but define parse_product(self, response, section), Python raises a missing or unexpected keyword-argument error. Give optional values defaults when a branch may omit them:

def parse_product(self, response, category=None, listing_url=None):
    ...

Add values before yielding the request

You can construct a request first and then edit its mapping. This is useful when a value is calculated in several steps.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
request = scrapy.Request(
    response.urljoin(details_url),
    callback=self.parse_details,
)
request.cb_kwargs["source_page"] = response.url
request.cb_kwargs["position"] = position
yield request

For ordinary spider-owned data, this is the documented recommended mechanism. It keeps callback inputs separate from framework metadata and makes the callback’s contract visible in its signature.

Passing a partially populated item to a detail callback

A common crawl follows a listing page, creates an item with the fields already available, and completes it on a detail page. Pass that object as one callback argument.

def parse_item(self, response):
    item = {
        "name": response.css("h1::text").get()
    }
    details_url = response.css("a.details::attr(href)").get()

    yield scrapy.Request(
        response.urljoin(details_url),
        callback=self.parse_details,
        cb_kwargs={"item": item},
    )

def parse_details(self, response, item):
    item["description"] = response.css(".description::text").get()
    item["sku"] = response.css(".sku::text").get()
    yield item

This is a shallow, in-memory handoff during the crawl. Treat the object as belonging to that request chain; do not assume that cloning a request creates independent nested structures.

When meta is the right field

Request.meta is intended for values consumed by Scrapy components, including downloader or spider middleware and extensions. A component can inspect a flag in meta without your callback having to accept another parameter.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
yield scrapy.Request(
    next_url,
    callback=self.parse_next,
    meta={"dont_retry": True, "audit_source": response.url},
)

Use cb_kwargs for your callback inputs and reserve meta for framework-facing controls or deliberately selected cross-request context. Never copy every key from one request’s metadata into an unrelated request. Scrapy or an extension may have inserted component-specific state; copying it can alter behavior. For example, propagating retry_times can reduce the retries available to the new request.

If a downstream callback also needs a small debugging value that a component does not consume, it can be placed in either field, but cb_kwargs communicates the intent more clearly. Keep the two namespaces separate in larger spiders so a middleware change cannot silently break callback signatures.

Errbacks: recover the same callback data

An errback receives a Failure, not a normal response. The failed request is available as failure.request, and its callback arguments remain in failure.request.cb_kwargs.

def parse(self, response):
    yield scrapy.Request(
        response.urljoin("/detail"),
        callback=self.parse_detail,
        errback=self.detail_error,
        cb_kwargs={"listing_url": response.url, "category": "books"},
    )

def parse_detail(self, response, listing_url, category):
    yield {"category": category, "listing_url": listing_url}

def detail_error(self, failure):
    request = failure.request
    args = request.cb_kwargs
    self.logger.error(
        "Could not fetch %s (category=%s, listing=%s)",
        request.url,
        args.get("category"),
        args.get("listing_url"),
    )

Reading from the request avoids relying on a response that does not exist. Use get() in an errback when different request branches may carry different keys.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Choose by scope and lifetime

Mechanism Intended reader Typical lifetime Use it for
cb_kwargs Your callback or errback One follow-up request chain Category, source URL, IDs, partially built items
meta Middleware, extensions, or intentionally selected request context The request and its descendants when explicitly copied Component flags and framework controls
spider.state The spider across batches Persisted job lifetime with JOBDIR Counters, checkpoints, and spider-wide state

Neither request field is a general shared-memory store. If several concurrent requests must update one global counter, protect that design against ordering assumptions; callback execution order is not a substitute for synchronization.

Copying, cloning, and JOBDIR persistence

cb_kwargs and meta are shallow-copied by Request.copy() and Request.replace(). The outer dictionaries are new, but nested lists, dictionaries, or item objects can still be shared. If a clone must have independent nested data, copy that value explicitly:

from copy import deepcopy

new_kwargs = deepcopy(old_request.cb_kwargs)
new_kwargs["item"]["attempt"] = 2
new_request = old_request.replace(cb_kwargs=new_kwargs)

When a spider uses JOBDIR, Scrapy serializes queued requests with Python’s pickle. Values in cb_kwargs and meta are deep-copied when written to and loaded from the job directory, so a callback receives a copy. Mutating that copy does not mutate the object that was originally queued. Every value must be serializable; an unserializable request may run during the current process but can be lost when the crawl pauses.

For spider-wide state that should survive clean pauses and resumes, use spider.state and Scrapy’s built-in state extension rather than threading a value through every request. Resume with the same Scrapy version that paused the job, and stop cleanly; an unclean stop can corrupt the job directory. This persistence model is separate from passing one item from a listing callback to a detail callback.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Inspect callback flow with the scrapy parse command

The command-line parser can invoke a callback and show the requests and items it yields. Supply callback arguments with --cbkwargs as a JSON string, or request metadata with --meta.

scrapy parse -c parse_product --cbkwargs '{"category":"books"}' https://example.org/product

Use this for a small, reproducible response before running a full crawl. It quickly exposes a misspelled callback name, a key that does not match the method signature, or a selector that produces no follow-up URL.

Troubleshooting callback handoffs

“unexpected keyword argument” or “missing positional argument”

Compare every key in cb_kwargs with the callback signature after response. Rename one side, or provide a default for an optional branch. Avoid passing an entire dictionary as several keys unless the callback actually needs each field.

Rank #4
ScrapTherapy® Cut the Scraps!: 7 Steps to Quilting Your Way through Your Stash
  • Country of Origin:US
  • CPSIA:N
  • Hazardous?:No
  • Tariff:4901990050

The value is present but changes unexpectedly

Check for nested mutable data shared by a request clone. Use deepcopy before modifying a cloned item, and remember that separate requests can complete in any order.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Middleware behavior changes after copying meta

Stop copying the complete metadata dictionary. Select only the keys your next request requires. Remove framework bookkeeping such as retry counters unless you deliberately understand its effect.

The errback cannot find the item

Read failure.request.cb_kwargs, not failure.value or a nonexistent response. Also ensure the request was created with an errback; otherwise the failure follows Scrapy’s normal error handling.

Paused jobs lose requests

Make every value in cb_kwargs and meta pickle-serializable, use a stable JOBDIR, resume with the same Scrapy version, and avoid unclean termination.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Performance and design guidance

  • Pass compact IDs or URLs rather than large response bodies. Large nested objects increase memory use while requests wait in the scheduler.
  • Use selectors to extract the fields needed by the next callback, then discard the response; do not retain whole responses in callback arguments.
  • Keep callback signatures explicit. It is easier to test parse_detail(response, item, source_url) than a callback that receives an opaque metadata bag.
  • Use meta only where a component needs the value. This prevents accidental coupling to middleware implementation details.
  • Use an item loader or a clearly defined item schema when many callbacks enrich the same record, and yield the completed item exactly once.

Or skip the browser setup

If the reason you are building a browser-based capture step around your crawler is simply to obtain page images or PDFs, ScreenshotNeo provides a one-request API instead. Its documented endpoint accepts a URL and returns PNG, JPEG, WebP, or PDF output:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Best Value
Scrap Quilt Secrets: 6 Design Techniques for Knockout Results
  • Suitable for all kinds of project works
  • Acid and toxic free
  • Designed for easy usage
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

See the ScreenshotNeo API documentation for all request options. The service accepts and removes cookie-consent banners, newsletter popups, and chat widgets before capture; bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing status. Its MCP server includes take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients. The free plan includes 1,000 screenshots per month without a card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.

FAQ

Can I pass positional arguments through a Scrapy request?

No. Use keyword keys in cb_kwargs and matching callback parameter names.

Should I pass a response object to the next callback?

No. Extract the needed values first. Responses are tied to their request and retaining them wastes memory.

Is spider.state shared by every callback?

Yes, it is spider-wide state, unlike request-scoped callback arguments. Its persistence depends on a configured, cleanly managed JOBDIR.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Frequently Asked Questions

What happens when a callback needs data from two earlier requests?

Combine the required values into one serializable structure and pass that structure in the next request’s cb_kwargs; do not depend on callback execution order.

Can callback arguments contain custom item classes?

They can during an in-memory crawl if the object is usable by your code, but a JOBDIR run also requires the value to be pickle-serializable.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Read next

Recommended PC Tool
Recommended PC Tool
Outdated Drivers Are Slowing You DownFree scan - exact matches
Windows Errors? Fix Them Before They SpreadFree repair scan

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.