Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Most Scrapy errors become easier to diagnose once you identify whether they originate during spider import, reactor setup, exception-driven control flow, or a network request. Start with the earliest meaningful exception in the full traceback; a later message may only describe the consequence. The fixes below reflect Scrapy 2.16–2.19 documentation, including settings and reactor details documented for 2.19.0. Check your installed version before relying on defaults or copied settings.

Start with the first useful error line

Keep the complete traceback and logs rather than focusing only on the final line. Locate the earliest relevant exception, then classify when it occurs: startup or spider loading, reactor installation, callback or item processing, or a live request/response. This separates an import problem from an intentional Scrapy exception or a network behavior issue.

  1. Preserve the traceback. Note the original exception type and the module or component named closest to its origin.
  2. Classify the stage. Startup failures usually point to imports or settings; request-time failures may involve callbacks, middleware, or the connection.
  3. Check the version and setting scope. Scrapy defaults can change between versions, and project settings, command defaults, and spider-level settings can affect the final configuration.
  4. Trace the component that raised it. Use the traceback and the exception’s documented purpose before suppressing or removing anything.

Scrapy’s Debugging Spiders documentation covers traffic inspection and debugging uncaught exceptions. Its exceptions reference distinguishes several exception types that may be expected control flow, not evidence of a broken crawl.

Reactor errors: installed reactor does not match

A mismatch usually means a Twisted reactor was installed before Scrapy could install the one selected by TWISTED_REACTOR. Importing twisted.internet.reactor can install a reactor as a side effect. Once installed, it cannot be replaced at runtime. The import may be in your code or in a dependency imported during startup, so the line that reports the mismatch may not be where the problem began.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Find and defer early imports

  1. Search project modules and startup dependencies for module-level imports such as from twisted.internet import reactor.
  2. If the reactor is needed only inside a method or function, move the import there so Scrapy can configure the reactor before that code runs.
  3. Check the configured TWISTED_REACTOR against the reactor already installed; do not treat changing the setting as a substitute for fixing an early import.

For example, avoid a module-level reactor import in a spider:

class ExampleSpider(scrapy.Spider):
    name = "example"

    async def start(self):
        from twisted.internet import reactor
        # Use reactor here only if this code path requires it.
        yield scrapy.Request("https://example.com")

Adapt imports and the spider’s start method to your project and Scrapy version. The important point is the import’s timing, not the example URL. Scrapy’s asyncio documentation discusses reactor installation and import side effects.

Runner APIs and reactor installation order

CrawlerRunner and AsyncCrawlerRunner require the matching reactor to be installed before the runner is used. By contrast, the CLI and process APIs can install a reactor when appropriate. Scrapy’s install_reactor() does not replace a reactor that is already installed, so calling it after an early import will not repair a mismatch. Check the API you use and arrange initialization before constructing or using a runner.

Defaults depend on Scrapy version

The Scrapy 2.19 settings reference lists twisted.internet.asyncioreactor.AsyncioSelectorReactor as the default TWISTED_REACTOR and notes a default change in version 2.13. Do not assume this is the default in every installed release: verify the settings reference for your version and the effective project configuration.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Errors involving reactor-free mode

TWISTED_REACTOR_ENABLED configures whether Scrapy uses a Twisted reactor. The asyncio troubleshooting reference describes errors involving an attempted reactor import when reactor use is disabled, a reactor already installed while Scrapy is configured without one, no reactor installed when Scrapy expects one, and a class that cannot run without the reactor.

  • Check whether the failing code path or dependency actually needs reactor-dependent functionality.
  • Look for imports that installed the reactor before Scrapy applied its configuration.
  • Determine whether a class or integration requires the traditional reactor and therefore cannot be used in reactor-free mode.
  • Do not try to enable or disable TWISTED_REACTOR_ENABLED per spider; the documentation says per-spider use is unsupported.

Resolve the conflict at the project or component level: align the configuration with the code’s actual requirements, or remove the premature/reactor-dependent import where it is not needed. Reactor-free mode is not a generic fix for a mismatch.

Spider import and settings errors

Scrapy’s spider loader normally reports failures loudly when importing spider classes from SPIDER_MODULES raises ImportError or SyntaxError. Follow the traceback into the named module: the spider file itself may be syntactically invalid, or one of its imports may fail because a dependency or another project module cannot load.

  1. Read the traceback from the originating exception upward to see which project module triggered it.
  2. Check syntax in the spider and imported modules.
  3. Verify that dependencies referenced in the traceback are installed and importable in the environment running Scrapy.
  4. Review the effective settings, including project settings in settings.py, command-specific defaults, and spider-level settings.

SPIDER_LOADER_WARN_ONLY = True changes an import failure into a warning rather than fixing the failed import. Use it only when that reporting behavior is intentional; it does not make the spider load successfully. The Scrapy settings reference documents this option and the reactor settings.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Exceptions that can be normal Scrapy control flow

Before treating an exception name as a bug, identify where it was raised and what that component is trying to do. Scrapy documents several exceptions with deliberate purposes:

Exception Meaning and where to investigate
CloseSpider(reason='cancelled') A spider callback can raise it to request that the spider stop. Check the callback’s stopping condition and reason.
DropItem An item pipeline stage raises it to stop processing that item. Inspect the pipeline condition and determine whether dropping the item is expected.
IgnoreRequest The scheduler or downloader middleware can raise it to indicate that a request should be ignored. Trace the request and the middleware or scheduling rule.
NotConfigured A component constructor can raise it to leave an extension, item pipeline, downloader middleware, or spider middleware disabled. Check the component’s configuration and whether it is intended to be enabled.
NotSupported Indicates an unsupported feature. Identify the feature and component that rejected it rather than assuming the whole crawl failed.
StopDownload(fail=True) A bytes_received or headers_received signal handler can raise it to stop downloading. With the default fail=True, the request errback runs; with fail=False, its callback runs. The response body may be truncated, and fail is keyword-only.

For StopDownload, make callback and errback logic tolerate partial response content when applicable. Removing or suppressing an exception without checking its intended control-flow role can change which items are processed, whether a spider stops, or which request handler runs. Definitions are in Scrapy’s exception reference.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

When logs do not explain request behavior

If startup succeeds but the request or response behavior remains unclear, inspect the traffic or configure a debugger to catch uncaught exceptions. The debugging method matters because some approaches observe the connection passively while others sit in the connection path.

Approach Connection effect What it helps with Trade-off
Passive packet capture Observes traffic without changing the spider’s connection path. Examining captured network traffic. The Scrapy guide describes it as non-interfering; it does not modify requests in transit.
Intercepting proxy such as mitmproxy Routes traffic through an additional proxy hop. Inspecting and modifying traffic. Because it changes the network path, it can affect low-level behavior; do not assume the proxied result exactly matches a direct connection.
Debugger catching uncaught exceptions Does not require a traffic proxy. Stopping at uncaught exceptions to inspect execution state. Configure the debugger to catch uncaught exceptions and reproduce the failing path.

Use the official debugging guide for its traffic-inspection approaches and debugger guidance. A proxy can reveal or alter traffic, but its extra connection hop means it is not a neutral observation method.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Or skip the browser setup

If you need a rendered screenshot of a page to inspect what a browser sees, ScreenshotNeo is a website screenshot API and MCP server. It is separate from Scrapy’s request debugging: it does not diagnose reactor imports, spider loading, or Scrapy callback exceptions.

One GET request returns an image or PDF. This cURL example saves a WebP capture of the example URL:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp

See the ScreenshotNeo API documentation for request options. Cookie banners, newsletter popups, and chat widgets are removed before the shot; each cleanup step can be turned off. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing status. An MCP server provides take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000.

Sign up for ScreenshotNeo’s free plan.

Common troubleshooting checks

  • “Installed reactor does not match.” Search for top-level Twisted reactor imports in project modules and dependencies. Defer the import that occurs too early, then check the configured reactor and installed Scrapy version.
  • Reactor import is forbidden or missing. Compare the component’s reactor requirements with TWISTED_REACTOR_ENABLED; check whether a dependency imported Twisted before Scrapy initialized. Do not configure this option as a per-spider switch.
  • Spider is not loading. Follow the original ImportError or SyntaxError to the module, syntax problem, or dependency. SPIDER_LOADER_WARN_ONLY changes reporting, not the underlying failure.
  • A Scrapy exception appears in the log. Check the exception’s purpose and origin. It may intentionally stop an item, ignore a request, disable a component, or stop a spider or download.
  • Logs do not explain the response. Inspect traffic or catch uncaught exceptions in a debugger. Remember that an intercepting proxy changes the connection path and may change behavior.

FAQ

Does an exception in Scrapy always mean the crawl crashed?

No. Some documented Scrapy exceptions implement intended control flow. The consequence depends on the exception and the component that raised it.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Can I change the reactor after Scrapy has installed one?

No. An already installed Twisted reactor cannot be switched at runtime; fix the initialization order or use configuration compatible with the reactor already selected.

Does SPIDER_LOADER_WARN_ONLY fix a broken spider import?

No. It changes the surfaced import failure to a warning, but the import problem remains.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.