Hardware FixRecommendedDevice not working? Your driver may be the problemCheck updates for common hardware issues.Fix DriversOctober DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsPC HealthRecommendedCrashes, freezes, slowdowns? Check your PC nowSpot repairable issues before they interrupt work.Check PC×
Skip to content
Docker

How to Monitor and Manage a Scrapy Spider: Stats, Live Control, Resuming Jobs, and Alerts

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The reliable way to monitor a Scrapy spider is to combine run-level statistics, live controls, durable job state, data-quality checks, and notifications. Scrapy’s built-in Stats Collector and logging extensions show whether work is progressing; the Telnet console lets you pause, resume, or stop a live engine; JOBDIR allows a cleanly stopped crawl to continue; and an extension such as Spidermon can detect malformed output and send alerts. For history across runs, dashboards, and operations at scale, export these signals to a monitoring system or use a managed deployment.

What to monitor in every crawl

A process that is still running is not necessarily healthy. A useful baseline combines progress, throughput, completion, and data-quality signals. Scrapy’s Stats Collector keeps a table for each open spider and closes it when the spider closes. The default MemoryStatsCollector retains the last run in memory; it is not a durable multi-run history.

Core counters

  • Timing: start time, finish time, elapsed duration, and the finish reason.
  • Responses: received response counts and HTTP status categories.
  • Items: scraped and dropped item counts.
  • Progress: pages crawled and item/page rates from LogStats.
  • Failures: retries, download errors, exceptions, and close reasons such as finished, cancelled, or shutdown.

Judge those values against the spider’s expected behavior. A product-catalog crawl may legitimately receive thousands of responses before producing an item, while a detail-page spider may expect nearly one valid item per response. Add custom stats for outcomes that matter to your domain:

self.crawler.stats.inc_value("items/validated")
self.crawler.stats.inc_value("items/missing_price")
self.crawler.stats.inc_value("responses/catalog")

Read counters from an extension or signal handler through crawler.stats. Export them to a time-series database, log pipeline, or dashboard when you need retention, cross-run comparisons, or threshold alerts. Do not assume the built-in collector provides that history automatically.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use rates, not only totals

Record snapshots at intervals. If response and item counters stop moving while the process remains alive, investigate a stalled download, exhausted concurrency, a blocked domain, or a callback that is waiting indefinitely. A falling item-to-response ratio can indicate a selector change or an anti-bot response even when the crawl eventually finishes.

Inspect a live spider safely

Telnet console

Scrapy’s Telnet console opens a Python shell inside the crawler process. It exposes objects including crawler, engine, spider, stats, and settings. Connect only locally or through a protected SSH tunnel or VPN: the Telnet transport is unencrypted, and a password does not encrypt the connection. Disable the console when you do not need it.

From the console, inspect current values:

>>> stats.get_stats()
>>> engine.running
>>> spider.name

Control the engine with the documented methods:

>>> engine.pause()    # stop scheduling new requests temporarily
>>> engine.unpause()  # continue scheduling
>>> engine.stop()     # close the crawler

Pause is useful before maintenance or while checking a suspected data problem. Stop deliberately rather than killing the process so Scrapy can persist state and emit its normal closing signals.

Lifecycle signals

Extensions can connect to spider_opened, spider_closed, engine_started, and engine_stopped. The spider-closed signal includes a reason, allowing your exporter or alerting code to distinguish a normal finish from cancellation or shutdown. Use these hooks to publish run metadata, flush metrics, and clean up temporary resources.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Pause and resume a long crawl with JOBDIR

Set JOBDIR to a private, distinct directory for one spider job. Scrapy stores scheduled requests, duplicate-filter state, and spider state there so the same command can resume after a clean stop.

scrapy crawl products -s JOBDIR=/var/lib/scrapy/jobs/products-2026-09-29

Run the command again with the same directory after the process has closed cleanly. Follow these rules:

  • Use one job directory for one job; never share it between different spiders or concurrent runs.
  • Restrict filesystem permissions. Untrusted users must not be able to write to the directory.
  • Resume with the same Scrapy version. After upgrading or downgrading, create a new directory rather than relying on implementation details of the old state.
  • Requests must be serializable for durable queues. Non-serializable requests remain in memory and can disappear when the process stops.
  • Enable SCHEDULER_DEBUG to log requests that cannot be serialized.
  • A sudden power loss or kill can corrupt state. Treat an unclean shutdown as a reason to validate the directory before trusting a resume.

“Pause” through Telnet and “resume later” through JOBDIR solve different problems: the former controls a live engine; the latter persists a job so a later process can continue.

Detect bad output, not just a dead process

A crawl can finish with a successful-looking exit while returning incomplete or malformed data. Define checks around the data your application actually needs:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • required-field coverage, such as non-empty product IDs and prices;
  • plausible item counts compared with the expected catalog size;
  • duplicate rates and rejected-item totals;
  • HTTP status or content changes that indicate a block page;
  • freshness, such as the newest timestamp seen in the source.

Spidermon can validate scraped output against schemas or models, evaluate conditions based on Scrapy stats, generate reports, and notify through email, Slack, Telegram, or Discord. Set thresholds from historical behavior and business requirements; there is no universal “healthy” item count.

For a smaller system, a custom extension can evaluate crawler.stats.get_stats() in spider_closed and fail the deployment or send a webhook when a required condition is false. Keep operational alerts separate from parser logs so an on-call person can see the reason, run identifier, and affected source immediately.

Build a practical monitoring workflow

  1. Define the baseline. Capture normal duration, response volume, item volume, error rate, and item-to-response ratio for each spider.
  2. Emit structured metrics. Add custom counters for validation, skips, blocked pages, and source-specific outcomes.
  3. Log periodically. Use built-in logging extensions such as LogStats, LogCount, and PeriodicLog to make stalled progress visible.
  4. Persist run history. Export final stats and lifecycle events to durable storage; retain enough history to compare the current run with recent normal runs.
  5. Alert on symptoms. Combine a no-progress timeout with data-quality thresholds so a slow but valid crawl is not confused with a parser failure.
  6. Attach recovery context. Include the spider name, job directory, finish reason, last counters, deployment version, and a link to logs in every notification.

Troubleshooting common failures

The process is alive but counters do not change

Check whether the engine is paused, whether concurrency slots are occupied by slow requests, and whether the downloader is waiting on a timeout. Inspect recent logs and response/error counters. If the source is blocked, stop safely, adjust request handling, and restart with a validated job directory.

The crawl finishes with too few items

Compare response and item counts with the baseline. Inspect status codes and sample bodies for consent pages, bot checks, changed markup, or an empty client-rendered shell. Add a parser validation counter and make the run fail when required fields fall below your agreed threshold.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Resume skips requests or loses work

Confirm that the same JOBDIR and Scrapy version are being used, that only one process owns the directory, and that the previous shutdown was clean. Review SCHEDULER_DEBUG output for non-serializable requests. If the state is corrupted, preserve it for diagnosis and start a new job directory rather than repeatedly retrying a damaged one.

Telnet will not connect

Verify that the console is enabled and that you are connecting to the correct local port. In remote environments, use an SSH tunnel or VPN instead of exposing Telnet publicly. If your security policy disallows an unencrypted console, disable it and rely on signals, metrics, and controlled deployment commands.

Alerts fire on every run

Inspect whether the threshold reflects a real baseline, whether a source changed, or whether the check runs before all items are processed. Alert on a combination of conditions and include the finish reason. Version parser and schema changes alongside the spider so an intentional change does not look like an unexplained outage.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Choose where the spider runs

Scrapy’s official deployment guidance describes three operating models:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Option Who operates infrastructure Best fit Trade-offs to examine
Scrapy Cloud by Zyte Managed service Teams wanting hosted scheduling, scaling, monitoring, and storage Vendor limits, data location, integrations, and current pricing
Scrapyd Your team Self-managed services with direct operational control Own server security, scheduling, logs, retention, and upgrades
Docker deployment Your team or platform team Containerized CI/CD and custom orchestration Cluster operations, persistent storage, observability, and job recovery

A Zyte vendor page observed on September 29, 2026 lists a Starter tier as “Free forever” with one hour of crawl time, one concurrent crawl, and seven-day data retention. The same page lists Professional from $9 per unit per month and defines a unit as 1 GB of RAM and one concurrent crawl; it describes unlimited crawl time and concurrent crawls plus 120-day retention for Professional. These are time-sensitive vendor claims: verify geography, billing, limits, and the live price before committing.

Choose based on operational ownership, required concurrency, retention, data-handling rules, integrations, and the maintenance capacity of your team—not on hosting branding alone.

Or skip the browser setup

If your monitoring workflow needs an image of a dashboard, status page, or crawl report, ScreenshotNeo returns a clean website screenshot or PDF through one GET request. It accepts cookie or consent banners as a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be disabled. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and whether the request was billed. Its MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients.

cURL:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

See the ScreenshotNeo documentation for options such as full-page capture, CSS selectors, custom headers, cookies, waits, request blocking, signed links, asynchronous jobs, webhooks, bulk capture, caching, and PDF settings. The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 screenshots. Create a free ScreenshotNeo account.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Operational checklist

  • Do stats include timing, finish reason, responses, items, drops, retries, and custom validation counters?
  • Can you view progress without exposing Telnet to the public internet?
  • Does each resumable job have a unique protected JOBDIR?
  • Are non-serializable requests and unclean shutdowns handled?
  • Do alerts cover both stalled execution and invalid output?
  • Are run history and logs retained outside process memory?
  • Is the deployment model appropriate for your concurrency, data, and maintenance requirements?

Frequently Asked Questions

Does Scrapy automatically keep a long-term history of every run?

No. The default MemoryStatsCollector retains the last run in memory. Export stats to durable storage or use a monitoring platform for historical comparisons.

Can I share one JOBDIR between spiders?

No. A job directory belongs to one job and should not be shared between different spiders or concurrent processes.

Is Scrapy Telnet safe to expose publicly?

No. Its transport is unencrypted. Keep it local or behind an SSH tunnel or VPN, or disable it.

What should an alert contain?

Include the spider and run identifiers, finish reason, key counters, threshold that failed, deployment version, and links to logs.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Read next

Recommended PC Tool
Recommended PC Tool
PC Slower Than It Used to Be?Free scan - under a minute
Crashes, No Sound, or Screen Glitches?Free driver scan

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.