October DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsClean PCRecommendedOne scan can reveal what keeps slowing WindowsLook for cleanup and repair opportunities.Run ScanOctober DealsAmazon USDeal season is back - check today's better picksAmazon US: current deals, useful picks and tech finds.See Picks×
Skip to content
Apify

How to Migrate from Scrapy to a Cloud Web Scraping SDK

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

You can usually keep your existing Scrapy spiders when moving scraping workloads to the cloud. The key is to decide what you are changing: where spiders run, how requests are fetched, or whether to wrap the project in a cloud platform’s SDK and runtime. Those are different migrations, with different effects on deployment, storage, scheduling, and compatibility.

Start with one representative spider, verify the target platform’s current Python and Scrapy requirements, and compare its output and failure behavior with your existing run before moving scheduled production crawls.

Choose what you are migrating

“Move Scrapy to the cloud” can mean three things. You might keep Scrapy and move only its hosting; keep your runtime and add a managed request layer; or adapt the project to another platform’s SDK and Actor lifecycle. Your spider code may remain largely intact in each case, but the surrounding operational responsibilities change.

Path What changes What usually stays Scrapy Best fit
Managed Scrapy hosting Deployment target, scheduling, monitoring, and resource controls Spiders, project structure, much of the existing configuration You want a hosted place to run and monitor crawls
Managed request/API layer How requests are downloaded, and potentially browser or proxy handling Scrapy scheduler, spider logic, deployment, and output pipeline You need help fetching target pages but want to retain your runtime
Cloud SDK/runtime wrapper Platform entry point, configuration, storage and lifecycle integration Scrapy spiders, if the project matches the SDK’s supported layout You want platform-specific storage, events, or runtime features

These are decision boundaries, not interchangeable product labels. An API integration is not automatically cloud hosting, and wrapping a project in an SDK is not a promise that every deployment script or storage assumption will work unchanged.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Path 1: Move hosting and keep Scrapy

Scrapy’s deployment documentation describes Scrapyd, an open-source server for running and monitoring spiders, and Zyte Scrapy Cloud, a hosted service. Scrapy Cloud is documented as compatible with Scrapyd and able to use the same scrapy.cfg configuration approach as scrapyd-deploy. Zyte also describes a self-hosted deployment route using shub: install it, log in, then deploy. Its hosting page describes scheduling, monitoring, dashboards, and capacity controls. See the Scrapy deployment documentation and Zyte Scrapy Cloud.

“No rewrites” is a vendor description, not a guarantee for your project. Custom deployment scripts, environment variables, package pins, persistent state, and output destinations still need to be checked. Hosting changes should not be confused with spider or downloader changes.

Example deployment with shub

For the Zyte Scrapy Cloud route, the product documentation gives this general sequence after preparing the project and its configuration:

  1. Install the shub command-line tool using the method supported by your Python environment.
  2. Run shub login and authenticate to the service.
  3. From the Scrapy project directory, use shub deploy to deploy the project.
  4. Use the service’s scheduling and monitoring controls for the pilot spider, then compare its results with the old runtime.

Check the current Scrapy Cloud documentation and plan terms before adopting this workflow; service commands, limits, and account requirements can change.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Plan limits and price to check

Zyte’s Scrapy Cloud page, accessed 2026-09-29, lists these vendor-published terms. They are volatile product details, not a like-for-like cost comparison against other platforms.

Plan or unit Published terms
Starter Free forever; one hour of crawl time, one concurrent crawl, and seven-day data retention
Professional From $9 per unit per month; unlimited crawl time and concurrent crawls, and 120-day data retention
Scrapy Unit 1 GB RAM and one concurrent crawl

The page does not specify geography for these terms. Confirm current pricing, the definition of billable usage, and retention requirements with Zyte before budgeting or moving production data.

Path 2: Keep your runtime and add a managed request layer

scrapy-zyte-api connects an existing Scrapy project to Zyte API for request/download handling. This can change how Scrapy obtains pages without moving the scheduler, deployment, or output storage to hosted Scrapy Cloud. Zyte describes Scrapy Cloud as running spiders and Zyte API as helping keep requests unblocked; the API can also be used with a self-hosted runtime. Treat service capabilities as vendor claims and run a target-site pilot rather than assuming bans or access failures disappear. The Zyte API tutorial provides service context.

Compatibility before installation

The stable setup page for scrapy-zyte-api is labeled version 0.34.0. It lists Python 3.10 or later, Scrapy 2.0.1 or later, and a Zyte API subscription; a free trial is described. The optional scrapy-poet integration requires Scrapy 2.6 or later. Check the current setup guide against your pinned environment before changing dependencies.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

For Scrapy 2.10 and later, the documented add-on setting is:

ADDONS = {"scrapy_zyte_api.Addon": 500}

The setup guide says the add-on enables transparent mode by default. Authentication uses the ZYTE_API_KEY environment variable; keep the key in a secret manager or other secure environment configuration, not in source control.

Install and configure the package

Install the integration in the same controlled environment where you run the project:

python -m pip install scrapy-zyte-api

Then add the documented setting to your Scrapy settings file and supply the key through your runtime environment. A minimal shell example for a local test is:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
export ZYTE_API_KEY="your_api_key
a"
scrapy crawl your_spider

Replace the example value and spider name with your own; use your deployment platform’s secret configuration for hosted runs. Follow the setup guide for any additional settings required by your project and API usage.

Check Twisted reactor and asyncio assumptions

The integration setup warns that switching to twisted.internet.asyncioreactor.AsyncioSelectorReactor may require project changes. Imports can install Twisted’s default reactor before it can be changed, and code that uses Twisted Deferreds may need deliberate integration with asyncio. A reactor already installed in a process cannot be swapped for that run, so perform compatibility tests in a fresh process. Inspect custom event-loop code and early imports before changing production settings.

Path 3: Wrap the project in a cloud Actor SDK

Apify’s Python SDK guide says its CLI can convert an existing Scrapy project into an Apify Actor with one command when the project follows a standard Scrapy layout, including a root-level scrapy.cfg. The process creates Actor files and directories, installs the SDK and dependencies, and updates Scrapy settings with platform components. This is a platform-specific integration, not generic zero-change deployment. Read the current Apify Scrapy guide and Python SDK overview before using it.

Version and platform checks

The SDK overview identifies version 4.0 and states that Python 3.11 or later is required. It describes Scrapy support as well as Actor lifecycle, storage, platform events, and proxy capabilities. Confirm the exact current guide and version before copying a command or code sample; preview documentation and older instructions may not match the SDK you install.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Before production use, validate how the wrapped Actor receives input, writes to storage, manages its request queue, and shuts down gracefully. The guide’s AsyncCrawlerRunner/asyncio bridging approach and platform-specific settings should be followed in the context of that current guide, not transplanted blindly into an older project.

How to plan a reversible migration

  1. Inventory the existing job. Record Python and Scrapy versions, pinned dependencies, custom middleware and extensions, pipelines and exporters, environment variables and secrets, persistent state, scheduler assumptions, request volume, and concurrency.
  2. Set the change boundary. Decide whether you are moving hosting only, changing request handling, or adopting a platform SDK/runtime. Estimate and test each scope separately.
  3. Check documented compatibility. Compare your pinned Python and Scrapy versions with the current service requirements. For Zyte integration, verify package setup and add-on support for your Scrapy version. For the Apify SDK v4.0 overview, account for Python 3.11 or later.
  4. Select one representative spider. Include a normal page, JavaScript-rendered page if relevant, retries, pagination, and the actual output pipeline. A trivial spider does not test production behavior.
  5. Run old and new paths against comparable work. Compare item schemas and counts, duplicate handling, retries and errors, crawl duration, memory and concurrency, logs, and downstream delivery. These are validation dimensions, not published vendor benchmarks.
  6. Roll out in small steps. Keep the previous deployment configuration, define how to roll back, migrate a small group of jobs, and only then change scheduled production crawls.

What to compare before switching

  • Spider behavior: Are callbacks, middleware, retry rules, pagination, and item pipelines behaving as before?
  • Request access: Does the workload need a managed fetch/API layer, proxies, or browser rendering, and what exactly does the chosen service handle?
  • Operations: Who owns scheduling, monitoring, concurrency, capacity, and restart behavior after the move?
  • Storage: Where do crawl results, logs, and persistent queues go, and how are they retained and retrieved?
  • Security: Are API keys and site credentials injected securely, and do custom headers or cookies follow the new platform’s handling rules?
  • Cost: How does the actual workload map to the provider’s unit, time, concurrency, request, or storage billing? There is no established like-for-like price or performance comparison across the services covered here.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Troubleshooting common migration failures

Dependency or Python version conflicts

Symptom: Installation fails, imports break, or a deployed job behaves differently from local development. Cause: The project’s pinned Python, Scrapy, or transitive package versions do not meet the platform integration’s documented requirements. Fix: compare the lock file or requirements against the current service guide; reproduce in a clean environment before changing production pins.

Twisted reactor already installed

Symptom: Reactor configuration errors or asyncio integration failures appear at startup. Cause: An import installed the default reactor before the requested reactor could be selected, or Deferred-based code is not correctly bridged. Fix: review import order and custom event-loop code, then test in a fresh process using the integration’s current setup instructions.

Actor starts but input or results are missing

Symptom: A wrapped job runs but does not receive expected parameters or deliver data to the expected destination. Cause: The cloud Actor’s input and storage lifecycle differ from the previous runtime’s assumptions. Fix: test input parsing and output persistence explicitly, and verify request-queue and shutdown behavior using the platform’s Scrapy guide.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Hosted crawl differs from the local run

Symptom: Items, retries, logs, or crawl duration differ after a hosting-only move. Cause: Environment variables, dependency versions, concurrency, scheduling, state, or storage configuration changed. Fix: compare the pilot’s runtime settings and outputs field by field; do not attribute the difference to the spider until configuration is matched.

Target pages still fail

Symptom: A managed fetch layer does not retrieve some target pages reliably. Cause: Target behavior varies, and an API integration does not establish universal access or eliminate every block. Fix: inspect response and crawl errors, verify the service configuration for the target, and expand the pilot only after representative pages and retry cases pass.

Or skip the browser setup

If the work is simply to capture a page as an image or PDF rather than migrate a Scrapy crawler, ScreenshotNeo is a website screenshot API and MCP server for developers. A GET request takes a URL and returns a PNG, JPEG, WebP, or PDF. Its API can remove cookie/consent banners, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks/CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, with response headers identifying the page verdict and billing status. An MCP server offers take_screenshot, get_page_info, and capture_pdf for AI agents and MCP clients.

For a quick call, first create an API key and replace the placeholder below. The ScreenshotNeo API documentation describes the request and options.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
curl -G "https://api.screenshotneo.com/v1/shot" 
  -d access_key=YOUR_API_KEY 
  --data-urlencode url=https://stripe.com 
  -o shot.webp

ScreenshotNeo includes 1,000 screenshots a month free with no card; paid plans start at $5 for 3,000. Sign up for free.

Frequently Asked Questions

Can I keep my existing Scrapy spiders when I move to the cloud?

Usually, yes. A hosting move or request-layer integration can retain the spiders, but project-specific deployment, environment, storage, and dependency assumptions still need testing.

Does adding a scraping API move my Scrapy project to cloud hosting?

No. A request/API integration changes how pages are fetched; it does not by itself move the spider scheduler, deployment, or output storage to a hosted runtime.

Which migration path should I try first?

Choose the smallest change that solves the problem: hosting-only for managed operations, an API layer for request handling, or an SDK wrapper when you need platform-specific runtime and storage features.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Read next

Recommended PC Tool
Recommended PC Tool
Windows Errors? Fix Them Before They SpreadFree repair scan
Crashes, No Sound, or Screen Glitches?Free driver scan

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.