Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

Use Selenium through Scrapy’s downloader middleware, not as a replacement for Scrapy. Keep normal pages on Scrapy’s regular Request; yield SeleniumRequest only when a page needs JavaScript, clicks, scrolling, or other browser behavior. The middleware opens a configured browser, waits or runs scripts, then hands your callback a response that can be parsed with the same CSS and XPath selectors you already use in Scrapy.

How the integration works

Scrapy continues to manage scheduling, duplicate filtering, concurrency, retries, callbacks, and item pipelines. Selenium WebDriver supplies the browser-rendering step. A typical request travels through this sequence:

  1. Your spider yields an ordinary Request for static HTML or a SeleniumRequest for a JavaScript-dependent page.
  2. scrapy_selenium.SeleniumMiddleware receives the Selenium request in Scrapy’s downloader middleware chain.
  3. The middleware navigates the configured browser to the URL, applies a delay or explicit wait, and optionally executes JavaScript.
  4. The middleware returns the browser’s current page source as a Scrapy response.
  5. Your callback extracts data with response.css() or response.xpath(). If a browser-only action is still needed, the Selenium driver is available in response.request.meta['driver'].

Selenium WebDriver drives browsers natively and can run either on the same machine as Scrapy or through a remote Selenium Server. WebDriver is a W3C Recommendation, while newer Selenium documentation also describes WebDriver BiDi for bidirectional browser events.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Install the project dependencies

Create or activate the virtual environment used by the crawler, then install Scrapy, Selenium, and the third-party middleware package:

python -m pip install scrapy selenium scrapy-selenium

The scrapy-selenium package is a community middleware rather than part of Scrapy core. Check its compatibility with the Scrapy, Selenium, browser, and driver versions you deploy; those combinations can change independently.

Choose a browser and driver

Install a Selenium-compatible browser such as Chrome, Firefox, or Edge. Traditional Selenium setups also require the matching driver executable. Selenium Manager, included with supported Selenium distributions, can discover, download, and cache drivers and supported browsers when they are missing. Selenium’s Python documentation describes this automatic behavior for Selenium 4.6.0 and later (released November 4, 2022).

For reproducible production builds, you can still install and pin the browser and driver in your container or host image. For local development, Selenium Manager is often the simplest starting point.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Configure Scrapy’s middleware

Add the browser settings to settings.py. The exact driver path is optional when Selenium Manager can resolve it.

SELENIUM_DRIVER_NAME = "chrome"

# Use this when you manage the driver yourself.
# SELENIUM_DRIVER_EXECUTABLE_PATH = "/usr/local/bin/chromedriver"

# Use this instead when the browser is exposed by a Selenium Server.
# SELENIUM_COMMAND_EXECUTOR = "http://selenium:4444/wd/hub"

SELENIUM_DRIVER_ARGUMENTS = ["--headless=new", "--no-sandbox", "--disable-dev-shm-usage"]

DOWNLOADER_MIDDLEWARES = {
    "scrapy_selenium.SeleniumMiddleware": 800,
}

Use either SELENIUM_DRIVER_EXECUTABLE_PATH for a local driver or SELENIUM_COMMAND_EXECUTOR for a remote WebDriver endpoint. Do not configure both for the same execution mode. Headless arguments are useful on servers without a desktop; remove them while diagnosing a visual browser problem.

Build a SeleniumRequest spider

This minimal spider renders a products page, waits for product elements, and then uses ordinary Scrapy extraction:

import scrapy
from scrapy_selenium import SeleniumRequest
from selenium.webdriver.common.by import By
from selenium.webdriver.support import expected_conditions as EC


class ProductSpider(scrapy.Spider):
    name = "products"

    def start_requests(self):
        yield SeleniumRequest(
            url="https://example.com/products",
            callback=self.parse,
            wait_until=EC.presence_of_all_elements_located(
                (By.CSS_SELECTOR, ".product")
            ),
            wait_time=10,
            dont_filter=True,
        )

    def parse(self, response):
        for row in response.css(".product"):
            yield {
                "name": row.css(".name::text").get(),
                "price": row.css(".price::text").get(),
            }

wait_time provides a bounded delay. wait_until is preferable when a known element signals that rendering is complete, because it can finish as soon as the condition is met. The request type also supports screenshots and a script argument for custom browser-side JavaScript.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Wait for asynchronous content correctly

A page may return an initial HTML shell while JavaScript fetches the data later. Sleeping for an arbitrary number of seconds can be either too short or wasteful. Prefer an explicit Selenium expected condition tied to the page state you need.

yield SeleniumRequest(
    url="https://example.com/catalog",
    callback=self.parse,
    wait_until=EC.element_to_be_clickable(
        (By.CSS_SELECTOR, "button.load-more")
    ),
    wait_time=15,
)

Other useful conditions include presence or visibility of an element, a title containing text, or a URL change. Import the condition and locator classes from Selenium, then pass the callable to wait_until. Set a timeout that reflects the site’s behavior and your crawler’s failure policy; a timeout should fail clearly rather than silently extracting an empty page.

Scroll, click, and execute JavaScript

Use the request’s script argument for a controlled action that must happen before extraction, such as scrolling a lazy-loaded page:

def scroll_page(driver):
    driver.execute_script("window.scrollTo(0, document.body.scrollHeight);")


yield SeleniumRequest(
    url="https://example.com/gallery",
    callback=self.parse,
    script=scroll_page,
    wait_time=3,
)

When you need multiple interactions, retrieve the driver in the callback. Keep the final extraction in Scrapy so selectors, item validation, and pipelines remain consistent:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
def parse(self, response):
    driver = response.request.meta["driver"]
    driver.find_element(By.CSS_SELECTOR, "button.accept").click()
    driver.execute_script("window.scrollTo(0, document.body.scrollHeight);")
    # If the action changes the DOM, read the updated page source.
    html = driver.page_source
    rendered = scrapy.http.HtmlResponse(
        url=driver.current_url,
        body=html,
        encoding="utf-8",
    )
    for value in rendered.css(".result::text").getall():
        yield {"value": value.strip()}

Prefer a dedicated SeleniumRequest wait or script when possible. Direct driver work in a callback is more stateful, so avoid leaving cookies, tabs, or modal dialogs that can affect the next request.

Mix ordinary and browser requests

Do not send every URL through a browser. Use Scrapy’s lightweight request path for pages whose data is present in the response HTML or available from a documented endpoint. Reserve Selenium for pages that genuinely require client-side rendering or interaction.

Approach Best use Operational considerations
Scrapy Request Static HTML, feeds, and direct HTTP responses Lowest setup and resource overhead; no browser execution
Local Selenium JavaScript rendering, clicks, scrolling, screenshots You maintain browser and driver processes; each session consumes substantially more resources than an HTTP request
Remote Selenium Centralized browsers, isolated workers, or distributed execution Requires a reachable Selenium Server or Grid and network-aware timeouts

Use separate queues or request metadata if you need to enforce a lower concurrency for browser traffic. Browser sessions are heavier, and a site can also impose stricter rate limits when it sees real browser behavior.

Run Selenium remotely

Start a Selenium Server or Grid that exposes a WebDriver endpoint, then point Scrapy at it:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
SELENIUM_DRIVER_NAME = "chrome"
SELENIUM_COMMAND_EXECUTOR = "http://selenium-grid.internal:4444/wd/hub"
SELENIUM_DRIVER_ARGUMENTS = ["--headless=new", "--no-sandbox"]

DOWNLOADER_MIDDLEWARES = {
    "scrapy_selenium.SeleniumMiddleware": 800,
}

The remote host must be able to reach the target sites, resolve DNS, and launch the requested browser. Treat the executor URL as infrastructure configuration rather than spider code. Add connection and page-load timeouts at the deployment layer, and ensure the Grid can provide an isolated session for concurrent requests.

Production design and reliability

Control browser lifecycle

Middleware commonly starts or reuses a configured browser. Verify how the version you deploy handles session reuse, then make cleanup part of shutdown so orphaned browser processes do not accumulate.

Make selectors and waits resilient

Prefer stable attributes and a specific readiness element over brittle positional selectors. Detect an empty result set and record the URL, wait condition, and browser error instead of treating an empty page as a successful scrape.

Handle consent dialogs and authentication

For sites that require a click before content appears, use a targeted script or driver interaction. Supply authentication through the site’s permitted mechanism, and avoid embedding secrets in spider source; use Scrapy settings or deployment secrets.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Respect crawl policy

Browser rendering does not change the site’s terms, robots policy, authentication requirements, or legal obligations. Apply Scrapy’s download delays, concurrency limits, and retry rules to browser requests as carefully as to ordinary requests.

Troubleshooting common failures

“ModuleNotFoundError: scrapy_selenium”

Install the package in the same environment that launches Scrapy: python -m pip install scrapy-selenium. Confirm the interpreter with python -m pip show scrapy-selenium.

Driver or browser cannot be found

Check SELENIUM_DRIVER_NAME, install a compatible browser, or let Selenium Manager resolve the driver on a supported Selenium version. If you pin a driver path, verify that the Scrapy process can execute it and that its version matches the browser.

The middleware never runs

Confirm the setting name and indentation, and ensure scrapy_selenium.SeleniumMiddleware appears in DOWNLOADER_MIDDLEWARES. Also verify that the spider yields SeleniumRequest, not a normal scrapy.Request.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The callback sees no data

The page may still be loading, the selector may be wrong, or the content may be inside an iframe. Add an explicit wait_until condition for a known element, inspect response.text, and switch into the required iframe through the driver before reading its DOM.

Timeouts on a remote Grid

Test the executor URL from the Scrapy host, check that a browser node is available, and increase the connection or page timeout only after confirming network reachability. A remote session can fail before navigation if the Grid is saturated.

Headless mode behaves differently

Temporarily remove headless arguments and reproduce with a visible browser. Compare viewport size, user-agent, permissions, and timing; then restore only the arguments you actually need.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup

If your goal is a clean image or PDF rather than a Scrapy crawl, ScreenshotNeo provides a single-call website screenshot API and MCP server. It accepts cookie and consent banners like a visitor, removes more than 60 known consent platforms plus newsletter popups and chat widgets before capture, and lets you turn each cleanup step off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed; response headers identify the page verdict and whether the request was billed.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use the API documentation at https://screenshotneo.com/docs/ for all options. A basic request is:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

ScreenshotNeo also offers an MCP server with take_screenshot, get_page_info, and capture_pdf for Claude, Cursor, and other MCP clients. It includes full-page and element capture, device presets, custom waits and scripts, request blocking, cookies and headers, geolocation, resizing, caching, signed links, asynchronous webhooks, bulk capture of up to 100 URLs per call, and HTML/CSS-to-image. The free plan includes 1,000 shots per month without a card; paid plans start at $5 for 3,000 shots. Sign up for the free plan.

Cost and maintenance decisions

Local Selenium avoids a hosted-browser bill but shifts browser installation, upgrades, crashes, and capacity planning to your infrastructure. Remote Selenium centralizes those concerns and can simplify parallel sessions, but adds network latency and Grid operations. In either model, browser requests consume more CPU and memory than ordinary Scrapy requests, so measure queue depth and session stability in your own deployment rather than assuming a fixed throughput.

Review compatibility whenever you upgrade Scrapy, Selenium, the middleware, or the browser. The middleware package is third-party, and no single version combination remains universally valid as those projects evolve.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Frequently Asked Questions

Can I use Selenium and Scrapy in the same spider?

Yes. Yield normal Scrapy requests for static pages and SeleniumRequest only for URLs that require browser rendering or interaction.

Where do I parse the rendered HTML?

Parse the response passed to your callback with Scrapy CSS or XPath selectors. Access response.request.meta[‘driver’] only when another browser action is necessary.

Should I use a fixed sleep or an explicit wait?

Use an explicit wait tied to a known page condition whenever possible; a fixed wait is less predictable and can either miss late content or waste time.

Can Selenium run on another machine?

Yes. Set SELENIUM_COMMAND_EXECUTOR to a reachable Selenium Server or Grid endpoint and ensure that remote host can access the target site.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.