Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

To scrape a JavaScript-rendered page with Scrapy, use Selenium only for requests that need a browser. Configure the Selenium downloader middleware, yield a SeleniumRequest, and wait for the element or state that means the page’s data is ready. Scrapy selectors can then parse the rendered response as usual. A page reaching readyState is not proof that its JavaScript content has appeared.

When Scrapy needs Selenium

Scrapy’s normal downloader fetches HTTP responses; it does not run the page’s JavaScript in a browser. If the HTML response already contains the information you need, parse it directly with Scrapy. If the page builds its results in JavaScript, requires a click to reveal them, or only exposes the relevant content after browser-side interaction, Selenium can render and interact with the page.

The useful division of work is selective: let Scrapy handle ordinary requests and use Selenium for the pages that need a browser. Middleware connects Scrapy’s request/response flow to a Selenium-controlled browser. A SeleniumRequest asks the middleware to load the URL; the rendered HTML is returned in a response that your callback can process with response.css() or response.xpath().

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Use Scrapy alone when the needed fields are already present in the downloaded HTML.
  • Use Selenium through Scrapy when JavaScript rendering or browser interaction is needed and you still want Scrapy’s request and parsing workflow.
  • Do not assume a browser is always better. Browser rendering adds browser and driver setup and resource use, so applying it to every request can make a crawl more operationally complex.

Install the package and prepare a browser

Choose a Selenium-compatible browser and make its driver available to the environment running your spider. The scrapy-selenium middleware uses a browser name, a driver executable path, and browser arguments in settings. A documented package variant, scrapy-selenium4, specifies support for Selenium 4 or later and documents browser/driver settings and optional remote command execution. Check the selected package’s documentation for the configuration expected by the exact version you install; do not mix configuration from different variants.

For the standard scrapy-selenium package, install Scrapy, Selenium, and the middleware package in the same Python environment as the project. The following is a typical starting point; use the browser and executable path that actually exist in your environment:

python -m pip install scrapy selenium scrapy-selenium

Save settings in the Scrapy project’s settings.py. Replace the example executable path and choose arguments appropriate for your environment:

DOWNLOADER_MIDDLEWARES = {
    "scrapy_selenium.SeleniumMiddleware": 800,
}

SELENIUM_DRIVER_NAME = "chrome"
SELENIUM_DRIVER_EXECUTABLE_PATH = "/path/to/chromedriver"
SELENIUM_DRIVER_ARGUMENTS = ["--headless"]

The middleware must be enabled for Scrapy to process SeleniumRequest as intended. The driver path must point to a usable driver, and the configured browser name and driver must be compatible with the installed browser. If you use a remote Selenium command executor, use the selected package’s documented remote configuration rather than assuming a local executable path applies.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Build a spider that waits for the page’s actual data

Waiting for an arbitrary number of seconds is a weak synchronization strategy. A page can take longer than expected and fail, or finish sooner and waste time. Selenium’s explicit waits let the spider wait for a condition tied to the information it needs. For example, wait until the results container is visible before parsing it.

import scrapy
from scrapy_selenium import SeleniumRequest
from selenium.webdriver.common.by import By
from selenium.webdriver.support import expected_conditions as EC


class ResultsSpider(scrapy.Spider):
    name = "results"
    start_urls = ["https://example.com/search"]

    def start_requests(self):
        for url in self.start_urls:
            yield SeleniumRequest(
                url=url,
                callback=self.parse_results,
                wait_time=10,
                wait_until=EC.visibility_of_element_located(
                    (By.CSS_SELECTOR, ".results")
                ),
            )

    def parse_results(self, response):
        for item in response.css(".results .result"):
            yield {
                "title": item.css(".title::text").get(),
                "url": item.css("a::attr(href)").get(),
            }

Replace the example URL and selectors with the target site’s values. The wait_time supplies the maximum wait duration for the request, while wait_until supplies the condition Selenium should satisfy. The condition above waits for a visible element matching .results; if the target data appears in a different element or state, use a condition that represents that state instead.

Once the middleware returns the rendered response, parse it with normal Scrapy selectors. If a callback needs to perform additional browser operations, the middleware documents access to the driver through response.request.meta["driver"]. Use that only where browser interaction is actually required; for ordinary extraction, the response HTML and Scrapy selectors are simpler.

Choose waits, page-load behavior, and timeouts deliberately

Explicit waits for dynamic elements

Navigation finishing does not mean JavaScript-generated content is ready. An application may populate results after the document loads, reveal a field after a click, or update the page asynchronously. An explicit wait expresses the condition the spider needs before continuing. Selenium’s Expected Conditions cover predicates such as element existence, visibility, visible text, title matching, and staleness. Choose the one that matches the page behavior: existence is not the same as visibility, and visibility is not proof that the expected text has arrived.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A fixed sleep can be too short on a slow response and unnecessarily long on a fast one. Selenium describes race conditions caused by moving to the next command before a page action has completed as a primary cause of flaky tests. Prefer a condition-based wait; use a fixed delay only when a known delay is itself the requirement and no meaningful state can be checked.

Page-load strategy is separate from the data-ready condition

Selenium offers three page-load strategies. normal waits for the load event, eager waits for DOMContentLoaded, and none does not block WebDriver on the page-load event. These choices affect when navigation returns, not whether an SPA has finished fetching and rendering the particular data you plan to scrape. Pair a strategy with an explicit condition for the needed content.

Three timeout controls serve different purposes

  • Implicit timeout: controls how long element searches wait before raising an error.
  • Page-load timeout: governs how long navigation may take.
  • Script timeout: governs asynchronous script execution.

These controls are independent. Increasing an element-search timeout will not necessarily fix a navigation timeout, and a page-load strategy does not replace a wait for the results your spider needs. Set values in keeping with the target site and the behavior being controlled; no universal timeout is established for all sites.

Use request-level rendering controls when needed

The middleware’s request pattern supports more than waiting. Its documented controls include wait_time, wait_until, screenshot=True, and a script argument. A script can perform a browser-side action such as scrolling with window.scrollTo. The screenshot option stores PNG bytes in the response metadata.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
yield SeleniumRequest(
    url=url,
    callback=self.parse_results,
    wait_time=10,
    wait_until=EC.visibility_of_element_located(
        (By.CSS_SELECTOR, ".results")
    ),
    screenshot=True,
)

Use scrolling only when the site’s behavior requires it—for example, when content is loaded as the page scrolls—and follow the action with a wait for the newly loaded state. A scroll command by itself does not establish that lazy content has finished loading. If you need a screenshot for diagnosis, retrieve the PNG bytes using the metadata key documented by the middleware version you installed.

Keep browser work bounded in a Scrapy crawl

A rendered request carries browser setup and execution overhead that an ordinary HTTP request does not. Restrict Selenium requests to pages that need JavaScript or interaction, and leave static pages on Scrapy’s normal downloader. This reduces unnecessary browser work and keeps the simpler path available for pages that do not need rendering.

There are no general benchmark figures that establish a Selenium crawl’s throughput for every site. Actual resource use and speed depend on the target page, browser, synchronization condition, and deployment. For a particular crawl, compare the choices that matter: whether rendering returns the content correctly, whether the wait reliably identifies readiness, the throughput you need, browser resource use, operational complexity, and whether a remote Selenium executor is appropriate.

Do not make a page-load strategy faster by removing the very readiness check your extraction depends on. If the page continues adding content after readyState completes, a shorter navigation wait may simply return an incomplete page sooner. The reliable stopping point is the state that corresponds to the fields you intend to extract.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Troubleshooting common failures

The spider sees empty or incomplete HTML

  • Likely cause: the initial response does not contain JavaScript-generated results, or the callback runs before the relevant content appears.
  • Fix: make that request a SeleniumRequest and wait for a target-specific element, visible text, or other expected state before parsing. Do not treat document readiness alone as proof that the application data is ready.

The middleware does not handle the request

  • Likely cause: the downloader middleware is missing, disabled, or configured for a different package variant.
  • Fix: confirm the middleware setting is enabled in the active Scrapy settings and matches the package you installed. Check that the spider yields SeleniumRequest for browser-rendered pages rather than a normal request.

The browser does not start

  • Likely cause: the driver executable path is wrong, the executable is unavailable to the running environment, or browser and driver configuration do not match.
  • Fix: verify the configured browser name, executable path, and browser arguments in the same environment that runs Scrapy. If using remote execution, check the selected middleware variant’s remote setup instead of troubleshooting a local driver path.

The wait times out although the page opened

  • Likely cause: the selector or condition does not describe the page’s actual ready state, the element is not visible, or the content did not load.
  • Fix: inspect the rendered page and choose a condition that matches the actual result element or expected text. Distinguish element existence from visibility, and confirm the selector is correct before simply extending the timeout.

The crawl is slow or uses more browser resources than expected

  • Likely cause: browser rendering is being applied to pages that do not need it, or waits use a broad delay instead of a precise condition.
  • Fix: route static pages through Scrapy’s normal downloader, reserve Selenium for dynamic pages, and use state-based waits. Review the selected page-load strategy and independent timeout settings against the target’s behavior.

Or skip the browser setup

If the task is to capture a page image or PDF rather than extract structured fields in a Scrapy spider, ScreenshotNeo offers a one-request screenshot API. It is not a replacement for Scrapy selectors or browser interactions inside a scraping workflow; it is an option when the desired output is a screenshot or PDF.

For example, this cURL request saves a WebP screenshot of Stripe:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

See the ScreenshotNeo API documentation for the request options and response details. Before capture, it accepts the cookie or consent banner as a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be turned off. Bot checks and CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing status. Its MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients. The free plan includes 1,000 screenshots a month without a card; paid plans start at $5 for 3,000.

Sign up for ScreenshotNeo’s free plan to try the API without a card.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

FAQ

Can I use Scrapy selectors on a Selenium-rendered page?

Yes. The middleware returns rendered HTML in the Scrapy response, so callbacks can use response.css() and response.xpath().

Should I render every URL with Selenium?

No. Keep ordinary pages on Scrapy’s downloader and use Selenium for requests that need JavaScript rendering or browser interaction.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.