Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
For a JavaScript-rendered page, first check whether you can request the data directly. Scrapy says reproducing the request that contains the desired data is the preferred approach when practical: it can return structured data with less parsing and network transfer. Use a browser when the relevant request is hard to reproduce or the task depends on browser-visible behavior. If you want browser rendering while retaining Scrapy’s scheduling and processing workflow, use scrapy-playwright.
Choose between reproducing a request and rendering a page
A page that fills in content with JavaScript does not necessarily require a browser. Open the page’s developer tools and inspect the Network panel while it loads and while you perform the action that reveals the data. Look for a request returning JSON or another response with the information you need. If it is understandable and repeatable, reproduce it with a normal Scrapy request and parse the response.
Scrapy calls reproducing requests that contain the desired data the preferred approach when possible. It can avoid parsing a rendered page and transferring browser resources such as scripts, stylesheets, and images. See Scrapy’s guide to selecting dynamically loaded content.
Free tools Windows power users keep installed
One-click scans. No signup required.
| Approach | Best fit | Trade-offs |
|---|---|---|
| Reproduce the underlying request | The data request is identifiable, repeatable, and supplies the fields you need. | Requires understanding the request and any parameters, headers, or pagination it uses; usually avoids browser rendering and makes structured data easier to parse. |
| Render or interact with a browser | The request is difficult to reproduce, or the task needs browser behavior such as clicking a control or reading content generated in the page. | Requires browser binaries and browser lifecycle management, and adds wait/action choices and browser resource costs. |
If browser automation is appropriate and the project is already a Scrapy spider, Scrapy recommends scrapy-playwright rather than launching Playwright directly inside a callback. The integration works as a download handler: selected requests go through Playwright, then their responses return to Scrapy’s request/response workflow. Direct Playwright use in a callback bypasses much of Scrapy, including components such as middleware and duplicate filtering.
#1 Best Overall
Install scrapy-playwright and its browser
The project README currently lists minimum requirements of Python 3.10, Scrapy 2.7, and Playwright 1.40. These floors and installation instructions can change; check the current scrapy-playwright README and your environment before relying on particular versions.
-
Install the integration in the same Python environment as your Scrapy project:
pip install scrapy-playwright. -
Install a Playwright browser binary if one is not already available:
playwright install. To install only a selected browser, use the browser-specific command documented by Playwright, for exampleplaywright install chromium.Recommended: Fix Windows Errors and Clear Junk Files in Minutes - Free Scan →Recommended: Crashes or Glitches? A Free Driver Scan Usually Finds the Culprit →Recommended: PC Feels Slow? A Free Scan Shows What's Dragging Windows Down →Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy. -
Keep browser binaries compatible with the installed Playwright version. Playwright documents that its browser binaries correspond to specific Playwright versions, so after updating Playwright you may need to rerun the browser installation command. See Playwright’s browser installation documentation.
Configure Scrapy’s download handler
Register the Playwright handler for both HTTP schemes in your project’s settings.py. Keeping the regular Scrapy handler as the fallback matches the project’s documented configuration pattern:
DOWNLOAD_HANDLERS = {
"http": "scrapy_playwright.handler.ScrapyPlaywrightDownloadHandler",
"https": "scrapy_playwright.handler.ScrapyPlaywrightDownloadHandler",
}
With the handler configured, requests still use the normal Scrapy workflow unless you opt them into browser handling. This lets a spider use ordinary Scrapy requests for pages that do not need rendering and Playwright for selected pages.
Opt selected requests into Playwright
Set the playwright request metadata flag to a truthy value. Here is a minimal spider that requests one page through a browser and parses a heading from the returned response:
import scrapy
class DynamicPageSpider(scrapy.Spider):
name = "dynamic_page"
start_urls = ["https://example.com"]
def start_requests(self):
for url in self.start_urls:
yield scrapy.Request(
url,
callback=self.parse,
meta={"playwright": True},
)
def parse(self, response):
yield {
"url": response.url,
"title": response.css("title::text").get(),
"headings": response.css("h1::text").getall(),
}
Replace https://example.com and the selectors with the target page and the fields you have verified. For a basic crawl that does not need customized startup settings, the documented default browser can handle opted-in requests. If you need a named browser context, set playwright_context in request metadata; the integration supports contexts so requests can use separate browser state.
The response delivered to parse remains a Scrapy response, so use the usual Scrapy selectors and item/yield patterns to extract its content. A rendered response only contains the page state available when Playwright returns it; the right interaction and wait condition depend on how the particular site populates its content.
Wait for content or click a control before extraction
Use PageMethod to ask the integration to perform page actions before it returns the final response. For example, to wait until a known content selector appears:
import scrapy
from scrapy_playwright.page import PageMethod
class WaitForContentSpider(scrapy.Spider):
name = "wait_for_content"
start_urls = ["https://example.com/products"]
def start_requests(self):
for url in self.start_urls:
yield scrapy.Request(
url,
callback=self.parse,
meta={
"playwright": True,
"playwright_page_methods": [
PageMethod("wait_for_selector", ".product-card"),
],
},
)
def parse(self, response):
for card in response.css(".product-card"):
yield {
"name": card.css(".product-name::text").get(),
"url": card.css("a::attr(href)").get(),
}
Choose a condition that represents the result you need, such as a selector appearing after a request completes. A fixed sleep may be appropriate for a known site behavior, but it is not a universal substitute for waiting on the right condition: a short delay can return too early, while a long delay wastes time. Consult the integration’s PageMethod documentation for its supported request metadata and page-action options.
The Tool Desk
Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →For an interaction, replace the wait method with an action such as clicking a button, then wait for the content that the action reveals. A simplified example for a “load more” control is:
Rank #3
"playwright_page_methods": [
PageMethod("click", "button.load-more"),
PageMethod("wait_for_selector", ".product-card:nth-child(21)"),
],
Use selectors that match the actual page and its expected result. If the control can be clicked repeatedly to fetch additional batches, each click and the resulting stopping condition need to be handled deliberately; one click does not guarantee all records have loaded.
Manage pages and contexts without stalling a crawl
Most requests do not require the callback to own a Playwright page. The integration closes pages automatically unless you explicitly request that a page be retained or included. If your code takes ownership of a page, it also takes responsibility for closing it, including when an error occurs.
-
Do not retain pages unless you need direct page access beyond the configured page actions.
Recommended Free Tools
Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy. -
If you retain a page, close it after the work is complete.
-
Attach an errback that closes an owned page when a request fails. The integration README warns that pages left open count toward per-context page limits and can eventually freeze a crawl.
-
Manage browser contexts and browser instances explicitly in code that creates or owns them. Playwright separates pages and contexts; see its Browser API documentation.
A retained-page request needs both a success path and a failure path. The exact page metadata and cleanup code depend on whether the callback receives the page directly or accesses it through the response; follow the current integration README for the API details rather than leaving a page open on an exception.
Or skip the browser setup
If your goal is to capture a webpage as an image or PDF rather than extract structured records into Scrapy items, ScreenshotNeo is a separate website screenshot API and MCP server for developers. A single GET request returns a PNG, JPEG, WebP, or PDF; it is not a replacement for a Scrapy spider that needs to crawl and parse data.
For example, save a WebP screenshot with cURL:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
See the ScreenshotNeo API documentation for request options and output formats. ScreenshotNeo accepts cookie and consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and each response says which result occurred in its X-Page-Verdict and X-Billed headers. Its MCP server provides take_screenshot, get_page_info, and capture_pdf tools for AI agents. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000 shots.
Sign up for ScreenshotNeo’s free plan to try 1,000 screenshots a month without a card.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Troubleshoot common setup and crawl problems
Import or handler errors
Symptom: Scrapy cannot import scrapy_playwright.handler.ScrapyPlaywrightDownloadHandler, or opted-in requests do not use a browser. Likely cause: the package is missing from the active environment or the download handler is not registered for the relevant scheme. Fix: install the package in the environment that runs Scrapy and verify both http and https entries in DOWNLOAD_HANDLERS.
Do these 3 things before closing this tab:
1Scan for outdated or missing drivers - takes under a minute2Clear out junk files and repair common Windows errors3Fix the driver behind crashes, sound loss and screen glitchesBrowser executable is missing or incompatible
Symptom: Playwright reports that a browser executable cannot be found or does not launch. Likely cause: the browser binary was not installed for the installed Playwright version, or Playwright has since been upgraded. Fix: run playwright install (or the selected-browser install command) in the same environment and check the Playwright browser instructions.
Extracted content is empty
Symptom: the response arrives but the desired selector has no values. Likely cause: the data has not loaded, the selector does not match the rendered DOM, or the page’s data is better obtained from an underlying request. Fix: inspect the Network panel first; if using a browser, wait for a site-specific selector or perform the required interaction before parsing. Confirm the selector against the final response content.
Best Value
The crawl slows or freezes after retaining pages
Symptom: later requests stop progressing after callbacks retain browser pages. Likely cause: pages were not closed and have consumed the configured per-context page capacity. Fix: close owned pages on both success and failure, using a request errback for failure cleanup; avoid retaining pages when PageMethod is sufficient.
Plan for performance, reliability, and cost
Request reproduction is generally the simpler path when the data endpoint is clear: it avoids the browser’s rendering work and can provide structured fields directly. Browser rendering adds setup and resource overhead, and reliability depends on using a wait condition tied to the page’s actual behavior and cleaning up retained pages. There is no universal performance winner: the appropriate choice depends on whether the underlying request can be reproduced and whether the task needs browser interaction.
Quick wins for a faster PC:
Scan for outdated or missing drivers - takes under a minuteDriver Scan →Repair Windows errors before they cause bigger problemsFix Now →Keep browser use selective. Route only pages that require rendering through Playwright, use ordinary Scrapy requests elsewhere, and request only the interaction and page state needed for extraction. Browser resources and page limits are operational constraints, not reasons to treat a fixed delay as a reliable loading strategy.
Frequently Asked Questions
Does scrapy-playwright replace Scrapy?
No. It supplies a Playwright-backed download handler for selected requests while returning responses through Scrapy’s workflow.
Can I use Playwright without scrapy-playwright in a Scrapy spider?
You can, but Scrapy recommends scrapy-playwright for browser automation in an existing spider because direct Playwright use in callbacks bypasses much of Scrapy’s normal processing.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

