What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

Use scrapy-playwright when a page needs a real browser to execute JavaScript, then parse the rendered response with your Scrapy spider as usual. Install the package and browser binaries, enable its download handler and Scrapy’s asyncio reactor, and opt individual requests in with meta={"playwright": True}. If the site exposes the data through a request you can reproduce directly, Scrapy recommends that lower-overhead route instead of rendering a whole page.

What scrapy-playwright does—and when to use it

scrapy-playwright connects Scrapy’s request-and-response workflow to Playwright for Python. For requests you mark, the integration opens a browser page, lets it load and run JavaScript, and returns a Scrapy response that you can handle in a callback. Unmarked requests continue through Scrapy’s ordinary downloader. See the scrapy-playwright README for the maintained configuration and options.

This is useful when the content you need appears only after JavaScript runs, depends on browser events, or requires browser-only output such as a screenshot. It is not automatically the best way to collect every page. Scrapy’s dynamic-content guidance says to reproduce the underlying data requests when practical: they can provide structured, complete data with less parsing time and network transfer. Scrapy’s documentation also says, “We recommend using scrapy-playwright for a better integration.”

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Try ordinary Scrapy requests first if the needed data is available from a reproducible HTTP request or API endpoint.
  • Use browser rendering if the necessary content or interaction depends on JavaScript execution or browser behavior that is difficult to reproduce directly.
  • Use a browser for browser-only output when the task specifically needs something such as a screenshot.

Browser rendering brings browser processes and their resource use into the crawl. The sources do not establish a general performance benchmark, so measure your own target and workload rather than assuming a particular speed or success rate.

Prerequisites and installation

The project README lists these minimum versions: Python 3.10 or newer, Scrapy 2.7 or newer, and Playwright 1.40 or newer. The commands below install the integration and then its browser binaries:

python -m pip install scrapy-playwright
playwright install

The second command downloads browsers for Playwright. If you only need particular engines, the project documents installing a subset, for example:

playwright install firefox chromium

Run installation in the Python environment used to run your Scrapy project. If the browser executable is missing at crawl time, rerun the browser-install command in the environment where Playwright is installed.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Enable the integration in Scrapy settings

Add the HTTPS handler and asyncio reactor to your project’s settings.py:

DOWNLOAD_HANDLERS = {
    "https": "scrapy_playwright.handler.ScrapyPlaywrightDownloadHandler",
}

TWISTED_REACTOR = "twisted.internet.asyncioreactor.AsyncioSelectorReactor"

Registering the HTTPS handler is normally sufficient because most modern sites use HTTPS. The handler is not a blanket instruction to render every request: the request’s playwright metadata flag controls whether it uses Playwright. An unmarked request continues through Scrapy’s regular downloader.

If your project also registers an HTTP handler, plan persistent browser-profile ownership carefully. The README warns that both handlers can attempt to open the same persistent profile, which can cause a conflict.

Build a minimal JavaScript-rendering spider

This example requests one HTTPS page through Playwright and extracts its title from the response after browser rendering. Save it as a spider in your Scrapy project, with the settings above enabled:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
import scrapy


class ExampleSpider(scrapy.Spider):
    name = "example"

    async def start(self):
        yield scrapy.Request(
            "https://example.org",
            meta={"playwright": True},
        )

    async def parse(self, response):
        yield {
            "url": response.url,
            "title": response.css("title::text").get(),
        }

Replace the example URL and extraction selector with the page and data you need. The key step is meta={"playwright": True}; without it, that request is not opted into browser rendering. The newer README examples use async def start. For older Scrapy versions, use the older start_requests entry point instead, while observing the package’s stated minimum-version requirements.

The callback can be asynchronous, as above, or written in the style appropriate to your Scrapy project. The returned object remains a Scrapy response, so ordinary Scrapy selectors and item-yielding patterns still apply.

Accessing the Playwright page and applying page methods

Retain a page only when you need it

If callback logic must interact with the live Playwright page, set playwright_include_page=True on the request. The page is then available as response.meta['playwright_page']. A retained page is a resource you own: close it when asynchronous work that needs it is complete.

async def parse(self, response):
    page = response.meta["playwright_page"]
    try:
        current_url = page.url
        yield {"url": current_url}
    finally:
        await page.close()

Do not retain a page simply to run a documented PageMethod operation. The integration can apply page methods without returning the page object to the callback.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Keep the response-versus-page distinction clear

Use the Scrapy response for normal extraction after navigation and rendering. Use the included Playwright page when you specifically need live page operations or page state. This distinction helps avoid keeping browser pages open unnecessarily while processing a crawl.

Contexts, sessions, and concurrency

A browser context represents a browser session boundary, useful when requests need separate state or session settings. Use playwright_context to select a named context for a request, and playwright_context_kwargs when the integration should create a new context with particular options. The PLAYWRIGHT_CONTEXTS setting configures contexts at startup; PLAYWRIGHT_MAX_CONTEXTS limits how many contexts can be active simultaneously.

Persistent contexts use a user_data_dir to hold profile data. Persistent profile paths should be assigned deliberately, especially if both HTTP and HTTPS handlers are registered: the handlers can contend over the same profile. For a crawl that hangs or runs out of resources, inspect context names, persistent-profile paths, and the configured maximum context count before changing unrelated parsing code.

More simultaneous browser contexts can increase resource demand. The project documentation establishes the relevant context controls, but does not provide a universal safe concurrency value; choose and verify a limit against your workload and runtime environment.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Browser engines, launch options, and remote browsers

PLAYWRIGHT_BROWSER_TYPE selects Chromium, Firefox, or WebKit. PLAYWRIGHT_LAUNCH_OPTIONS passes browser launch arguments, including headless mode and a timeout. The integration also documents remote-browser connection settings:

  • PLAYWRIGHT_CDP_URL connects over the Chrome DevTools Protocol and requires Chromium.
  • PLAYWRIGHT_CONNECT_URL connects to a browser using Playwright’s connection mechanism.
  • The README says the CDP and connect URL settings cannot be used together.

Start with the default local-browser setup unless you have a concrete need for a different engine, launch configuration, or remote browser. If you change browser configuration, make sure it matches the target and the browser service you actually operate.

Other useful controls

The integration supports more than basic rendering. Use these capabilities when a crawl requirement calls for them, and consult the README for exact metadata keys, settings, and usage details:

  • Request-header processing and custom browser providers.
  • Page methods, downloads, and screenshots.
  • Access to response-related data through Playwright metadata.

For example, a page method may be appropriate when the page needs a specific browser operation before extraction, while a screenshot capability is relevant when the output itself is visual. Avoid adding advanced controls to a minimal spider until the simpler request, handler, and selector path is working.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Troubleshooting empty responses, startup errors, and hangs

The spider returns empty or unrendered HTML

  • Confirm that the particular request has meta={"playwright": True}. Enabling the handler in settings alone does not opt every request into browser rendering.
  • Check that the HTTPS handler is configured and the asyncio reactor setting matches the documented configuration.
  • Confirm the target actually requires browser rendering. If the page’s data comes from an API request you can reproduce directly, use that route or determine whether the browser is waiting for content that appears later.

Playwright cannot find a browser executable

Install browser binaries with playwright install in the environment used by the project. If the deployment environment differs from the development environment, install the required browser binaries there too. To install only selected engines, the project documents commands such as playwright install firefox chromium.

Configuration or compatibility errors at startup

Check the documented minimum versions: Python 3.10, Scrapy 2.7, and Playwright 1.40. Also check spelling and placement of DOWNLOAD_HANDLERS and TWISTED_REACTOR in the settings loaded by the spider. The project README is the authority for current package requirements and configuration.

A callback hangs or browser resources accumulate

  • If you set playwright_include_page=True, close the retained page once page-dependent work is done.
  • Check whether context names are being reused as intended and whether the context limit is constraining the crawl.
  • Review persistent profile paths and avoid having both handlers attempt to own the same profile.
  • Consider whether browser rendering is necessary for each affected request; direct HTTP requests avoid browser work when they can provide the needed data.

Cost, overhead, and reliability considerations

A browser must be launched or reached, a page navigated, and relevant browser work completed before extraction. That makes browser rendering operationally heavier than a direct Scrapy request when both can retrieve the same information. The official guidance supports preferring reproducible underlying data requests where practical, but it does not specify a numeric performance advantage, benchmark, or success rate.

For reliability, make the crawl’s browser dependencies explicit: install binaries in the runtime environment, configure the correct handler and reactor, opt in only the requests that need rendering, and release pages retained for callback work. Treat concurrency and persistent profiles as operational choices to validate under your own workload rather than assuming a universal setting.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Or skip the browser setup

If your goal is a screenshot rather than a Scrapy crawl, ScreenshotNeo is a website screenshot API and MCP server for developers. Its GET endpoint returns an image or PDF for a URL; the API can also be used by an AI agent through its MCP server. The following cURL example requests a WebP screenshot. See the ScreenshotNeo documentation for request parameters and response details.

curl -G "https://api.screenshotneo.com/v1/shot" 
  -d access_key=YOUR_API_KEY 
  --data-urlencode url=https://stripe.com 
  -o shot.webp

ScreenshotNeo accepts cookie or consent banners as a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each of those steps can be turned off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and the response identifies page verdict and billing status in headers. Its MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients.

The Free plan includes 1,000 screenshots per month with no card required. Paid plans start at $5 for 3,000 shots, and every feature is available on every plan. Sign up for ScreenshotNeo’s free plan to try it without a card.

Frequently Asked Questions

Can I use scrapy-playwright for every request in a Scrapy spider?

Yes, but it is opt-in per request: mark requests that need browser rendering with the playwright metadata key. Unmarked requests use Scrapy’s regular downloader.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Do I need to include the Playwright page in every response?

No. Set playwright_include_page=True only when callback logic needs the live page; page methods can be applied without retaining it.

Which browser engines can scrapy-playwright use?

The documented browser types are Chromium, Firefox, and WebKit. Remote CDP connections require Chromium.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.