Scrapy Splash is not a browser library you install into Scrapy. It is a Scrapy client that sends requests to a separate Splash HTTP rendering service, usually running in Docker. A working integration therefore has two parts: install scrapy-splash in a Python 3.10-or-newer environment, start Splash, and configure Scrapy’s middleware, argument deduplication, and request fingerprinter. Use render.html or render.json for simple pages; use /execute or /run when Lua must control navigation, JavaScript, cookies, or the returned data.
How the Scrapy Splash architecture works
Scrapy remains responsible for scheduling, parsing, retries, pipelines, and item storage. Splash is a separate HTTP service that loads a URL in its WebKit-based rendering engine and returns HTML, JSON, a Lua result, or another requested representation. The scrapy-splash package supplies request classes and middleware that connect the two.
- Scrapy client: your spider and its normal downloader.
- Splash server: a long-running service listening on an address such as
http://localhost:8050. - Rendering request: a
SplashRequestcontaining the target URL and renderer arguments. - Response: the rendered result, with Splash metadata available to the spider.
This separation matters operationally: installing the Python package does not start a renderer, and running a container without configuring Scrapy does not make ordinary Request objects execute JavaScript.
Prerequisites and installation
Use a dedicated Python environment
Current Scrapy installation guidance requires Python 3.10 or newer (CPython or PyPy). Create an isolated environment so Scrapy, scrapy-splash, and project dependencies do not conflict with system packages.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
#1 Best Overall
python3 -m venv .venv
. .venv/bin/activate
python -m pip install --upgrade pip
pip install scrapy scrapy-splash
Start Splash with Docker
The documented quick start publishes Splash’s port 8050 on the host:
docker run -p 8050:8050 scrapinghub/splash
For diagnostics, add verbose logging:
docker run -p 8050:8050 scrapinghub/splash -v2
Keep the container running while the spider operates. In a remote deployment, replace localhost with the reachable service address and restrict the port with your network controls; do not expose an unauthenticated renderer to the public internet.
Configure Scrapy correctly
Set the service address and the integration components in your project settings. The middleware priorities below are the documented arrangement; changing them casually can break cookies, compression, or request processing.
SPLASH_URL = 'http://localhost:8050'
DOWNLOADER_MIDDLEWARES = {
'scrapy_splash.SplashCookiesMiddleware': 723,
'scrapy_splash.SplashMiddleware': 725,
'scrapy.downloadermiddlewares.httpcompression.HttpCompressionMiddleware': 810,
}
SPIDER_MIDDLEWARES = {
'scrapy_splash.SplashDeduplicateArgsMiddleware': 100,
}
REQUEST_FINGERPRINTER_CLASS = 'scrapy_splash.SplashRequestFingerprinter'
SPLASH_URL tells the client where to send rendering calls. SplashCookiesMiddleware keeps Splash cookie handling compatible with Scrapy’s request flow. SplashMiddleware converts a SplashRequest into the correct HTTP call. The compression priority prevents response decompression from occurring in the wrong stage. Argument deduplication avoids scheduling equivalent rendering arguments repeatedly, while SplashRequestFingerprinter ensures Splash-specific arguments participate in duplicate filtering.
Free tools Windows power users keep installed
One-click scans. No signup required.
Verify the service before debugging a spider
Open http://localhost:8050 in a browser or send a simple HTTP request to confirm the container is reachable. If the connection is refused, fix Docker, port mapping, or the host name before changing Lua or Scrapy settings.
Choose the right Splash endpoint
| Endpoint | Use it for | What you provide |
|---|---|---|
render.html |
Straightforward JavaScript rendering where the final page HTML is enough | URL plus renderer arguments such as wait time |
render.json |
HTML and structured response metadata in one result | URL and JSON renderer arguments |
/execute |
Custom navigation, JavaScript evaluation, cookies, waits, and computed values | A Lua source string containing main(splash) |
/run |
Running a Lua script supplied or managed as a named resource | Lua script and its arguments |
The Splash API documentation describes execute and run as the most versatile endpoints because they can execute arbitrary Lua rendering scripts. Start with render.html when no interaction is needed; move to execute when the page requires a click, a custom wait condition, a cookie round trip, or a non-HTML result.
Write and run a Lua script
The basic pattern
A Splash Lua script defines main(splash). Navigate with splash:go, wait or evaluate JavaScript if necessary, then return a string, value, or table.
lua_source = '''
function main(splash)
assert(splash:go(splash.args.url))
return splash:evaljs("document.title")
end
'''
yield SplashRequest(
url='https://example.com',
endpoint='execute',
args={'lua_source': lua_source},
callback=self.parse_title,
)
def parse_title(self, response):
self.logger.info("Rendered title: %s", response.text)
Here splash.args.url is populated from the request’s URL. assert turns a failed navigation into a Lua traceback instead of silently returning an unusable page. splash:evaljs evaluates JavaScript in the loaded document and returns its result.
Quick wins for a faster PC:
Clear out junk files and repair common Windows errorsFree Scan →Scan for outdated or missing drivers - takes under a minuteDriver Scan →Repair Windows errors before they cause bigger problemsFix Now →Return rendered HTML
lua_source = '''
function main(splash)
assert(splash:go(splash.args.url))
splash:wait(2)
return splash:html()
end
'''
yield SplashRequest(
url='https://example.com/app',
endpoint='execute',
args={'lua_source': lua_source},
callback=self.parse_page,
)
def parse_page(self, response):
title = response.css('title::text').get()
self.logger.info("Title: %s", title)
The wait is deliberately explicit. Replace it with a condition based on page state when possible; a fixed delay increases latency and can still be too short for a slow application.
Return a custom table
lua_source = '''
function main(splash)
assert(splash:go(splash.args.url))
return {
title = splash:evaljs("document.title"),
html = splash:html()
}
end
'''
yield SplashRequest(
url='https://example.com',
endpoint='execute',
args={'lua_source': lua_source},
callback=self.parse_result,
)
def parse_result(self, response):
data = response.data
title = data.get('title')
html = data.get('html')
Returning a table is useful when the spider needs both the page and computed values. Confirm the exact response shape in your project logs before writing parsing code around it.
Rank #3
Cookies and sessions
Splash is stateless per request. A later request does not automatically inherit the browser state of an earlier one. For a session, pass incoming cookies into Lua, navigate, and return the updated cookie jar:
lua_source = '''
function main(splash)
splash:init_cookies(splash.args.cookies)
assert(splash:go(splash.args.url))
return {
cookies = splash:get_cookies(),
html = splash:html()
}
end
'''
yield SplashRequest(
url='https://example.com/account',
endpoint='execute',
args={'lua_source': lua_source, 'cookies': []},
session_id='account-session',
callback=self.parse_session,
)
def parse_session(self, response):
cookies = response.data.get('cookies', [])
html = response.data.get('html', '')
On subsequent requests, provide the cookies returned by the previous response. The session_id on the Scrapy side groups related Splash requests, but it does not remove the need to pass and update cookie data in Lua.
Do these 3 things before closing this tab:
1Clear out junk files and repair common Windows errors2Fix the driver behind crashes, sound loss and screen glitches3Repair Windows errors before they cause bigger problemsPOST requests and cached arguments
Splash 1.8 or newer is required for the http_method and body POST arguments. With /execute, the Lua script must pass those values to splash:go; merely adding them to a Scrapy request is not enough.
lua_source = '''
function main(splash)
assert(splash:go(
splash.args.url,
splash.args.http_method,
splash.args.body
))
return splash:html()
end
'''
yield SplashRequest(
url='https://example.com/form',
endpoint='execute',
args={
'lua_source': lua_source,
'http_method': 'POST',
'body': 'field=value',
},
callback=self.parse_page,
)
Splash 2.1 or newer supports server-side caching of large static arguments such as lua_source. That can reduce repeated request traffic and disk-queue duplication when the same script is used across many URLs. It does not cache the target site’s changing content.
Compatibility: when Splash is a poor fit
WebKit is the central limitation
The Splash FAQ warns that target sites may be incompatible with its WebKit engine. Modern sites can depend on browser APIs, JavaScript syntax, security behavior, or rendering features that an older WebKit implementation does not provide. A page that works in a current Chrome-based browser can therefore fail, render partially, or remain blank in Splash.
Scrapy’s dynamic-content guidance positions Splash for JavaScript-rendered pages, but recommends a modern headless browser when the job requires on-the-fly DOM interaction or multiple windows. Use the simplest renderer that supports the site rather than forcing every workflow through Lua.
Match project versions deliberately
- Use Python 3.10 or newer for current Scrapy installation guidance.
- Use Splash 1.8 or newer for POST handling with
http_methodandbody. - Use Splash 2.1 or newer when you need cached large arguments.
- Check Scrapy release notes when upgrading: backward incompatibilities are called out there, and deprecated features are generally retained for at least one year.
Troubleshooting checklist
Connection refused or timeout
- Cause: the container is stopped, port 8050 is not published, or
SPLASH_URLpoints to the wrong host. - Fix: run
docker ps, confirm-p 8050:8050, test the service directly, and use the container or service DNS name from inside a containerized Scrapy deployment.
Ordinary requests work but JavaScript is missing
- Cause: the spider uses
scrapy.Requestinstead ofSplashRequest, or the Splash middleware is absent. - Fix: use a Splash endpoint explicitly and verify
DOWNLOADER_MIDDLEWARES,SPIDER_MIDDLEWARES, andSPLASH_URL.
Duplicate filtering behaves unexpectedly
- Cause: Splash arguments are not included in request fingerprints or equivalent arguments are not deduplicated.
- Fix: set
REQUEST_FINGERPRINTER_CLASStoscrapy_splash.SplashRequestFingerprinterand enableSplashDeduplicateArgsMiddleware.
Lua traceback or empty output
- Cause: navigation failed, a selector is not ready, JavaScript returned an unexpected type, or the script omitted a return value.
- Fix: run Splash with
-v2, inspect the complete request and endpoint, log the Lua traceback, assertsplash:go, and return a known diagnostic value before adding more interactions.
The page is blank or broken only in Splash
- Cause: target-site incompatibility with Splash’s WebKit engine, anti-bot behavior, or an unsupported browser feature.
- Fix: compare the same URL in a current headless browser, simplify the page interaction, or migrate that workflow to a modern browser renderer.
POST data is ignored
- Cause: Splash is older than 1.8, or the Lua script does not pass
http_methodandbodytosplash:go. - Fix: upgrade Splash and use the explicit Lua call shown above.
Performance, reliability, and operating cost
Rendering is substantially heavier than downloading raw HTML. Limit concurrency to what the Splash host can sustain, reuse a single script rather than sending large Lua text unnecessarily, and prefer condition-based waits over long fixed sleeps. Cache only content that is safe to reuse; stale rendered pages can be more damaging than slower fresh requests.
Record the target URL, endpoint, renderer arguments, Splash version, and Lua traceback for failed jobs. Treat authentication cookies as secrets, keep them out of logs, and do not share a session identifier across unrelated accounts. Because Splash is a separate service, monitor both Scrapy’s queue and the renderer’s CPU, memory, and disk queue.
Or skip the browser setup
If you need a screenshot or PDF rather than a self-managed Scrapy renderer, ScreenshotNeo provides a website screenshot API and MCP server. A single GET request returns PNG, JPEG, WebP, or PDF. It accepts cookie and consent banners as a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each cleanup step can be disabled. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and the response identifies the page verdict and billing status in X-Page-Verdict and X-Billed headers.
cURL:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python:
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
See the ScreenshotNeo documentation for all options. It supports full-page captures with lazy images loaded, CSS-selector element captures, dark mode, 12 device presets and arbitrary viewports, retina scale, PDF paper sizes and page ranges, HTML/CSS-to-image, custom CSS and JavaScript, clicks, selector or network-idle waits, request blocking, headers, cookies, user agents, authorization, timezone, geolocation, transparent backgrounds, resizing, chosen cache TTLs, signed image links, asynchronous webhooks, bulk capture of up to 100 URLs per call, a usage API, and an OpenAPI specification. An MCP server exposes take_screenshot, get_page_info, and capture_pdf to Claude, Cursor, and other MCP clients.
Windows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallOutdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchThe Free plan includes 1,000 screenshots per month with no card. Paid plans start at $5 for 3,000 shots; yearly billing gives two months free, and every feature is available on every plan. Create a free ScreenshotNeo account.
Best Value
Frequently Asked Questions
Does installing scrapy-splash install Splash itself?
No. The package installs the Scrapy-side client and middleware; Splash must run separately, commonly in the official Docker container.
Can Splash preserve login state between requests?
Not automatically. Splash is stateless per request; pass cookies into Lua, return updated cookies, and carry them forward with an appropriate Scrapy session.
Which endpoint should I start with for a JavaScript page?
Use render.html when you only need rendered HTML. Choose execute or run when you need custom Lua navigation, JavaScript evaluation, cookies, interactions, or a custom result.
When should I replace Splash with a modern browser?
Replace it when the target depends on browser features unavailable in Splash’s WebKit engine, requires multiple windows, or needs complex live DOM interaction.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




