October DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsClean PCRecommendedOne scan can reveal what keeps slowing WindowsLook for cleanup and repair opportunities.Run ScanOctober DealsAmazon USDeal season is back - check today's better picksAmazon US: current deals, useful picks and tech finds.See Picks×
Skip to content
Docker

Scrapy Splash Guide: Setup, Lua, and Compatibility

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Scrapy Splash is not a browser library you install into Scrapy. It is a Scrapy client that sends requests to a separate Splash HTTP rendering service, usually running in Docker. A working integration therefore has two parts: install scrapy-splash in a Python 3.10-or-newer environment, start Splash, and configure Scrapy’s middleware, argument deduplication, and request fingerprinter. Use render.html or render.json for simple pages; use /execute or /run when Lua must control navigation, JavaScript, cookies, or the returned data.

How the Scrapy Splash architecture works

Scrapy remains responsible for scheduling, parsing, retries, pipelines, and item storage. Splash is a separate HTTP service that loads a URL in its WebKit-based rendering engine and returns HTML, JSON, a Lua result, or another requested representation. The scrapy-splash package supplies request classes and middleware that connect the two.

  • Scrapy client: your spider and its normal downloader.
  • Splash server: a long-running service listening on an address such as http://localhost:8050.
  • Rendering request: a SplashRequest containing the target URL and renderer arguments.
  • Response: the rendered result, with Splash metadata available to the spider.

This separation matters operationally: installing the Python package does not start a renderer, and running a container without configuring Scrapy does not make ordinary Request objects execute JavaScript.

Prerequisites and installation

Use a dedicated Python environment

Current Scrapy installation guidance requires Python 3.10 or newer (CPython or PyPy). Create an isolated environment so Scrapy, scrapy-splash, and project dependencies do not conflict with system packages.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
python3 -m venv .venv
. .venv/bin/activate
python -m pip install --upgrade pip
pip install scrapy scrapy-splash

Start Splash with Docker

The documented quick start publishes Splash’s port 8050 on the host:

docker run -p 8050:8050 scrapinghub/splash

For diagnostics, add verbose logging:

docker run -p 8050:8050 scrapinghub/splash -v2

Keep the container running while the spider operates. In a remote deployment, replace localhost with the reachable service address and restrict the port with your network controls; do not expose an unauthenticated renderer to the public internet.

Configure Scrapy correctly

Set the service address and the integration components in your project settings. The middleware priorities below are the documented arrangement; changing them casually can break cookies, compression, or request processing.

SPLASH_URL = 'http://localhost:8050'

DOWNLOADER_MIDDLEWARES = {
    'scrapy_splash.SplashCookiesMiddleware': 723,
    'scrapy_splash.SplashMiddleware': 725,
    'scrapy.downloadermiddlewares.httpcompression.HttpCompressionMiddleware': 810,
}

SPIDER_MIDDLEWARES = {
    'scrapy_splash.SplashDeduplicateArgsMiddleware': 100,
}

REQUEST_FINGERPRINTER_CLASS = 'scrapy_splash.SplashRequestFingerprinter'

SPLASH_URL tells the client where to send rendering calls. SplashCookiesMiddleware keeps Splash cookie handling compatible with Scrapy’s request flow. SplashMiddleware converts a SplashRequest into the correct HTTP call. The compression priority prevents response decompression from occurring in the wrong stage. Argument deduplication avoids scheduling equivalent rendering arguments repeatedly, while SplashRequestFingerprinter ensures Splash-specific arguments participate in duplicate filtering.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Verify the service before debugging a spider

Open http://localhost:8050 in a browser or send a simple HTTP request to confirm the container is reachable. If the connection is refused, fix Docker, port mapping, or the host name before changing Lua or Scrapy settings.

Choose the right Splash endpoint

Endpoint Use it for What you provide
render.html Straightforward JavaScript rendering where the final page HTML is enough URL plus renderer arguments such as wait time
render.json HTML and structured response metadata in one result URL and JSON renderer arguments
/execute Custom navigation, JavaScript evaluation, cookies, waits, and computed values A Lua source string containing main(splash)
/run Running a Lua script supplied or managed as a named resource Lua script and its arguments

The Splash API documentation describes execute and run as the most versatile endpoints because they can execute arbitrary Lua rendering scripts. Start with render.html when no interaction is needed; move to execute when the page requires a click, a custom wait condition, a cookie round trip, or a non-HTML result.

Write and run a Lua script

The basic pattern

A Splash Lua script defines main(splash). Navigate with splash:go, wait or evaluate JavaScript if necessary, then return a string, value, or table.

lua_source = '''
function main(splash)
    assert(splash:go(splash.args.url))
    return splash:evaljs("document.title")
end
'''

yield SplashRequest(
    url='https://example.com',
    endpoint='execute',
    args={'lua_source': lua_source},
    callback=self.parse_title,
)

def parse_title(self, response):
    self.logger.info("Rendered title: %s", response.text)

Here splash.args.url is populated from the request’s URL. assert turns a failed navigation into a Lua traceback instead of silently returning an unusable page. splash:evaljs evaluates JavaScript in the loaded document and returns its result.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Return rendered HTML

lua_source = '''
function main(splash)
    assert(splash:go(splash.args.url))
    splash:wait(2)
    return splash:html()
end
'''

yield SplashRequest(
    url='https://example.com/app',
    endpoint='execute',
    args={'lua_source': lua_source},
    callback=self.parse_page,
)

def parse_page(self, response):
    title = response.css('title::text').get()
    self.logger.info("Title: %s", title)

The wait is deliberately explicit. Replace it with a condition based on page state when possible; a fixed delay increases latency and can still be too short for a slow application.

Return a custom table

lua_source = '''
function main(splash)
    assert(splash:go(splash.args.url))
    return {
        title = splash:evaljs("document.title"),
        html = splash:html()
    }
end
'''

yield SplashRequest(
    url='https://example.com',
    endpoint='execute',
    args={'lua_source': lua_source},
    callback=self.parse_result,
)

def parse_result(self, response):
    data = response.data
    title = data.get('title')
    html = data.get('html')

Returning a table is useful when the spider needs both the page and computed values. Confirm the exact response shape in your project logs before writing parsing code around it.

Cookies and sessions

Splash is stateless per request. A later request does not automatically inherit the browser state of an earlier one. For a session, pass incoming cookies into Lua, navigate, and return the updated cookie jar:

lua_source = '''
function main(splash)
    splash:init_cookies(splash.args.cookies)
    assert(splash:go(splash.args.url))
    return {
        cookies = splash:get_cookies(),
        html = splash:html()
    }
end
'''

yield SplashRequest(
    url='https://example.com/account',
    endpoint='execute',
    args={'lua_source': lua_source, 'cookies': []},
    session_id='account-session',
    callback=self.parse_session,
)

def parse_session(self, response):
    cookies = response.data.get('cookies', [])
    html = response.data.get('html', '')

On subsequent requests, provide the cookies returned by the previous response. The session_id on the Scrapy side groups related Splash requests, but it does not remove the need to pass and update cookie data in Lua.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

POST requests and cached arguments

Splash 1.8 or newer is required for the http_method and body POST arguments. With /execute, the Lua script must pass those values to splash:go; merely adding them to a Scrapy request is not enough.

lua_source = '''
function main(splash)
    assert(splash:go(
        splash.args.url,
        splash.args.http_method,
        splash.args.body
    ))
    return splash:html()
end
'''

yield SplashRequest(
    url='https://example.com/form',
    endpoint='execute',
    args={
        'lua_source': lua_source,
        'http_method': 'POST',
        'body': 'field=value',
    },
    callback=self.parse_page,
)

Splash 2.1 or newer supports server-side caching of large static arguments such as lua_source. That can reduce repeated request traffic and disk-queue duplication when the same script is used across many URLs. It does not cache the target site’s changing content.

Compatibility: when Splash is a poor fit

WebKit is the central limitation

The Splash FAQ warns that target sites may be incompatible with its WebKit engine. Modern sites can depend on browser APIs, JavaScript syntax, security behavior, or rendering features that an older WebKit implementation does not provide. A page that works in a current Chrome-based browser can therefore fail, render partially, or remain blank in Splash.

Scrapy’s dynamic-content guidance positions Splash for JavaScript-rendered pages, but recommends a modern headless browser when the job requires on-the-fly DOM interaction or multiple windows. Use the simplest renderer that supports the site rather than forcing every workflow through Lua.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Match project versions deliberately

  • Use Python 3.10 or newer for current Scrapy installation guidance.
  • Use Splash 1.8 or newer for POST handling with http_method and body.
  • Use Splash 2.1 or newer when you need cached large arguments.
  • Check Scrapy release notes when upgrading: backward incompatibilities are called out there, and deprecated features are generally retained for at least one year.

Troubleshooting checklist

Connection refused or timeout

  • Cause: the container is stopped, port 8050 is not published, or SPLASH_URL points to the wrong host.
  • Fix: run docker ps, confirm -p 8050:8050, test the service directly, and use the container or service DNS name from inside a containerized Scrapy deployment.

Ordinary requests work but JavaScript is missing

  • Cause: the spider uses scrapy.Request instead of SplashRequest, or the Splash middleware is absent.
  • Fix: use a Splash endpoint explicitly and verify DOWNLOADER_MIDDLEWARES, SPIDER_MIDDLEWARES, and SPLASH_URL.

Duplicate filtering behaves unexpectedly

  • Cause: Splash arguments are not included in request fingerprints or equivalent arguments are not deduplicated.
  • Fix: set REQUEST_FINGERPRINTER_CLASS to scrapy_splash.SplashRequestFingerprinter and enable SplashDeduplicateArgsMiddleware.

Lua traceback or empty output

  • Cause: navigation failed, a selector is not ready, JavaScript returned an unexpected type, or the script omitted a return value.
  • Fix: run Splash with -v2, inspect the complete request and endpoint, log the Lua traceback, assert splash:go, and return a known diagnostic value before adding more interactions.

The page is blank or broken only in Splash

  • Cause: target-site incompatibility with Splash’s WebKit engine, anti-bot behavior, or an unsupported browser feature.
  • Fix: compare the same URL in a current headless browser, simplify the page interaction, or migrate that workflow to a modern browser renderer.

POST data is ignored

  • Cause: Splash is older than 1.8, or the Lua script does not pass http_method and body to splash:go.
  • Fix: upgrade Splash and use the explicit Lua call shown above.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Performance, reliability, and operating cost

Rendering is substantially heavier than downloading raw HTML. Limit concurrency to what the Splash host can sustain, reuse a single script rather than sending large Lua text unnecessarily, and prefer condition-based waits over long fixed sleeps. Cache only content that is safe to reuse; stale rendered pages can be more damaging than slower fresh requests.

Record the target URL, endpoint, renderer arguments, Splash version, and Lua traceback for failed jobs. Treat authentication cookies as secrets, keep them out of logs, and do not share a session identifier across unrelated accounts. Because Splash is a separate service, monitor both Scrapy’s queue and the renderer’s CPU, memory, and disk queue.

Or skip the browser setup

If you need a screenshot or PDF rather than a self-managed Scrapy renderer, ScreenshotNeo provides a website screenshot API and MCP server. A single GET request returns PNG, JPEG, WebP, or PDF. It accepts cookie and consent banners as a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each cleanup step can be disabled. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and the response identifies the page verdict and billing status in X-Page-Verdict and X-Billed headers.

cURL:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

See the ScreenshotNeo documentation for all options. It supports full-page captures with lazy images loaded, CSS-selector element captures, dark mode, 12 device presets and arbitrary viewports, retina scale, PDF paper sizes and page ranges, HTML/CSS-to-image, custom CSS and JavaScript, clicks, selector or network-idle waits, request blocking, headers, cookies, user agents, authorization, timezone, geolocation, transparent backgrounds, resizing, chosen cache TTLs, signed image links, asynchronous webhooks, bulk capture of up to 100 URLs per call, a usage API, and an OpenAPI specification. An MCP server exposes take_screenshot, get_page_info, and capture_pdf to Claude, Cursor, and other MCP clients.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The Free plan includes 1,000 screenshots per month with no card. Paid plans start at $5 for 3,000 shots; yearly billing gives two months free, and every feature is available on every plan. Create a free ScreenshotNeo account.

Frequently Asked Questions

Does installing scrapy-splash install Splash itself?

No. The package installs the Scrapy-side client and middleware; Splash must run separately, commonly in the official Docker container.

Can Splash preserve login state between requests?

Not automatically. Splash is stateless per request; pass cookies into Lua, return updated cookies, and carry them forward with an appropriate Scrapy session.

Which endpoint should I start with for a JavaScript page?

Use render.html when you only need rendered HTML. Choose execute or run when you need custom Lua navigation, JavaScript evaluation, cookies, interactions, or a custom result.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

When should I replace Splash with a modern browser?

Replace it when the target depends on browser features unavailable in Splash’s WebKit engine, requires multiple windows, or needs complex live DOM interaction.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Read next

Recommended PC Tool
Recommended PC Tool
Outdated Drivers Are Slowing You DownFree scan - exact matches
PC Slower Than It Used to Be?Free scan - under a minute

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.