Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

Short answer: Scrapy does not execute page JavaScript by itself. First inspect the response Scrapy receives and identify the request or embedded data that JavaScript uses. Reproduce that request with Scrapy whenever possible; it is usually more complete and lighter than rendering a browser. Use a headless browser through scrapy-playwright when the data cannot be reproduced reliably or when you need a browser-only result such as a screenshot.

What “JavaScript-rendered” means in Scrapy

A browser can display products, prices or comments that are absent from the HTML response downloaded by Scrapy. The browser received an initial document, then JavaScript made additional requests, transformed embedded state, or inserted elements into the DOM. Scrapy selectors only see the response body unless you arrange another way to obtain that data.

That does not automatically mean you need Selenium or Playwright. The most reliable workflow is to discover the data source and request it directly. Rendering is the fallback for genuinely browser-dependent behavior.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use this decision path

  1. Inspect the raw response. Run scrapy fetch --nolog https://example.com/page from your project and save or inspect the returned body. Search for the value you want, JSON-looking state, API URLs and <script> elements.
  2. Inspect network requests in a normal browser. Open developer tools, select the Network panel, reload the page and filter by Fetch/XHR. Find the request whose response contains the records, then note its URL, method, query parameters, request body, headers and pagination fields.
  3. Reproduce that request in a Scrapy callback. Request reproduction is Scrapy’s preferred approach when feasible. It normally returns structured data directly and avoids the parsing and transfer overhead of a full browser page.
  4. Parse embedded state without rendering. Data may be in a JSON script element, an external JavaScript file or a JavaScript object that is not strict JSON.
  5. Render with a browser only when necessary. Choose scrapy-playwright when the site requires browser execution, complex interaction, or a screenshot.

Approach 1: reproduce the data request

Suppose the page calls https://example.com/api/products?page=1 and returns JSON. Create a Scrapy spider that requests the endpoint instead of scraping the post-rendered page:

import scrapy

class ProductsSpider(scrapy.Spider):
    name = "products"

    def start_requests(self):
        yield scrapy.Request(
            "https://example.com/api/products?page=1",
            headers={"Accept": "application/json"},
            callback=self.parse_products,
        )

    def parse_products(self, response):
        payload = response.json()
        for product in payload.get("items", []):
            yield {
                "id": product.get("id"),
                "name": product.get("name"),
                "price": product.get("price"),
            }

        next_url = payload.get("next")
        if next_url:
            yield response.follow(next_url, callback=self.parse_products)

Copy the actual request details from your browser rather than guessing them. POST endpoints require the same JSON or form body. If the server checks a session, pass the necessary cookies or use a Scrapy session flow that first visits the page and then sends the API request. Keep only headers that are required; a copied browser request often contains irrelevant or short-lived fields.

When the endpoint needs a token

Look for a token in the initial HTML, a cookie, a script configuration object or a preceding request. Request the page first, extract the token, and then issue the API request. Do not hard-code a token that changes between sessions.

class ProductsSpider(scrapy.Spider):
    name = "products"

    def start_requests(self):
        yield scrapy.Request("https://example.com/catalog", callback=self.open_catalog)

    def open_catalog(self, response):
        token = response.css("meta[name='api-token']::attr(content)").get()
        yield scrapy.Request(
            "https://example.com/api/products",
            method="POST",
            headers={
                "Accept": "application/json",
                "Content-Type": "application/json",
                "X-Api-Token": token or "",
            },
            body='{"page":1}',
            callback=self.parse_products,
        )

    def parse_products(self, response):
        for item in response.json().get("items", []):
            yield item

Validate the response status, content type and schema before extracting fields. A successful HTTP status can still contain an authentication error or an HTML challenge page.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Approach 2: parse JavaScript already delivered in the response

JSON in a script element

Many frameworks place a serialized state object in a script tag. Select its text and decode it as JSON:

import json
import scrapy

class StateSpider(scrapy.Spider):
    name = "state"
    start_urls = ["https://example.com/page"]

    def parse(self, response):
        raw = response.css("script#__DATA__::text").get()
        if not raw:
            self.logger.warning("State script was not found")
            return
        state = json.loads(raw)
        for row in state.get("products", []):
            yield {"name": row.get("name")}

The selector and script identifier are site-specific. Inspect the saved response rather than assuming a framework’s conventional name.

JavaScript objects that are not strict JSON

Single quotes, trailing commas, comments, unquoted keys and expressions make a JavaScript object invalid JSON. A JavaScript-object parser such as chompjs can handle JSON-like values. For code where selectors are more convenient, js2xml can convert JavaScript into XML that Scrapy selectors can query. Use the smallest parsing scope possible and treat untrusted text as data, not executable code.

External JavaScript files

If the needed value is in an external file, request it and inspect response.text. Minified bundles can contain application defaults, but runtime values are often fetched separately; do not assume a bundle is the authoritative data source.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Approach 3: render with scrapy-playwright

Use a browser when reproducing requests is too difficult, the page depends on browser APIs, or your task itself requires rendered behavior. Scrapy’s guidance recommends scrapy-playwright for better integration. Calling Playwright directly from a standalone script can bypass much of Scrapy’s middleware and duplicate filtering.

Install and configure

The current Scrapy 2.19 installation guidance requires Python 3.10 or later. Create a virtual environment, install Scrapy and the integration package, and install a Playwright browser:

python -m venv .venv
# macOS/Linux
. .venv/bin/activate
# Windows PowerShell: .venvScriptsActivate.ps1
pip install scrapy scrapy-playwright
playwright install chromium

In settings.py, enable the download handler and middleware:

DOWNLOAD_HANDLERS = {
    "http": "scrapy_playwright.handler.ScrapyPlaywrightDownloadHandler",
    "https": "scrapy_playwright.handler.ScrapyPlaywrightDownloadHandler",
}

TWISTED_REACTOR = "twisted.internet.asyncioreactor.AsyncioSelectorReactor"

PLAYWRIGHT_BROWSER_TYPE = "chromium"
PLAYWRIGHT_DEFAULT_NAVIGATION_TIMEOUT = 30_000

Keep the browser context bounded for a crawl. Set concurrency and page limits appropriate to the machine, and close pages you create manually.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Request a rendered page

import scrapy

class RenderedSpider(scrapy.Spider):
    name = "rendered"

    def start_requests(self):
        yield scrapy.Request(
            "https://example.com/catalog",
            meta={
                "playwright": True,
                "playwright_page_methods": [
                    {"method": "wait_for_selector", "args": [".product-card"]}
                ],
            },
            callback=self.parse,
        )

    def parse(self, response):
        for card in response.css(".product-card"):
            yield {
                "name": card.css(".name::text").get(),
                "price": card.css(".price::text").get(),
            }

The integration renders the page and returns the resulting content to the callback. Replace the selector and wait condition with one that represents the data you need. A fixed delay can help with pages that have no stable selector, but waiting for a meaningful condition is generally more deterministic.

Clicks, scrolling and page methods

For a “Load more” control, use a Playwright page method to click it, then wait for the next records. For infinite scrolling, scroll in bounded steps and stop when a known count or end marker appears. Full-page lazy images may require scrolling before extraction. Avoid unbounded loops: record how many pages or items have been processed and stop on a missing next button or unchanged result set.

Intercepting responses instead of scraping the DOM

Even in a browser session, the cleanest result may be a JSON response. Capture or inspect that response and parse its payload rather than extracting text from a deeply nested rendered DOM. This keeps the browser for authentication or interaction while preserving structured data extraction.

Choosing between requests, parsing and rendering

Situation Preferred method Why
Data appears in HTML Normal Scrapy selectors No JavaScript work is required.
Data comes from a reproducible API request Scrapy request reproduction Structured results with less parsing and network transfer.
State is embedded in a script JSON, chompjs or js2xml parsing Extracts delivered data without starting a browser.
Request signing, browser APIs or complex interaction is essential scrapy-playwright Provides browser execution while retaining Scrapy integration.
Only a screenshot or PDF is required Browser capture service or browser automation The deliverable is browser-rendered rather than structured records.

Reliability, performance and cost controls

  • Prefer stable endpoints. API schemas and pagination are easier to validate than CSS classes generated by a frontend build.
  • Cache during development. Repeated browser launches slow iteration and can trigger site defenses. Use Scrapy’s normal caching and throttling settings where appropriate.
  • Limit browser concurrency. Each page consumes substantially more memory than an HTTP request. Set conservative concurrency and close contexts or pages you explicitly create.
  • Wait on evidence. A selector, response event or network-idle condition is more useful than an arbitrary multi-second sleep, though network idle can be unsuitable for pages with long-lived connections.
  • Log the raw failure. Save status, content type and a short response sample when JSON parsing fails. This distinguishes an expired token, bot check, timeout and changed schema.
  • Respect access rules. Follow the target site’s terms, robots policy and applicable law. Authentication and rate limits are part of the site’s contract, not problems to bypass.

Common failures and fixes

Selectors return nothing

Confirm whether the response contains the data at all. If not, inspect Fetch/XHR requests. If the request is present but the DOM is populated later, enable Playwright and wait for a specific result selector.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #4
ScrapTherapy® Cut the Scraps!: 7 Steps to Quilting Your Way through Your Stash
  • Country of Origin:US
  • CPSIA:N
  • Hazardous?:No
  • Tariff:4901990050

You receive HTML instead of JSON

Check the final URL, status code, cookies, authorization and required headers. A login page, consent page or bot check can have a successful status while violating your expected schema.

JSON decoding fails

Print the first part of the response and its content type. Remove anti-JSON prefixes only when you have confirmed the site’s format. If the value is a JavaScript object with comments or unquoted keys, use a parser intended for JavaScript rather than json.loads.

Playwright times out

Verify that Chromium is installed, increase the navigation timeout only for a known slow page, and replace a brittle selector with a stable one. Check whether a consent dialog blocks the page or whether the URL redirects to authentication.

Items are duplicated

Pagination may repeat records, or browser requests may be scheduled in addition to normal Scrapy requests. Use a stable item key, follow the API’s cursor exactly and keep duplicate filtering enabled unless you have a specific reason to change it.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Direct Playwright code behaves differently from Scrapy

Standalone Playwright does not automatically use Scrapy’s middleware, scheduling and duplicate filtering. Move browser requests into scrapy-playwright when those Scrapy components matter.

Best Value
Sale
Scrap Quilt Secrets: 6 Design Techniques for Knockout Results
  • Suitable for all kinds of project works
  • Acid and toxic free
  • Designed for easy usage
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup

ScreenshotNeo provides a website screenshot API and MCP server when the result you need is a clean image or PDF rather than scraped records. It accepts consent banners before capture and removes more than 60 known consent platforms, newsletter popups and chat widgets; each cleanup step can be disabled. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads and cache hits are not billed, and response headers identify the page verdict and billing status. Its MCP server exposes take_screenshot, get_page_info and capture_pdf to Claude, Cursor and other MCP clients.

One GET request is enough:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

See the ScreenshotNeo documentation for the full option set, including full-page and element captures, device and viewport choices, retina scale, dark mode, PDF paper settings, custom CSS or JavaScript, clicks, selector waits, request blocking, headers, cookies, user agents, authorization, timezone, geolocation, transparent backgrounds, resizing, cache TTLs, signed links, asynchronous webhooks, bulk capture and the usage API.

Python:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

The Free plan includes 1,000 shots per month with no card. Paid plans start at $5 for 3,000 shots; Growth is $15 for 15,000, Pro $39 for 60,000, Scale $99 for 250,000 and Business $249 for 1,000,000. Yearly billing gives two months free, and every feature is included on every plan. Create a free ScreenshotNeo account to start without a card.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

FAQ

Does Scrapy execute JavaScript by default?

No. Standard Scrapy downloads HTTP responses and parses them; JavaScript execution requires request analysis, a parser for embedded code or browser integration.

Should I use Selenium instead of Playwright?

The current Scrapy guidance specifically recommends scrapy-playwright for browser integration. Choose another automation stack only when its browser support or existing project requirements justify the change.

Can I scrape an authenticated single-page application?

Often, yes: reproduce its authenticated API calls if you can establish the session legitimately. If authentication depends on browser interaction, use scrapy-playwright for that flow and then extract the resulting data or responses.

When is a screenshot API the wrong tool?

Use Scrapy and an API request when you need structured records, pagination, joins or large-scale data processing. A screenshot API is aimed at rendered visual output, page information and PDF capture.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Frequently Asked Questions

Can JavaScript be executed without a headless browser?

Sometimes. If the needed value is embedded in the response, parse it; if JavaScript calls a reproducible endpoint, request that endpoint directly. A browser is needed only for behavior that cannot be reproduced reliably or for browser-only output.

Why does a browser show data that Scrapy cannot find?

The browser may have made additional Fetch/XHR requests or transformed script state after the initial response. Compare the raw Scrapy response with the browser’s Network panel to locate the missing source.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.