Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use FormRequest to submit known form fields, FormRequest.from_response when the form (and its hidden tokens) was first downloaded, Scrapy’s default cookie middleware to preserve web sessions, and HttpAuthMiddleware for HTTP Basic authentication. These mechanisms solve different problems: a login form is an application-level exchange, while Basic auth is an HTTP challenge. If the page is populated by JavaScript, inspect the browser’s network request and reproduce that request instead of guessing at HTML fields.

Choose the mechanism that matches the site

Situation Scrapy approach Verify
Known form endpoint and fields FormRequest Action URL, field names, method, encoding and the response outcome
Form is present in a downloaded response FormRequest.from_response Correct form selection, hidden fields, tokens and submit control
Site creates a cookie-backed session Default CookiesMiddleware Later requests use the same cookie session
HTTP Basic challenge HttpAuthMiddleware or request metadata Credentials are restricted to the protected domain
Data arrives through XHR, fetch or another browser request Reproduce the observed request Method, URL, body, headers, tokens and access authorization

Do not send Basic credentials merely because a site has a username-and-password form, and do not submit a form when the endpoint is protected by HTTP Basic. Confirm that you are authorized to access the target service and comply with its terms.

Submit a known form with FormRequest

FormRequest is a Request subclass whose formdata values are URL-encoded. Without an explicit method, form data is sent as a POST body. Set method="GET" when the values belong in the query string.

POST form data

import scrapy

class SearchSpider(scrapy.Spider):
    name = "search_example"

    def start_requests(self):
        yield scrapy.FormRequest(
            "https://example.org/search",
            formdata={"q": "scrapy"},
            callback=self.parse_results,
        )

    def parse_results(self, response):
        for result in response.css(".result"):
            yield {
                "title": result.css("::text").get(),
                "url": result.css("a::attr(href)").get(),
            }

GET form data

yield scrapy.FormRequest(
    "https://example.org/search",
    method="GET",
    formdata={"q": "scrapy", "page": "2"},
    callback=self.parse_results,
)

Use GET only when the endpoint expects parameters in the URL and those values are acceptable to expose in URLs and logs. For either method, inspect the actual form action and control names in the HTML; a label such as “Email” does not guarantee that the submitted field is named email.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Submit a form from a response

Login and checkout forms commonly contain hidden CSRF values, session identifiers, return URLs or a submit button that changes the server-side action. Download the page first, then call FormRequest.from_response. It copies fields from the selected HTML form and lets you override only values such as the username and password.

import scrapy

class LoginSpider(scrapy.Spider):
    name = "example_login"

    def start_requests(self):
        yield scrapy.Request(
            "https://example.org/login",
            callback=self.parse_login,
        )

    def parse_login(self, response):
        yield scrapy.FormRequest.from_response(
            response,
            formdata={
                "username": "USER_FROM_SECURE_CONFIG",
                "password": "SECRET_FROM_SECURE_CONFIG",
            },
            callback=self.after_login,
        )

    def after_login(self, response):
        # Replace this with a target-specific success check.
        if response.css(".account-home"):
            yield scrapy.Request(
                "https://example.org/account",
                callback=self.parse_account,
            )

    def parse_account(self, response):
        yield {"url": response.url, "title": response.css("title::text").get()}

Select the right form and submit control

If the page has more than one form, select it explicitly using the available form selector or index supported by your installed Scrapy version. If a button determines the action, include the corresponding submit field. The current stable documentation is identified as Scrapy 2.19.0, while detailed request documentation can be served from the master branch; check the version installed in your project before copying a helper name or argument verbatim.

Keep credentials out of source and logs

Load secrets from protected environment or deployment configuration, not from a committed spider. Avoid printing request bodies, authorization headers or passwords. A successful HTTP status is not proof of login: check an account-only marker, the redirect destination, an authenticated endpoint or a site-specific error message.

Keep a login session with CookiesMiddleware

CookiesMiddleware is enabled by default. It records cookies received from a response and sends matching cookies on subsequent requests, providing the session continuity a browser normally supplies.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
class AccountSpider(scrapy.Spider):
    name = "account"

    def start_requests(self):
        yield scrapy.Request("https://example.org/login", callback=self.parse_login)

    def parse_login(self, response):
        yield scrapy.FormRequest.from_response(
            response,
            formdata={"username": "USER", "password": "PASS"},
            callback=self.after_login,
        )

    def after_login(self, response):
        # The middleware carries the session cookie into this request.
        yield scrapy.Request("https://example.org/private", callback=self.parse_private)

    def parse_private(self, response):
        yield {"private_page": response.css("h1::text").get()}

Send custom cookies correctly

Use the request’s cookies argument for cookies you deliberately provide:

yield scrapy.Request(
    "https://example.org/private",
    cookies={"session_id": "VALUE_FROM_SECURE_CONFIG"},
    callback=self.parse_private,
)

Do not set a raw Cookie header and expect the middleware to manage it. Scrapy’s documented behavior is to drop that header for cookie handling; the cookies argument is the supported path.

Inspect cookie flow safely

Set COOKIES_DEBUG = True to log cookies sent and received while diagnosing a session. Treat those logs as sensitive: a session cookie may grant account access. Enable this only in access-controlled logs and disable it after troubleshooting. Set COOKIES_ENABLED = False only when you intentionally do not want cookie state.

Use HTTP Basic authentication

Scrapy’s HTTP authentication middleware “authenticates requests using Basic access authentication (aka. HTTP auth).” Configure stable credentials in settings:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
HTTPAUTH_USER = "api-user"
HTTPAUTH_PASS = "PASSWORD_FROM_SECRET_STORE"
HTTPAUTH_DOMAIN = "protected.example.org"

The domain setting is a safety boundary, not an optional convenience. If it is unset or None, credentials can be sent to every request, including unrelated hosts in a multi-domain crawl.

Override credentials per request

For a request-specific identity, pass metadata:

yield scrapy.Request(
    "https://protected.example.org/report",
    meta={
        "http_user": "api-user",
        "http_pass": "PASSWORD_FROM_SECRET_STORE",
        "http_auth_domain": "protected.example.org",
    },
    callback=self.parse_report,
)

Use settings when credentials remain stable for the spider run; use metadata when a particular request needs an override. Never broaden the domain merely to make a failing request pass.

When the browser, not HTML, performs the login

A page can display a form while JavaScript sends a different JSON or multipart request. In browser developer tools, open the Network panel, perform the login or search, and identify the request that returns the needed response. Record its HTTP method, URL, request body, relevant headers, cookies and tokens. Then reproduce that request with Scrapy.

Start with the smallest equivalent request and add only requirements demonstrated by the browser: an Accept header, JSON body, CSRF header, authorization value or a preceding token request. Browser automation is not automatically required; reproducing the underlying request is often simpler, although complex flows can require substantial developer effort. Scrapy can also construct a request from a copied cURL command, which helps preserve exact method, URL and body details.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Respect token and origin checks

  • Fetch the page that issues a CSRF token before posting it.
  • Preserve the cookie that binds the token to the session.
  • Match the expected content type, such as form encoding versus JSON.
  • Send an Origin or Referer header only when the target requires it and you are authorized to make the request.

Scrapy’s security guidance notes that a Referer can disclose crawled URLs to another site. The default policy avoids sending a referrer from HTTPS to HTTP; a stricter policy such as same-origin or no-referrer may be appropriate for your crawl.

Diagnose a login that appears to fail

HTTP 200, but still logged out

Check for the account page’s distinctive content, a successful redirect, an authenticated endpoint response or an explicit error message. Many login pages return 200 for both success and failure.

Invalid credentials or missing fields

Compare the submitted names with the HTML controls or the browser Network request. Include the actual submit button value when the server branches on it. For a response-derived form, ensure hidden fields were preserved by from_response.

CSRF or expired-session error

Request a fresh login page immediately before submission, retain its cookies, and submit the current hidden token. Do not cache a token indefinitely or mix cookies from separate sessions.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Redirect reaches the public site

Inspect the redirect chain and cookie headers with controlled debugging. A missing or rejected session cookie, a wrong cookie domain, or a login endpoint on another host can break continuity. Verify that subsequent requests stay within the intended authentication domain.

Basic credentials leak or challenge repeats

Set HTTPAUTH_DOMAIN or the per-request http_auth_domain explicitly and verify the hostname exactly. Confirm that the endpoint really uses Basic auth rather than a form or bearer-token scheme.

Data exists only after JavaScript

Capture the successful XHR or fetch request and reproduce its method, URL, body and required headers. If the request requires a short-lived token, model the token-fetch and data-fetch sequence instead of hard-coding one observed value.

Operational, performance and security considerations

  • Reuse a session when the target expects a sequence of requests, but isolate sessions when accounts or permissions differ.
  • Keep timeouts, retries and concurrency within the target’s documented limits; authentication failures should not be retried blindly.
  • Record a non-sensitive success signal for each login so downstream items are not collected anonymously.
  • Restrict logs containing cookies, authorization headers, request bodies and URLs with sensitive query parameters.
  • Follow the target service’s authorization requirements, robots policy where applicable, and applicable law.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup

If your goal is to obtain a clean image or PDF of a page while developing or documenting a crawl, ScreenshotNeo provides a single HTTP call rather than a hand-built browser session. Before capture it accepts cookie or consent banners and removes more than 60 known consent platforms, newsletter popups and chat widgets; each cleanup step can be disabled. Only clean shots are billed: bot checks or CAPTCHAs, blank pages, timeouts, failed loads and cache hits cost nothing, with the result identifying the page verdict and billing status in X-Page-Verdict and X-Billed headers. Its MCP server exposes take_screenshot, get_page_info and capture_pdf to Claude, Cursor and other MCP clients.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

See the ScreenshotNeo documentation for parameters. cURL:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

ScreenshotNeo supports full-page captures with lazy images loaded, CSS-selector element capture, dark mode, 12 device presets or custom viewports, retina scale, PDF paper sizes and page ranges, custom CSS and JavaScript, clicks, selector or network-idle waits, request blocking, custom headers and cookies, user-agent, authorization, timezone, geolocation, transparent backgrounds, resizing, chosen cache TTLs, signed image links, asynchronous webhooks, bulk capture of up to 100 URLs per call, a usage API and an OpenAPI specification. Parameters commonly used by other screenshot APIs also work. Every feature is on every plan: 1,000 shots per month are free with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account to start.

Frequently Asked Questions

Does Scrapy manage cookies automatically?

Yes. CookiesMiddleware is enabled by default, stores cookies received from a site and sends matching cookies on later requests. Use the request-level cookies argument for deliberate custom cookies.

How can I see the cookies being sent and received from Scrapy?

Enable COOKIES_DEBUG temporarily. Protect the resulting logs because session cookies can grant access, then turn the setting off.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Should I use FormRequest or HttpAuthMiddleware for a login page?

Use FormRequest for an HTML application form. Use HttpAuthMiddleware only when the server uses HTTP Basic authentication; it does not fill out a site’s login form.

Why does from_response matter when I already know the username and password?

It carries hidden inputs and other form controls supplied by the current response, including tokens or return values that a plain request may omit.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.