Use FormRequest to submit known form fields, FormRequest.from_response when the form (and its hidden tokens) was first downloaded, Scrapy’s default cookie middleware to preserve web sessions, and HttpAuthMiddleware for HTTP Basic authentication. These mechanisms solve different problems: a login form is an application-level exchange, while Basic auth is an HTTP challenge. If the page is populated by JavaScript, inspect the browser’s network request and reproduce that request instead of guessing at HTML fields.
Choose the mechanism that matches the site
| Situation | Scrapy approach | Verify |
|---|---|---|
| Known form endpoint and fields | FormRequest |
Action URL, field names, method, encoding and the response outcome |
| Form is present in a downloaded response | FormRequest.from_response |
Correct form selection, hidden fields, tokens and submit control |
| Site creates a cookie-backed session | Default CookiesMiddleware |
Later requests use the same cookie session |
| HTTP Basic challenge | HttpAuthMiddleware or request metadata |
Credentials are restricted to the protected domain |
| Data arrives through XHR, fetch or another browser request | Reproduce the observed request | Method, URL, body, headers, tokens and access authorization |
Do not send Basic credentials merely because a site has a username-and-password form, and do not submit a form when the endpoint is protected by HTTP Basic. Confirm that you are authorized to access the target service and comply with its terms.
Submit a known form with FormRequest
FormRequest is a Request subclass whose formdata values are URL-encoded. Without an explicit method, form data is sent as a POST body. Set method="GET" when the values belong in the query string.
POST form data
import scrapy
class SearchSpider(scrapy.Spider):
name = "search_example"
def start_requests(self):
yield scrapy.FormRequest(
"https://example.org/search",
formdata={"q": "scrapy"},
callback=self.parse_results,
)
def parse_results(self, response):
for result in response.css(".result"):
yield {
"title": result.css("::text").get(),
"url": result.css("a::attr(href)").get(),
}
GET form data
yield scrapy.FormRequest(
"https://example.org/search",
method="GET",
formdata={"q": "scrapy", "page": "2"},
callback=self.parse_results,
)
Use GET only when the endpoint expects parameters in the URL and those values are acceptable to expose in URLs and logs. For either method, inspect the actual form action and control names in the HTML; a label such as “Email” does not guarantee that the submitted field is named email.
#1 Best Overall
Submit a form from a response
Login and checkout forms commonly contain hidden CSRF values, session identifiers, return URLs or a submit button that changes the server-side action. Download the page first, then call FormRequest.from_response. It copies fields from the selected HTML form and lets you override only values such as the username and password.
import scrapy
class LoginSpider(scrapy.Spider):
name = "example_login"
def start_requests(self):
yield scrapy.Request(
"https://example.org/login",
callback=self.parse_login,
)
def parse_login(self, response):
yield scrapy.FormRequest.from_response(
response,
formdata={
"username": "USER_FROM_SECURE_CONFIG",
"password": "SECRET_FROM_SECURE_CONFIG",
},
callback=self.after_login,
)
def after_login(self, response):
# Replace this with a target-specific success check.
if response.css(".account-home"):
yield scrapy.Request(
"https://example.org/account",
callback=self.parse_account,
)
def parse_account(self, response):
yield {"url": response.url, "title": response.css("title::text").get()}
Select the right form and submit control
If the page has more than one form, select it explicitly using the available form selector or index supported by your installed Scrapy version. If a button determines the action, include the corresponding submit field. The current stable documentation is identified as Scrapy 2.19.0, while detailed request documentation can be served from the master branch; check the version installed in your project before copying a helper name or argument verbatim.
Keep credentials out of source and logs
Load secrets from protected environment or deployment configuration, not from a committed spider. Avoid printing request bodies, authorization headers or passwords. A successful HTTP status is not proof of login: check an account-only marker, the redirect destination, an authenticated endpoint or a site-specific error message.
Keep a login session with CookiesMiddleware
CookiesMiddleware is enabled by default. It records cookies received from a response and sends matching cookies on subsequent requests, providing the session continuity a browser normally supplies.
Free tools Windows power users keep installed
One-click scans. No signup required.
class AccountSpider(scrapy.Spider):
name = "account"
def start_requests(self):
yield scrapy.Request("https://example.org/login", callback=self.parse_login)
def parse_login(self, response):
yield scrapy.FormRequest.from_response(
response,
formdata={"username": "USER", "password": "PASS"},
callback=self.after_login,
)
def after_login(self, response):
# The middleware carries the session cookie into this request.
yield scrapy.Request("https://example.org/private", callback=self.parse_private)
def parse_private(self, response):
yield {"private_page": response.css("h1::text").get()}
Send custom cookies correctly
Use the request’s cookies argument for cookies you deliberately provide:
Rank #2
yield scrapy.Request(
"https://example.org/private",
cookies={"session_id": "VALUE_FROM_SECURE_CONFIG"},
callback=self.parse_private,
)
Do not set a raw Cookie header and expect the middleware to manage it. Scrapy’s documented behavior is to drop that header for cookie handling; the cookies argument is the supported path.
Inspect cookie flow safely
Set COOKIES_DEBUG = True to log cookies sent and received while diagnosing a session. Treat those logs as sensitive: a session cookie may grant account access. Enable this only in access-controlled logs and disable it after troubleshooting. Set COOKIES_ENABLED = False only when you intentionally do not want cookie state.
Use HTTP Basic authentication
Scrapy’s HTTP authentication middleware “authenticates requests using Basic access authentication (aka. HTTP auth).” Configure stable credentials in settings:
Recommended Free Tools
HTTPAUTH_USER = "api-user"
HTTPAUTH_PASS = "PASSWORD_FROM_SECRET_STORE"
HTTPAUTH_DOMAIN = "protected.example.org"
The domain setting is a safety boundary, not an optional convenience. If it is unset or None, credentials can be sent to every request, including unrelated hosts in a multi-domain crawl.
Override credentials per request
For a request-specific identity, pass metadata:
yield scrapy.Request(
"https://protected.example.org/report",
meta={
"http_user": "api-user",
"http_pass": "PASSWORD_FROM_SECRET_STORE",
"http_auth_domain": "protected.example.org",
},
callback=self.parse_report,
)
Use settings when credentials remain stable for the spider run; use metadata when a particular request needs an override. Never broaden the domain merely to make a failing request pass.
When the browser, not HTML, performs the login
A page can display a form while JavaScript sends a different JSON or multipart request. In browser developer tools, open the Network panel, perform the login or search, and identify the request that returns the needed response. Record its HTTP method, URL, request body, relevant headers, cookies and tokens. Then reproduce that request with Scrapy.
Start with the smallest equivalent request and add only requirements demonstrated by the browser: an Accept header, JSON body, CSRF header, authorization value or a preceding token request. Browser automation is not automatically required; reproducing the underlying request is often simpler, although complex flows can require substantial developer effort. Scrapy can also construct a request from a copied cURL command, which helps preserve exact method, URL and body details.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Respect token and origin checks
- Fetch the page that issues a CSRF token before posting it.
- Preserve the cookie that binds the token to the session.
- Match the expected content type, such as form encoding versus JSON.
- Send an
OriginorRefererheader only when the target requires it and you are authorized to make the request.
Scrapy’s security guidance notes that a Referer can disclose crawled URLs to another site. The default policy avoids sending a referrer from HTTPS to HTTP; a stricter policy such as same-origin or no-referrer may be appropriate for your crawl.
Diagnose a login that appears to fail
HTTP 200, but still logged out
Check for the account page’s distinctive content, a successful redirect, an authenticated endpoint response or an explicit error message. Many login pages return 200 for both success and failure.
Invalid credentials or missing fields
Compare the submitted names with the HTML controls or the browser Network request. Include the actual submit button value when the server branches on it. For a response-derived form, ensure hidden fields were preserved by from_response.
CSRF or expired-session error
Request a fresh login page immediately before submission, retain its cookies, and submit the current hidden token. Do not cache a token indefinitely or mix cookies from separate sessions.
The Tool Desk
Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Redirect reaches the public site
Inspect the redirect chain and cookie headers with controlled debugging. A missing or rejected session cookie, a wrong cookie domain, or a login endpoint on another host can break continuity. Verify that subsequent requests stay within the intended authentication domain.
Basic credentials leak or challenge repeats
Set HTTPAUTH_DOMAIN or the per-request http_auth_domain explicitly and verify the hostname exactly. Confirm that the endpoint really uses Basic auth rather than a form or bearer-token scheme.
Data exists only after JavaScript
Capture the successful XHR or fetch request and reproduce its method, URL, body and required headers. If the request requires a short-lived token, model the token-fetch and data-fetch sequence instead of hard-coding one observed value.
Operational, performance and security considerations
- Reuse a session when the target expects a sequence of requests, but isolate sessions when accounts or permissions differ.
- Keep timeouts, retries and concurrency within the target’s documented limits; authentication failures should not be retried blindly.
- Record a non-sensitive success signal for each login so downstream items are not collected anonymously.
- Restrict logs containing cookies, authorization headers, request bodies and URLs with sensitive query parameters.
- Follow the target service’s authorization requirements, robots policy where applicable, and applicable law.
Or skip the browser setup
If your goal is to obtain a clean image or PDF of a page while developing or documenting a crawl, ScreenshotNeo provides a single HTTP call rather than a hand-built browser session. Before capture it accepts cookie or consent banners and removes more than 60 known consent platforms, newsletter popups and chat widgets; each cleanup step can be disabled. Only clean shots are billed: bot checks or CAPTCHAs, blank pages, timeouts, failed loads and cache hits cost nothing, with the result identifying the page verdict and billing status in X-Page-Verdict and X-Billed headers. Its MCP server exposes take_screenshot, get_page_info and capture_pdf to Claude, Cursor and other MCP clients.
See the ScreenshotNeo documentation for parameters. cURL:
Best Value
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python:
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
ScreenshotNeo supports full-page captures with lazy images loaded, CSS-selector element capture, dark mode, 12 device presets or custom viewports, retina scale, PDF paper sizes and page ranges, custom CSS and JavaScript, clicks, selector or network-idle waits, request blocking, custom headers and cookies, user-agent, authorization, timezone, geolocation, transparent backgrounds, resizing, chosen cache TTLs, signed image links, asynchronous webhooks, bulk capture of up to 100 URLs per call, a usage API and an OpenAPI specification. Parameters commonly used by other screenshot APIs also work. Every feature is on every plan: 1,000 shots per month are free with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account to start.
Frequently Asked Questions
Does Scrapy manage cookies automatically?
Yes. CookiesMiddleware is enabled by default, stores cookies received from a site and sends matching cookies on later requests. Use the request-level cookies argument for deliberate custom cookies.
How can I see the cookies being sent and received from Scrapy?
Enable COOKIES_DEBUG temporarily. Protect the resulting logs because session cookies can grant access, then turn the setting off.
Do these 3 things before closing this tab:
1Scan for outdated or missing drivers - takes under a minute2Repair Windows errors before they cause bigger problems3Fix the driver behind crashes, sound loss and screen glitchesShould I use FormRequest or HttpAuthMiddleware for a login page?
Use FormRequest for an HTML application form. Use HttpAuthMiddleware only when the server uses HTTP Basic authentication; it does not fill out a site’s login form.
Why does from_response matter when I already know the username and password?
It carries hidden inputs and other form controls supplied by the current response, including tokens or return values that a plain request may omit.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

