Short answer: for a small, static site, use Python’s HTTP client and an HTML parser. For a multi-page or production crawl, use Scrapy. If the data is rendered by JavaScript, first look for an API or data in the initial response; use browser automation only when that route is unavailable. In every case, check access rules, keep requests bounded, validate what you extract, and treat downloaded content as untrusted.
What web scraping in Python actually involves
Web scraping is the process of requesting web pages and turning their responses into structured data. A typical Python scraper has four stages:
- Request: send an HTTP request with a clear user agent, timeout, and conservative rate.
- Inspect: check the status, content type, size, and whether the response contains the data you need.
- Parse: use stable HTML or JSON selectors to extract fields.
- Persist and monitor: save structured output with its source URL and retrieval time, then detect missing fields or changed markup.
A screenshot is not the same as extracted data. A scraper should normally preserve the values it parsed, while a visual capture is useful for audits, archives, or checking how a page looked at a particular time.
Requests and Beautiful Soup or Scrapy?
Both approaches are valid. Choose according to the crawl rather than by habit.
#1 Best Overall
| Need | HTTP client plus HTML parser | Scrapy |
|---|---|---|
| One-off page or a few URLs | Usually the simplest code and easiest debugging | More setup than necessary |
| Many pages and follow-up links | You must build scheduling and retry logic | Built-in crawling lifecycle, scheduling, concurrency, and callbacks |
| Exports and caching | Add your own storage and cache code | Feed exports and caching are integrated |
| Cookies, sessions, and authentication | Manage them explicitly in your client | Supported through requests, sessions, and middleware |
| Robots.txt handling | Implement a check yourself | RobotsTxtMiddleware can filter forbidden requests when ROBOTSTXT_OBEY is enabled |
| Maintenance burden | Low for a bounded script | Lower for a long-running crawl because common controls are centralized |
Scrapy’s basic lifecycle sends Request objects through a downloader. Returned Response objects go to spider callbacks, which yield items and more requests. That model is why it scales more naturally than a loop that manually manages every URL.
Check permission, limits, and safety first
Legality is specific to the target site and your jurisdiction. Before collecting anything, review the site’s terms, access controls, privacy obligations, and applicable law. A publicly reachable page is not automatically unrestricted for every use.
- Read
/robots.txtand identify the paths your crawler intends to request. - Do not cross an authentication boundary or attempt to defeat a bot check, CAPTCHA, or other access control.
- Use an honest user-agent string that identifies your application and an address for contact when appropriate.
- Set a conservative request rate and concurrency. Stop when the site signals overload or blocks access.
- Collect only the fields you need, especially when pages contain personal information.
- Keep credentials, cookies, and authorization headers out of logs and exported data.
Robots rules are an access preference, not a legal opinion. Scrapy’s middleware filters requests forbidden by the robots exclusion standard when enabled; parser behavior can differ for wildcard rules and for rules with different specificity, so test the exact paths you plan to visit.
A small, complete scraper with Requests and Beautiful Soup
Install the dependencies
python -m pip install requests beautifulsoup4
Fetch, parse, validate, and record provenance
The following script extracts article headings and links from one page. It uses a timeout, checks the response, limits the body size it accepts, and records the URL and retrieval time.
from datetime import datetime, timezone
from urllib.parse import urljoin
import requests
from bs4 import BeautifulSoup
URL = 'https://www.python.org/'
HEADERS = {
'User-Agent': 'ExampleResearchBot/1.0 (contact: [email protected])'
}
MAX_BYTES = 5_000_000
def fetch(url, attempts=3):
last_error = None
for attempt in range(attempts):
try:
response = requests.get(url, headers=HEADERS, timeout=20)
response.raise_for_status()
content_length = response.headers.get('Content-Length')
if content_length and int(content_length) > MAX_BYTES:
raise ValueError('response is larger than the configured limit')
if len(response.content) > MAX_BYTES:
raise ValueError('response is larger than the configured limit')
return response
except (requests.RequestException, ValueError) as error:
last_error = error
raise RuntimeError(f'failed after {attempts} attempts: {last_error}')
response = fetch(URL)
soup = BeautifulSoup(response.text, 'html.parser')
rows = []
for heading in soup.select('h1, h2, h3'):
text = ' '.join(heading.get_text(' ', strip=True).split())
if not text:
continue
rows.append({
'heading': text,
'source_url': response.url,
'retrieved_at': datetime.now(timezone.utc).isoformat(),
})
for link in soup.select('a[href]'):
href = urljoin(response.url, link['href'])
label = ' '.join(link.get_text(' ', strip=True).split())
if label:
rows.append({
'link_text': label,
'url': href,
'source_url': response.url,
'retrieved_at': datetime.now(timezone.utc).isoformat(),
})
for row in rows:
print(row)
Replace the selector with one that reflects the target’s actual markup. Prefer semantic attributes, stable data attributes, or a narrow container over a long chain of positional selectors. Validate required fields before writing a record; an empty string should not silently become a successful row.
When retries help—and when they do not
Retry transient network failures and temporary server responses with a bounded number of attempts. Do not blindly retry a denial, an authentication failure, or a response that violates your access policy. Add a delay between retries and cache successful responses so a rerun does not repeatedly hit the same pages.
Rank #2
Build a multi-page crawler with Scrapy
Scrapy is a Python framework for crawling sites and extracting structured data. It includes selectors, feed exports, caching, cookies and sessions, authentication, crawl-depth controls, and robots.txt support.
Create a project and spider
python -m pip install scrapy
scrapy startproject catalog
cd catalog
scrapy genspider products example.org
Replace the generated spider with this bounded example. It follows only links under the allowed domain and emits one item per page.
Free tools Windows power users keep installed
One-click scans. No signup required.
import scrapy
class ProductsSpider(scrapy.Spider):
name = 'products'
allowed_domains = ['example.org']
start_urls = ['https://example.org/catalog/']
custom_settings = {
'ROBOTSTXT_OBEY': True,
'FEEDS': {'products.jsonl': {'format': 'jsonlines', 'overwrite': True}},
}
def parse(self, response):
for card in response.css('article.product'):
name = card.css('h2::text').get()
price = card.css('.price::text').get()
if name and price:
yield {
'name': name.strip(),
'price': price.strip(),
'source_url': response.url,
}
for href in response.css('a.next::attr(href)').getall():
yield response.follow(href, callback=self.parse)
Run it with scrapy crawl products. Set ROBOTSTXT_OBEY to true for a crawler that should automatically filter requests disallowed by robots.txt. Keep the allowed domain narrow and add explicit depth or pagination bounds when the site can generate an unbounded number of URLs.
How to handle JavaScript-rendered pages
Do not assume that a blank HTML response means the data is inaccessible. First determine whether the needed values are present in an API response or in the initial HTML. If they are, call that endpoint directly with an HTTP client and parse the returned JSON or HTML. This is usually cheaper and simpler than running a browser.
- Request the page once and inspect the source, script data, and network requests in your normal browser developer tools.
- Identify the documented or clearly public data request that supplies the fields you need.
- Reproduce that request with the minimum headers, cookies, and parameters required by the site.
- Respect the same robots, terms, authentication, privacy, and rate constraints as the page request.
- Validate that the API response still contains the required fields before exporting it.
If the values exist only after client-side execution, browser automation may be necessary. It adds operational cost and complexity: you must manage navigation waits, dynamic selectors, browser resources, sessions, and failures that do not occur with a direct HTTP request. Keep that path bounded, and do not use it to bypass a bot check or CAPTCHA.
Robots.txt in Python
For a small script, Python’s standard library can check a site’s robots file before requesting a URL.
from urllib.parse import urlparse
from urllib.robotparser import RobotFileParser
def allowed_by_robots(url, user_agent):
parts = urlparse(url)
robots_url = f'{parts.scheme}://{parts.netloc}/robots.txt'
parser = RobotFileParser(robots_url)
parser.read()
return parser.can_fetch(user_agent, url)
url = 'https://example.org/catalog/'
agent = 'ExampleResearchBot/1.0'
if not allowed_by_robots(url, agent):
raise SystemExit(f'robots.txt disallows {url}')
This check is only one part of responsible access. Rules can use wildcards and overlapping patterns; a parser’s treatment of specificity matters. Keep a copy of the robots file and the decision made for each crawl if you need an audit trail.
Make a scraper resilient to site changes
Use stable selectors and explicit schemas
Define the fields you expect, their types, and which are mandatory. Prefer a selector tied to a semantic element or data attribute. After parsing, reject or quarantine rows missing required values instead of publishing partial records.
Separate collection from export
Write raw or normalized records to a durable format such as JSON Lines or a database, including source URL, retrieval time, and parser version. This lets you reprocess data without downloading every page again.
Cache and schedule conservatively
Cache successful responses and use a bounded queue. Separate transient retries from permanent failures, and keep concurrency low enough that the target remains responsive. Scrapy’s integrated scheduling, caching, and feed exports reduce the amount of infrastructure you need to maintain for a larger crawl.
Recommended Free Tools
Monitor schema drift
Track field-population rates, HTTP statuses, response sizes, and the number of pages visited. Alert when a required selector returns nothing, a page suddenly grows far beyond its normal size, or a crawl begins receiving denials. Save a representative failing response for debugging, subject to your privacy and retention rules.
Security rules for scraped responses
Anything returned by a server is untrusted input. Scrapy explicitly warns that responses can be tampered with in transit or come from a compromised server. Never pass response text to eval, exec, or pickle.loads.
- Parse as data, not executable code.
- Limit response sizes and avoid decompression bombs.
- Keep credentials and cookies in a secret store, not in page content or logs.
- Prevent cross-domain leakage when attaching authorization headers.
- Protect crawler consoles, including Scrapy’s telnet console, from untrusted networks.
- Sanitize filenames and paths derived from URLs before writing files.
Troubleshooting common failures
| Symptom | Likely cause | Fix |
|---|---|---|
| 403 or repeated denials | Access policy, rate, or authentication boundary | Stop, review terms and robots rules, reduce request pressure, and obtain authorized access. Do not try to evade the control. |
| HTML contains no expected fields | Data is rendered by JavaScript or the selector changed | Inspect the initial response and network calls; use the permitted data endpoint or update and test the selector. |
| Rows are empty but requests succeed | Selector is too broad, too narrow, or points at a template | Print a small sample, anchor the selector to a stable container, and enforce required-field validation. |
| Timeouts and connection resets | Slow target, overloaded server, or excessive concurrency | Use a finite timeout, bounded retries with delay, caching, and lower concurrency. |
| Duplicate records | Pagination, redirects, or repeated links | Canonicalize URLs, track visited URLs, and deduplicate on a stable record key. |
| Scrapy follows too many URLs | Unbounded calendars, query strings, or broad link rules | Restrict allowed domains and paths, set crawl-depth or pagination bounds, and filter tracking parameters. |
| Sensitive data appears in output | Overly broad extraction or logging | Reduce fields, redact logs, review retention, and remove data you do not need. |
Or skip the browser setup
When your goal is a reliable visual capture rather than parsed fields, ScreenshotNeo provides a single HTTP endpoint. It accepts cookie and consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each cleanup step can be turned off. Only clean shots are billed: bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and the response identifies the result with X-Page-Verdict and X-Billed headers. It can return PNG, JPEG, WebP, or PDF.
One-call examples
See the full parameter reference in the ScreenshotNeo documentation.
Do these 3 things before closing this tab:
1Fix the driver behind crashes, sound loss and screen glitches2Repair Windows errors before they cause bigger problems3Scan for outdated or missing drivers - takes under a minutecurl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' }); const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
Options that matter for scraping workflows
- Full-page capture with lazy images loaded, or one element selected by CSS selector.
- Dark mode, 12 device presets, arbitrary viewport sizes, and retina scale.
- Wait for a selector, a delay, or network idle; click an element before capture; hide selectors; and inject custom CSS or JavaScript.
- Block ads, trackers, requests, or resource types; provide custom headers, cookies, user agents, and Authorization; set timezone and geolocation.
- PDF paper size, margins, landscape mode, and page ranges; transparent backgrounds and image resizing.
- Configurable caching TTL, signed links for public
<img>tags, asynchronous jobs with signed webhooks, bulk capture of up to 100 URLs per call, a usage API, and an OpenAPI specification. - Parameter names used by other screenshot APIs also work, which can simplify migration.
ScreenshotNeo also offers an MCP server for AI agents, including Claude, Cursor, and other MCP clients, with take_screenshot, get_page_info, and capture_pdf tools. Every feature is on every plan. The Free plan includes 1,000 shots per month without a card; paid plans start at $5 for 3,000 shots. Other listed plans are Starter $5/3,000, Growth $15/15,000, Pro $39/60,000, Scale $99/250,000, and Business $249/1,000,000; yearly billing gives two months free.
Create a free ScreenshotNeo account to get 1,000 screenshots a month with no card.
FAQ
Can I scrape a site that requires a login?
Only when you are authorized and the site’s terms and applicable law permit it. Keep credentials private and do not expose session cookies to another domain.
Should I save the complete HTML response?
Save it only when your retention, privacy, and storage policies allow it. Otherwise, retain the fields needed for verification plus the source URL, retrieval time, and parser version.
Quick wins for a faster PC:
Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Repair Windows errors before they cause bigger problemsFix Now →Is a robots.txt file a guarantee that scraping is legal?
No. It is an automated access signal. Legal permissibility still depends on the site’s terms, access controls, privacy duties, and the law governing your activity.
Best Value
When should I stop retrying?
Stop after a bounded number of attempts, and stop immediately for an access denial, authentication failure, or policy violation. Retries are for transient faults, not for defeating controls.
Frequently Asked Questions
Can I scrape a site that requires a login?
Only when you are authorized and the site’s terms and applicable law permit it. Keep credentials private and do not expose session cookies to another domain.
Should I save the complete HTML response?
Save it only when your retention, privacy, and storage policies allow it. Otherwise, retain the fields needed for verification plus the source URL, retrieval time, and parser version.
Is a robots.txt file a guarantee that scraping is legal?
No. It is an automated access signal. Legal permissibility still depends on the site’s terms, access controls, privacy duties, and the law governing your activity.
When should I stop retrying?
Stop after a bounded number of attempts, and stop immediately for an access denial, authentication failure, or policy violation. Retries are for transient faults, not for defeating controls.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

