Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Short answer: for a small, static site, use Python’s HTTP client and an HTML parser. For a multi-page or production crawl, use Scrapy. If the data is rendered by JavaScript, first look for an API or data in the initial response; use browser automation only when that route is unavailable. In every case, check access rules, keep requests bounded, validate what you extract, and treat downloaded content as untrusted.

What web scraping in Python actually involves

Web scraping is the process of requesting web pages and turning their responses into structured data. A typical Python scraper has four stages:

  1. Request: send an HTTP request with a clear user agent, timeout, and conservative rate.
  2. Inspect: check the status, content type, size, and whether the response contains the data you need.
  3. Parse: use stable HTML or JSON selectors to extract fields.
  4. Persist and monitor: save structured output with its source URL and retrieval time, then detect missing fields or changed markup.

A screenshot is not the same as extracted data. A scraper should normally preserve the values it parsed, while a visual capture is useful for audits, archives, or checking how a page looked at a particular time.

Requests and Beautiful Soup or Scrapy?

Both approaches are valid. Choose according to the crawl rather than by habit.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Need HTTP client plus HTML parser Scrapy
One-off page or a few URLs Usually the simplest code and easiest debugging More setup than necessary
Many pages and follow-up links You must build scheduling and retry logic Built-in crawling lifecycle, scheduling, concurrency, and callbacks
Exports and caching Add your own storage and cache code Feed exports and caching are integrated
Cookies, sessions, and authentication Manage them explicitly in your client Supported through requests, sessions, and middleware
Robots.txt handling Implement a check yourself RobotsTxtMiddleware can filter forbidden requests when ROBOTSTXT_OBEY is enabled
Maintenance burden Low for a bounded script Lower for a long-running crawl because common controls are centralized

Scrapy’s basic lifecycle sends Request objects through a downloader. Returned Response objects go to spider callbacks, which yield items and more requests. That model is why it scales more naturally than a loop that manually manages every URL.

Check permission, limits, and safety first

Legality is specific to the target site and your jurisdiction. Before collecting anything, review the site’s terms, access controls, privacy obligations, and applicable law. A publicly reachable page is not automatically unrestricted for every use.

  • Read /robots.txt and identify the paths your crawler intends to request.
  • Do not cross an authentication boundary or attempt to defeat a bot check, CAPTCHA, or other access control.
  • Use an honest user-agent string that identifies your application and an address for contact when appropriate.
  • Set a conservative request rate and concurrency. Stop when the site signals overload or blocks access.
  • Collect only the fields you need, especially when pages contain personal information.
  • Keep credentials, cookies, and authorization headers out of logs and exported data.

Robots rules are an access preference, not a legal opinion. Scrapy’s middleware filters requests forbidden by the robots exclusion standard when enabled; parser behavior can differ for wildcard rules and for rules with different specificity, so test the exact paths you plan to visit.

A small, complete scraper with Requests and Beautiful Soup

Install the dependencies

python -m pip install requests beautifulsoup4

Fetch, parse, validate, and record provenance

The following script extracts article headings and links from one page. It uses a timeout, checks the response, limits the body size it accepts, and records the URL and retrieval time.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
from datetime import datetime, timezone
from urllib.parse import urljoin

import requests
from bs4 import BeautifulSoup

URL = 'https://www.python.org/'
HEADERS = {
    'User-Agent': 'ExampleResearchBot/1.0 (contact: [email protected])'
}
MAX_BYTES = 5_000_000


def fetch(url, attempts=3):
    last_error = None
    for attempt in range(attempts):
        try:
            response = requests.get(url, headers=HEADERS, timeout=20)
            response.raise_for_status()
            content_length = response.headers.get('Content-Length')
            if content_length and int(content_length) > MAX_BYTES:
                raise ValueError('response is larger than the configured limit')
            if len(response.content) > MAX_BYTES:
                raise ValueError('response is larger than the configured limit')
            return response
        except (requests.RequestException, ValueError) as error:
            last_error = error
    raise RuntimeError(f'failed after {attempts} attempts: {last_error}')


response = fetch(URL)
soup = BeautifulSoup(response.text, 'html.parser')
rows = []
for heading in soup.select('h1, h2, h3'):
    text = ' '.join(heading.get_text(' ', strip=True).split())
    if not text:
        continue
    rows.append({
        'heading': text,
        'source_url': response.url,
        'retrieved_at': datetime.now(timezone.utc).isoformat(),
    })

for link in soup.select('a[href]'):
    href = urljoin(response.url, link['href'])
    label = ' '.join(link.get_text(' ', strip=True).split())
    if label:
        rows.append({
            'link_text': label,
            'url': href,
            'source_url': response.url,
            'retrieved_at': datetime.now(timezone.utc).isoformat(),
        })

for row in rows:
    print(row)

Replace the selector with one that reflects the target’s actual markup. Prefer semantic attributes, stable data attributes, or a narrow container over a long chain of positional selectors. Validate required fields before writing a record; an empty string should not silently become a successful row.

When retries help—and when they do not

Retry transient network failures and temporary server responses with a bounded number of attempts. Do not blindly retry a denial, an authentication failure, or a response that violates your access policy. Add a delay between retries and cache successful responses so a rerun does not repeatedly hit the same pages.

Build a multi-page crawler with Scrapy

Scrapy is a Python framework for crawling sites and extracting structured data. It includes selectors, feed exports, caching, cookies and sessions, authentication, crawl-depth controls, and robots.txt support.

Create a project and spider

python -m pip install scrapy
scrapy startproject catalog
cd catalog
scrapy genspider products example.org

Replace the generated spider with this bounded example. It follows only links under the allowed domain and emits one item per page.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
import scrapy


class ProductsSpider(scrapy.Spider):
    name = 'products'
    allowed_domains = ['example.org']
    start_urls = ['https://example.org/catalog/']
    custom_settings = {
        'ROBOTSTXT_OBEY': True,
        'FEEDS': {'products.jsonl': {'format': 'jsonlines', 'overwrite': True}},
    }

    def parse(self, response):
        for card in response.css('article.product'):
            name = card.css('h2::text').get()
            price = card.css('.price::text').get()
            if name and price:
                yield {
                    'name': name.strip(),
                    'price': price.strip(),
                    'source_url': response.url,
                }

        for href in response.css('a.next::attr(href)').getall():
            yield response.follow(href, callback=self.parse)

Run it with scrapy crawl products. Set ROBOTSTXT_OBEY to true for a crawler that should automatically filter requests disallowed by robots.txt. Keep the allowed domain narrow and add explicit depth or pagination bounds when the site can generate an unbounded number of URLs.

How to handle JavaScript-rendered pages

Do not assume that a blank HTML response means the data is inaccessible. First determine whether the needed values are present in an API response or in the initial HTML. If they are, call that endpoint directly with an HTTP client and parse the returned JSON or HTML. This is usually cheaper and simpler than running a browser.

  1. Request the page once and inspect the source, script data, and network requests in your normal browser developer tools.
  2. Identify the documented or clearly public data request that supplies the fields you need.
  3. Reproduce that request with the minimum headers, cookies, and parameters required by the site.
  4. Respect the same robots, terms, authentication, privacy, and rate constraints as the page request.
  5. Validate that the API response still contains the required fields before exporting it.

If the values exist only after client-side execution, browser automation may be necessary. It adds operational cost and complexity: you must manage navigation waits, dynamic selectors, browser resources, sessions, and failures that do not occur with a direct HTTP request. Keep that path bounded, and do not use it to bypass a bot check or CAPTCHA.

Robots.txt in Python

For a small script, Python’s standard library can check a site’s robots file before requesting a URL.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
from urllib.parse import urlparse
from urllib.robotparser import RobotFileParser


def allowed_by_robots(url, user_agent):
    parts = urlparse(url)
    robots_url = f'{parts.scheme}://{parts.netloc}/robots.txt'
    parser = RobotFileParser(robots_url)
    parser.read()
    return parser.can_fetch(user_agent, url)


url = 'https://example.org/catalog/'
agent = 'ExampleResearchBot/1.0'
if not allowed_by_robots(url, agent):
    raise SystemExit(f'robots.txt disallows {url}')

This check is only one part of responsible access. Rules can use wildcards and overlapping patterns; a parser’s treatment of specificity matters. Keep a copy of the robots file and the decision made for each crawl if you need an audit trail.

Make a scraper resilient to site changes

Use stable selectors and explicit schemas

Define the fields you expect, their types, and which are mandatory. Prefer a selector tied to a semantic element or data attribute. After parsing, reject or quarantine rows missing required values instead of publishing partial records.

Separate collection from export

Write raw or normalized records to a durable format such as JSON Lines or a database, including source URL, retrieval time, and parser version. This lets you reprocess data without downloading every page again.

Cache and schedule conservatively

Cache successful responses and use a bounded queue. Separate transient retries from permanent failures, and keep concurrency low enough that the target remains responsive. Scrapy’s integrated scheduling, caching, and feed exports reduce the amount of infrastructure you need to maintain for a larger crawl.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Monitor schema drift

Track field-population rates, HTTP statuses, response sizes, and the number of pages visited. Alert when a required selector returns nothing, a page suddenly grows far beyond its normal size, or a crawl begins receiving denials. Save a representative failing response for debugging, subject to your privacy and retention rules.

Security rules for scraped responses

Anything returned by a server is untrusted input. Scrapy explicitly warns that responses can be tampered with in transit or come from a compromised server. Never pass response text to eval, exec, or pickle.loads.

  • Parse as data, not executable code.
  • Limit response sizes and avoid decompression bombs.
  • Keep credentials and cookies in a secret store, not in page content or logs.
  • Prevent cross-domain leakage when attaching authorization headers.
  • Protect crawler consoles, including Scrapy’s telnet console, from untrusted networks.
  • Sanitize filenames and paths derived from URLs before writing files.

Troubleshooting common failures

Symptom Likely cause Fix
403 or repeated denials Access policy, rate, or authentication boundary Stop, review terms and robots rules, reduce request pressure, and obtain authorized access. Do not try to evade the control.
HTML contains no expected fields Data is rendered by JavaScript or the selector changed Inspect the initial response and network calls; use the permitted data endpoint or update and test the selector.
Rows are empty but requests succeed Selector is too broad, too narrow, or points at a template Print a small sample, anchor the selector to a stable container, and enforce required-field validation.
Timeouts and connection resets Slow target, overloaded server, or excessive concurrency Use a finite timeout, bounded retries with delay, caching, and lower concurrency.
Duplicate records Pagination, redirects, or repeated links Canonicalize URLs, track visited URLs, and deduplicate on a stable record key.
Scrapy follows too many URLs Unbounded calendars, query strings, or broad link rules Restrict allowed domains and paths, set crawl-depth or pagination bounds, and filter tracking parameters.
Sensitive data appears in output Overly broad extraction or logging Reduce fields, redact logs, review retention, and remove data you do not need.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup

When your goal is a reliable visual capture rather than parsed fields, ScreenshotNeo provides a single HTTP endpoint. It accepts cookie and consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each cleanup step can be turned off. Only clean shots are billed: bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and the response identifies the result with X-Page-Verdict and X-Billed headers. It can return PNG, JPEG, WebP, or PDF.

One-call examples

See the full parameter reference in the ScreenshotNeo documentation.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' }); const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Options that matter for scraping workflows

  • Full-page capture with lazy images loaded, or one element selected by CSS selector.
  • Dark mode, 12 device presets, arbitrary viewport sizes, and retina scale.
  • Wait for a selector, a delay, or network idle; click an element before capture; hide selectors; and inject custom CSS or JavaScript.
  • Block ads, trackers, requests, or resource types; provide custom headers, cookies, user agents, and Authorization; set timezone and geolocation.
  • PDF paper size, margins, landscape mode, and page ranges; transparent backgrounds and image resizing.
  • Configurable caching TTL, signed links for public <img> tags, asynchronous jobs with signed webhooks, bulk capture of up to 100 URLs per call, a usage API, and an OpenAPI specification.
  • Parameter names used by other screenshot APIs also work, which can simplify migration.

ScreenshotNeo also offers an MCP server for AI agents, including Claude, Cursor, and other MCP clients, with take_screenshot, get_page_info, and capture_pdf tools. Every feature is on every plan. The Free plan includes 1,000 shots per month without a card; paid plans start at $5 for 3,000 shots. Other listed plans are Starter $5/3,000, Growth $15/15,000, Pro $39/60,000, Scale $99/250,000, and Business $249/1,000,000; yearly billing gives two months free.

Create a free ScreenshotNeo account to get 1,000 screenshots a month with no card.

FAQ

Can I scrape a site that requires a login?

Only when you are authorized and the site’s terms and applicable law permit it. Keep credentials private and do not expose session cookies to another domain.

Should I save the complete HTML response?

Save it only when your retention, privacy, and storage policies allow it. Otherwise, retain the fields needed for verification plus the source URL, retrieval time, and parser version.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Is a robots.txt file a guarantee that scraping is legal?

No. It is an automated access signal. Legal permissibility still depends on the site’s terms, access controls, privacy duties, and the law governing your activity.

When should I stop retrying?

Stop after a bounded number of attempts, and stop immediately for an access denial, authentication failure, or policy violation. Retries are for transient faults, not for defeating controls.

Frequently Asked Questions

Can I scrape a site that requires a login?

Only when you are authorized and the site’s terms and applicable law permit it. Keep credentials private and do not expose session cookies to another domain.

Should I save the complete HTML response?

Save it only when your retention, privacy, and storage policies allow it. Otherwise, retain the fields needed for verification plus the source URL, retrieval time, and parser version.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Is a robots.txt file a guarantee that scraping is legal?

No. It is an automated access signal. Legal permissibility still depends on the site’s terms, access controls, privacy duties, and the law governing your activity.

When should I stop retrying?

Stop after a bounded number of attempts, and stop immediately for an access denial, authentication failure, or policy violation. Retries are for transient faults, not for defeating controls.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.