Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

Start with one permitted page, a few clearly defined fields, and a single request. Fetch the HTML, parse it with a selector, print the result, and compare it with the source page. Move to a crawler such as Scrapy only when you need multiple pages, link following, scheduling, or repeatable exports.

There is no universal answer to “Is web scraping legal?” The answer can depend on your jurisdiction, the site’s terms, the content, personal-data rules, and your intended use. Check the target’s published guidance and obtain advice for a specific commercial or jurisdictional project.

What web scraping is—and the smallest useful first project

Web scraping is the controlled retrieval of a web page followed by extraction of selected values from its response. A beginner project should have a narrow scope:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • One target site and one or a few URLs.
  • A short field list, such as article title, price, or publication date.
  • A documented reason to access and reuse the data.
  • A review step that checks extracted records against the source page.

Do not begin by crawling an entire domain. A page can redirect, return an error, expose different markup to an HTTP client than to a browser, or require JavaScript before content appears.

Before your first request: permission, scope and site behavior

Check the site’s published rules

Read the site’s terms and any crawler guidance before writing code. Robots.txt is an access signal for crawlers, not a complete legal permission. Google’s robots.txt guidance also notes that blocking a URL does not guarantee that the URL will stay out of search results. Scraping rules can vary by location, contract, copyright, privacy obligations and purpose; the available evidence does not resolve a particular case.

Keep the request narrow

Use the smallest number of requests and collect only fields you need. Identify your user agent where appropriate, use conservative delays for multiple requests, and stop when the site indicates that access should not continue.

Protect your own machine

If URLs come from users, files or another untrusted source, validate the URL scheme and, where appropriate, the host before requesting it. Scrapy’s security documentation highlights this as a defense against server-side request forgery (SSRF) and related risks. Never let an untrusted input make your program request internal network addresses.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Install a minimal Python toolkit

For a one-page exercise, Python’s requests library plus Beautiful Soup is enough.

  1. Install Python 3 and create an isolated environment: python -m venv .venv.
  2. Activate it: on macOS/Linux, source .venv/bin/activate; on Windows, .venvScriptsactivate.
  3. Install dependencies: python -m pip install requests beautifulsoup4.

Use a real target URL that you are allowed to access. The example below uses https://example.com/ as a harmless placeholder; replace it only after checking your target’s rules.

Fetch and parse one page with Python

This complete script requests one page, checks the HTTP result, parses the HTML, extracts the document title and first heading, and prints a compact record.

import requests
from bs4 import BeautifulSoup
from urllib.parse import urlparse

url = "https://example.com/"
parsed = urlparse(url)
if parsed.scheme not in {"http", "https"} or not parsed.netloc:
    raise ValueError("Only an absolute HTTP(S) URL is allowed")

response = requests.get(
    url,
    headers={"User-Agent": "LearningScraper/1.0 (contact: [email protected])"},
    timeout=20,
)
response.raise_for_status()

soup = BeautifulSoup(response.text, "html.parser")
title = soup.select_one("title")
heading = soup.select_one("h1")
record = {
    "url": response.url,
    "title": title.get_text(" ", strip=True) if title else None,
    "heading": heading.get_text(" ", strip=True) if heading else None,
}
print(record)

The final URL matters because a redirect may have changed the page you received. A missing selector should produce an explicit null or warning, not silently become a false value.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Inspect the response before choosing selectors

Confirm status and content

Print response.status_code, response.url, and response.headers.get("content-type"). Save a small response sample while developing so you can see whether you received HTML, an error document, or something else.

Find stable selectors

Use your browser’s developer tools to inspect the element, then prefer stable attributes or semantic elements over deeply nested positional selectors. For example, article h2 a is generally easier to maintain than a selector containing many generated class names.

Handle repeated records

For a list, select the repeated container and extract each field inside it:

cards = []
for card in soup.select("article"):
    link = card.select_one("h2 a")
    if not link:
        continue
    cards.append({
        "title": link.get_text(" ", strip=True),
        "url": link.get("href"),
    })
print(cards)

Inspect several records manually. A selector can return plausible text from the wrong element after a layout change.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

When static requests are not enough

An HTTP client sees the server response. If the visible data is inserted by browser-side JavaScript, the initial HTML may not contain it. First inspect the response and the page’s network behavior; do not assume that adding random delays to requests will execute JavaScript. A browser automation approach may be needed for a permitted target, with greater cost and operational complexity.

When to move to Scrapy

Scrapy is a Python crawling and extraction framework. Its documented workflow starts requests from URLs, sends responses to callbacks, supports CSS and XPath selectors, and exports data in several formats. Choose it when the job has multiple linked pages, request scheduling, reusable project structure, crawl controls, or recurring exports. Keep a direct request plus parser for a one-off or very small extraction.

Create a project and spider

python -m pip install scrapy
scrapy startproject catalog
cd catalog
scrapy genspider products example.com

Edit the generated spider (the exact module path is shown by the command) so that it yields structured items:

import scrapy

class ProductsSpider(scrapy.Spider):
    name = "products"
    allowed_domains = ["example.com"]
    start_urls = ["https://example.com/"]

    def parse(self, response):
        for card in response.css("article"):
            link = card.css("h2 a::attr(href)").get()
            yield {
                "title": card.css("h2 a::text").get(default="").strip(),
                "url": response.urljoin(link) if link else None,
            }
        for next_page in response.css("a.next::attr(href)").getall():
            yield response.follow(next_page, callback=self.parse)

Run it with a feed export such as scrapy crawl products -O items.json. Test selectors in Scrapy’s interactive shell before starting a larger crawl.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Enable robots.txt behavior deliberately

Scrapy does not respect robots.txt merely because it is installed. Its robots middleware must be enabled and ROBOTSTXT_OBEY set in the project settings:

ROBOTSTXT_OBEY = True

Verify the setting in the configuration you actually run. This technical setting still does not answer whether your planned use is authorized.

Designing a responsible crawler

Bound the crawl

  • Use an explicit allowed-domain list.
  • Set a maximum depth or page count.
  • Follow only the links needed for the task.
  • Cache responses during development instead of repeatedly hitting the site.
  • Record failures and stop or back off on repeated errors.

Validate and normalize data

Normalize whitespace, parse dates with an explicit assumption, and preserve the source URL. Keep raw values when transformation could lose meaning. Deduplicate records using a stable key rather than an array position.

Review output as a dataset

Check for missing fields, duplicate URLs, unexpected content types, redirects, and sudden changes in record counts. A successful HTTP status does not prove that the extracted data is correct.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Common problems and fixes

Symptom Likely cause Fix
403 or 429 response Access policy, rate limit, or bot defense Stop, review permission and published guidance, reduce scope and rate; do not try to bypass controls.
Empty selector result Wrong selector or JavaScript-rendered content Inspect the actual response HTML, test a simpler selector, and determine whether browser execution is required.
Timeout Slow server, network issue, or overloaded crawl Use a bounded timeout, retry conservatively where allowed, and reduce concurrency.
Wrong page saved Redirect, login page, or consent interstitial Check final URL, status, content type and representative text before parsing.
Scrapy follows unwanted links Broad link rule or missing domain restriction Use allowed_domains, narrow callbacks and explicit limits.
SSRF exposure Unvalidated user-supplied URL Allow only HTTP(S), validate hosts, and block private or internal destinations as appropriate to your environment.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup

If your goal is a clean screenshot rather than HTML-field extraction, ScreenshotNeo provides a single-call website screenshot API. Before capture it accepts cookie or consent banners and removes more than 60 known consent platforms, newsletter popups and chat widgets; each cleanup step can be turned off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads and cache hits are not billed, and the response identifies the page verdict and billing status in X-Page-Verdict and X-Billed headers.

It supports PNG, JPEG, WebP and PDF, plus full-page capture with lazy images loaded, CSS-selector element capture, dark mode, device presets or custom viewports, retina scale, custom CSS and JavaScript, clicks, waits, blocked resources, headers, cookies, user agents, authorization, timezone, geolocation, transparent backgrounds, resizing, chosen cache TTLs, signed image links, asynchronous webhooks, bulk capture of up to 100 URLs per call, a usage API and an OpenAPI specification. An MCP server exposes take_screenshot, get_page_info and capture_pdf to Claude, Cursor and other MCP clients.

See the ScreenshotNeo documentation for parameters and authentication. cURL:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

The Free plan includes 1,000 screenshots per month with no card. Paid plans start at $5 for 3,000 shots; yearly billing provides two months free. Create a free ScreenshotNeo account to try it.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Further reading

The publisher-hosted preview of Web Scraping with Python covers Beautiful Soup, Scrapy and legal considerations. The preview does not establish a current edition or availability in a particular marketplace, so verify those details before buying. Scrapy’s official documentation also provides a free learning path from its overview through the tutorial and community resources.

Frequently Asked Questions

Should I save the whole HTML page?

Save only what your task and retention policy require. During development, a limited local sample can help debug selectors; remove unnecessary personal or sensitive data.

Can I ignore robots.txt if a page is public?

No conclusion follows merely from public visibility. Configure your crawler deliberately, review the site’s rules and terms, and get jurisdiction-specific advice for consequential use.

How do I know whether my parser still works tomorrow?

Run a small scheduled check against permitted pages, assert required fields, and alert on missing selectors, changed status codes or unusual record counts.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The Bottom Line

Begin with one authorized URL and a tiny Python parser. Inspect and verify the output before expanding scope. Adopt Scrapy when scheduling, link traversal and exports justify a crawler, and configure robots.txt handling and URL validation explicitly.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.