Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

Use the simplest permitted tool that can reach the data. For one or a few server-rendered pages, fetch HTML with Requests and parse it with Beautiful Soup. For a multi-page crawl, use Scrapy. If a page gets its data from a background request, reproduce that request when appropriate; use a headless browser such as Playwright only when the rendered DOM is genuinely required. In every case, separate fetching, parsing, normalization, validation, and storage, and keep the collection lawful, identifiable, and low impact.

1. Choose a permitted target and define the output

Start with a site you own, a target for which you have permission, or a site that explicitly supports the intended use. Check for an official API or documented feed before writing a scraper. Read the site’s terms and its robots.txt; robots rules guide crawler behavior but do not grant legal permission. The legal answer can depend on the target, data, jurisdiction, access method, contracts, and intended use, so this tutorial is not jurisdiction-specific legal advice.

Write down the fields before choosing selectors. A small catalog might require:

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • title, a non-empty string;
  • author, which may be missing;
  • detail_url, resolved to an absolute HTTPS URL;
  • collected_at, an ISO 8601 timestamp.

Restrict the crawl to those fields and the pages needed to obtain them. A narrow specification makes it easier to detect markup changes, remove duplicates, and stop when the job is complete.

2. Think in five stages

A maintainable scraper is a pipeline rather than one large selector expression:

  1. Fetch: request a URL with a finite timeout and identifiable client headers.
  2. Parse: turn the response into a searchable document.
  3. Normalize: trim text, standardize URLs and dates, and convert types.
  4. Validate: reject or flag records that do not meet your schema.
  5. Store: write structured output with enough logging to diagnose failures.

Keeping these stages separate lets you replace Requests with Scrapy, or HTML parsing with a browser, without rewriting validation and storage.

3. Fetch a static page with Requests and Beautiful Soup

Install the libraries

python -m pip install requests beautifulsoup4

Requests handles HTTP; Beautiful Soup searches the returned markup. The following example uses an illustrative URL, not a tested target or a permission recommendation. Substitute a page you are authorized to access.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
import requests
from bs4 import BeautifulSoup

url = 'https://example.com/page'
headers = {'User-Agent': 'ExampleResearchBot/1.0 (contact: [email protected])'}
response = requests.get(url, headers=headers, timeout=15)
response.raise_for_status()
soup = BeautifulSoup(response.text, 'html.parser')

print(soup.title.get_text(strip=True) if soup.title else 'No title')

The timeout bounds how long the client waits. raise_for_status() turns HTTP errors into visible exceptions instead of allowing an error page to flow into your parser. A descriptive User-Agent gives the operator a way to identify the client.

Inspect markup before writing selectors

Use browser developer tools or save a response fixture, then select stable attributes such as a semantic element, a meaningful class, or a data attribute. Avoid selectors tied to generated class names or a fragile position in the DOM. Never assume a node exists:

def text_or_none(node):
    return node.get_text(' ', strip=True) if node else None

record = {
    'title': text_or_none(soup.select_one('h1')),
    'author': text_or_none(soup.select_one('[data-author]')),
}

link = soup.select_one('a[data-detail]')
record['detail_url'] = link.get('href') if link else None
print(record)

Beautiful Soup supports CSS-style searches and tree navigation. Return None for an absent field, then decide during validation whether that is acceptable. Scrapy’s tutorial makes the same maintainability point: extraction should remain useful even when some elements are not found.

4. Normalize, validate, deduplicate, and store

Raw markup often contains whitespace, relative links, inconsistent casing, and optional values. Normalize before storing, and keep the original URL available for diagnostics.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
from datetime import datetime, timezone
from urllib.parse import urljoin, urlparse

BASE_URL = 'https://example.com/'

def normalize_record(raw):
    title = (raw.get('title') or '').strip()
    author = (raw.get('author') or '').strip() or None
    href = (raw.get('detail_url') or '').strip()
    detail_url = urljoin(BASE_URL, href) if href else None
    return {
        'title': title,
        'author': author,
        'detail_url': detail_url,
        'collected_at': datetime.now(timezone.utc).isoformat(),
    }

def validate(record):
    if not record['title']:
        return False, 'missing title'
    if record['detail_url']:
        parsed = urlparse(record['detail_url'])
        if parsed.scheme not in {'http', 'https'}:
            return False, 'unsupported URL scheme'
    return True, None

clean = normalize_record(record)
ok, reason = validate(clean)
if ok:
    print(clean)
else:
    print('Rejected:', reason)

For a real job, write valid records as JSON Lines or into a database with a unique key. Keep a rejected-record log containing the URL and reason. Deduplicate on a stable identifier (often a canonical URL), not on display text alone. Add a small fixture-based regression test: save representative HTML, run the parser against it, and assert that required fields still appear after selector changes.

5. Follow pagination with Scrapy

Once a job needs many pages, link following, retries, structured crawl state, and exports, Scrapy provides the crawler workflow instead of forcing you to hand-roll a queue. A spider defines starting requests, callback methods, selectors, and yielded items.

Create a project

python -m pip install scrapy
scrapy startproject quotes_project
cd quotes_project
scrapy genspider quotes quotes.toscrape.com

The practice domain above is used in Scrapy’s tutorial examples. Treat tutorial code as illustrative unless you have confirmed that the target permits your use.

Write a resilient spider

import scrapy
from urllib.parse import urljoin

class QuotesSpider(scrapy.Spider):
    name = 'quotes'
    allowed_domains = ['quotes.toscrape.com']
    start_urls = ['https://quotes.toscrape.com/']

    def parse(self, response):
        for card in response.css('div.quote'):
            quote = card.css('span.text::text').get()
            author = card.css('small.author::text').get()
            href = card.css('span a::attr(href)').get()
            yield {
                'quote': quote.strip() if quote else None,
                'author': author.strip() if author else None,
                'author_url': urljoin(response.url, href) if href else None,
            }

        next_href = response.css('li.next a::attr(href)').get()
        if next_href:
            yield response.follow(next_href, callback=self.parse)

Run it with scrapy crawl quotes -O quotes.jsonl. The callback yields one item per card and follows a next link only when one exists. That conditional pattern is important: a missing selector should end pagination cleanly, not crash the crawl. Scrapy’s selector API offers .get() and .getall(); prefer those safe-return methods over indexing an assumed first match.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Inspect selectors interactively

Use the Scrapy shell against an authorized response while refining selectors:

scrapy shell 'https://quotes.toscrape.com/'
response.css('div.quote span.text::text').getall()
response.xpath('//li[@class="next"]/a/@href').get()

Check the result for empty lists, unexpected encodings, and links that leave the allowed host. Configure item validation and export formats in the project rather than relying on ad hoc print statements.

6. Handle JavaScript-rendered pages without guessing

Find the data source first

View the browser’s Network panel, reload the page, and identify the request that supplies the records. Inspect its method, query parameters, request body, response format, pagination, and required headers. If reproducing that request is permitted and practical, it is usually simpler and lighter than rendering a whole browser. Scrapy’s guidance is to locate the data source before reaching for browser automation.

Use a headless browser only when needed

Choose Playwright for Python, or a Scrapy integration, when the required information exists only after scripts execute in the browser DOM, or when the interaction itself is part of the permitted workflow. This is a rendering technique, not a way to defeat bot checks, CAPTCHAs, authentication barriers, or other restrictions. Stop if access is denied or disallowed.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
python -m pip install playwright
playwright install chromium
from playwright.sync_api import sync_playwright

with sync_playwright() as p:
    browser = p.chromium.launch(headless=True)
    page = browser.new_page()
    page.goto('https://example.com/page', wait_until='networkidle', timeout=30000)
    title = page.locator('h1').first.text_content()
    print((title or '').strip())
    browser.close()

Use explicit waits for a known selector when network idle is not meaningful. Limit concurrency and browser lifetime; each rendered page consumes substantially more CPU and memory than an HTTP request.

7. Be polite, secure, and explicit about boundaries

  • Enable robots.txt compliance in your crawler. Scrapy provides middleware for this; verify the setting is enabled for the project.
  • Use a descriptive User-Agent, finite timeouts, sensible concurrency, and a delay appropriate to the site.
  • Cache responses during development so selector work does not repeatedly hit the server.
  • Stop on repeated access denials, explicit opt-outs, or terms that prohibit the activity.
  • Protect API keys, cookies, and authorization headers; never commit them to source control.
  • If URLs come from users, feeds, or files, validate the scheme and hostname before fetching. Restrict private IP ranges and redirects where appropriate to reduce SSRF risk.
  • Do not expose a crawler control endpoint to an untrusted network.

Robots Exclusion Protocol guidance describes crawler instructions, not authorization. Treat it as one operational signal alongside permission, terms, and applicable law.

8. Which Python library should you use?

Need Starting point Why
One or a few static pages Requests + Beautiful Soup Small setup with a clear fetch/parse split.
Many pages with pagination and exports Scrapy Spiders, callbacks, selectors, link following, and crawl state are built in.
Dynamic data supplied by a discoverable request Reproduce the relevant request Usually avoids unnecessary browser rendering.
Data available only in a rendered DOM Playwright or a Scrapy browser integration Executes page scripts when request-level extraction is not practical.

Compare candidates by page complexity, crawl scale, control over requests, setup effort, and operational or security requirements. There is no universal fastest or best library.

9. Reliability, performance, and cost controls

Make failures observable

Log URL, status code, elapsed time, retry count, parser version, and validation outcome. Separate transient network errors from permanent HTTP responses. Save a small sample of failed responses where permitted, but redact credentials and personal data.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Control load and memory

Use bounded concurrency, connection reuse, response-size limits, and streaming exports for large jobs. Follow pagination until the site signals completion; do not generate URLs indefinitely. Browser jobs should reuse a controlled browser context rather than launching a new process for every item.

Test for change

Run fixture tests in continuous integration and alert when required fields fall below an expected completeness threshold. A successful HTTP response does not prove that the page contains the data you need: templates, consent screens, login pages, and error documents can all return status 200.

10. Troubleshooting common failures

Timeouts or connection errors

Confirm DNS and network access, keep a finite timeout, reduce concurrency, and retry only transient failures with backoff. Do not solve a timeout by removing the timeout.

HTTP 403, 429, or a bot challenge

Verify permission, identify your client, slow the crawl, and check the site’s documented access method. Do not present bypassing a challenge as a scraping technique. If access is not permitted, stop.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Empty selectors

Save the response and inspect whether it is the expected page, a consent screen, a login page, or an error document. Re-check the selector against current markup and use safe accessors that tolerate missing nodes.

Content appears in a browser but not in Requests

Inspect Network requests for the underlying JSON or HTML endpoint. Reproduce it only when authorized. If no practical request exists and the DOM is permitted for extraction, use Playwright with an explicit wait.

Duplicate or malformed records

Canonicalize URLs, deduplicate before storage, validate required fields, and record rejected rows with reasons. Keep normalization deterministic so reruns produce comparable output.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup

If your task is to capture a rendered page rather than build a data parser, ScreenshotNeo provides a website screenshot API and MCP server. One GET request returns PNG, JPEG, WebP, or PDF. It accepts cookie and consent banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets; each cleanup step can be turned off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers report the page verdict and billing status.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

See the ScreenshotNeo API documentation for all parameters. A minimal cURL request is:

curl -G 'https://api.screenshotneo.com/v1/shot' -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python:

import requests
r = requests.get('https://api.screenshotneo.com/v1/shot', params={'access_key': 'YOUR_API_KEY', 'url': 'https://stripe.com'}, timeout=90)
open('shot.webp', 'wb').write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Options cover full-page captures with lazy images loaded, CSS-element capture, dark mode, 12 device presets or any viewport, retina scale, PDF paper sizes, margins, landscape and page ranges, custom CSS and JavaScript, click-before-capture, selector hiding, waits for selectors, delays or network idle, request and resource blocking, custom headers, cookies, user agents and Authorization, timezone, geolocation, transparent backgrounds, resizing, chosen cache TTLs, signed links for public <img> tags, asynchronous jobs with signed webhooks, bulk capture of up to 100 URLs per call, a usage API, and an OpenAPI specification. Common parameter names from other screenshot APIs also work.

An MCP server exposes take_screenshot, get_page_info, and capture_pdf to Claude, Cursor, and other MCP clients. Pricing is: Free, 1,000 shots per month with no card; Starter, $5 for 3,000; Growth, $15 for 15,000; Pro, $39 for 60,000; Scale, $99 for 250,000; and Business, $249 for 1,000,000. Yearly billing gives two months free, and every feature is available on every plan.

Create a free ScreenshotNeo account to get 1,000 screenshots a month without a card.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

FAQ

Is scraping the same as crawling?

Crawling describes discovering and requesting pages; scraping describes extracting structured data from them. A project can do either or both.

Should I save raw HTML?

Keeping limited, permission-appropriate fixtures helps reproduce parser bugs and test selector changes. Apply retention, access controls, and redaction when responses contain personal or confidential data.

How do I know a parser still works?

Run fixture-based tests and monitor required-field completeness, validation rejects, HTTP status patterns, and unexpected content types.

Can a scraper use authenticated pages?

Only with explicit authorization and secure credential handling. Follow the service’s documented API or export mechanism when one exists.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Frequently Asked Questions

Is scraping the same as crawling?

Crawling is discovering and requesting pages; scraping is extracting structured data. A project may do either or both.

Should I save raw HTML?

Limited, permission-appropriate fixtures help reproduce parser bugs and test selector changes; protect and redact any sensitive content.

How do I know a parser still works?

Use fixture tests and monitor required-field completeness, validation rejects, HTTP status patterns, and unexpected content types.

Can a scraper use authenticated pages?

Only with explicit authorization and secure credential handling, preferably through the service’s documented API or export path.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.