Driver FixRecommendedSound, Wi-Fi or graphics acting up? Check drivers firstFind missing or outdated drivers fast.Check DriversOctober DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsPC HealthRecommendedCrashes, freezes, slowdowns? Check your PC nowSpot repairable issues before they interrupt work.Check PC×
Skip to content
crawling

What Is Web Scraping? A Complete Guide to How It Works

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Web scraping is the programmatic collection of information from websites, followed by cleaning and storing that information in a usable format such as JSON, XML, CSV or a database. A typical scraper discovers URLs, sends HTTP requests, parses HTML or browser-rendered pages, extracts selected fields, follows links or pagination, validates the results and records where each value came from.

The difficult part is not downloading one page. Reliable scraping requires choosing an authorized access method, controlling request volume, handling JavaScript and failures, respecting robots.txt and site terms, and monitoring selectors when a site changes. This guide explains the complete workflow and includes a runnable Python example.

What web scraping includes

The National Network of Libraries of Medicine describes web scraping as programmatically and systematically collecting information on the web and processing it into analyzable formats that can be serialized, such as JSON or XML, and stored for later use. In practical terms, a scraper turns pages designed for people into records designed for analysis.

  • Fetching: requesting an HTML document, JSON response, feed or rendered page.
  • Parsing: reading the response and locating elements with CSS selectors, XPath or a structured-data parser.
  • Structuring: mapping values into named fields such as title, price, date and URL.
  • Quality control: normalizing text, dates and currencies; removing duplicates; checking required fields.
  • Storage: writing records to JSON, CSV, an object store or a database with source and retrieval metadata.

Scraping is different from copying a page once by hand. It is repeatable, scheduled and explicit about the fields and pages it collects.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How a web-scraping workflow works

1. Define the dataset and permission

Write down the fields you need, the URL scope, how often data should be refreshed, how long it will be retained and who may use it. Check whether the publisher offers an API, feed or downloadable dataset; that channel is normally more stable and easier to govern than parsing presentation HTML. Decide in advance what you will do with personal data, credentials and records that are later removed.

2. Discover URLs

Seed the job with known pages, a sitemap, a feed, search results or links exposed by the site. Keep an allowlist of domains and URL patterns. Canonicalize URLs and remove tracking parameters when they do not identify a distinct resource, otherwise the same page can be downloaded repeatedly.

3. Schedule requests

A crawler maintains a queue, prioritizes URLs, limits concurrency and avoids duplicate work. Scrapy calls this component the scheduler. A production queue also records retry counts, next-attempt times and the reason a URL was deferred.

4. Download responses

The downloader handles HTTP headers, cookies, compression, redirects, timeouts, retries and connection limits. Identify your client honestly where appropriate, use a bounded timeout and treat 429, 403 and repeated 5xx responses as signals to slow down or stop rather than as invitations to increase traffic.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

5. Parse and select fields

Use an HTTP parser when the required data is in the response HTML or embedded structured data. Select fields with stable CSS or XPath expressions and record the selector version used. If JavaScript creates the content only after load, use an authorized browser-rendering step or look for an official endpoint that supplies the data.

6. Follow pagination and links

Extract the next-page URL or an API cursor, enqueue it, and stop at a defined page limit or when no next link remains. Keep a visited set keyed by canonical URL. For link discovery, restrict following to the intended domain and path patterns.

7. Normalize and validate

Strip incidental whitespace, parse dates with an explicit timezone, convert currencies only when the source currency is known, normalize Unicode and distinguish an absent value from an empty string. Validate required fields and reject or quarantine records that fail schema checks. Preserve the original value when a transformation could lose information.

8. Store or export

Write items through a pipeline to JSON, CSV, S3-compatible object storage or a database. Store the source URL, retrieval timestamp, HTTP status, parser version and relevant transformation details alongside the extracted fields. This provenance makes an individual value auditable.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

9. Monitor drift

Track status codes, response sizes, field completeness, duplicate rates, selector failures and crawl volume. Alert when a normally populated field becomes empty or a site redesign changes the HTML. A successful process that silently produces empty records is a failed scraper.

Crawling versus scraping

Activity Primary purpose Typical output
Crawling Discovering, queuing and fetching pages or links Responses, URLs and crawl metadata
Scraping Selecting and structuring information from responses Validated records such as JSON rows

One program often does both: a crawler schedules requests while a spider parses each response into items. The distinction is useful because URL discovery and field extraction have different failure modes and controls.

Choose the least complex access method

Official API, feed or export

Prefer an official interface when it supplies the fields you need. It usually has a documented schema, clearer rate limits and less selector maintenance. Confirm authentication, pagination, retention and licensing terms.

Static HTML or embedded data

Direct HTTP parsing is efficient when the server response already contains the content. Inspect the response, not only what a browser displays, and check for JSON-LD or other structured data before writing fragile selectors.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

JavaScript-rendered pages

Use browser automation or rendering only when the required data is absent from the initial response. Rendering costs more CPU and time and introduces browser, cookie and consent states. Do not use it to bypass a login, paywall or anti-bot control without explicit authorization.

Authenticated or restricted areas

Treat credentials, paywalls and bot checks as access boundaries. Obtain permission, use the provider’s supported API or export, protect secrets and define a stop condition if the operator blocks the account.

Scrape a simple page with Python

This example requests a page, extracts article cards, follows a small number of pagination links and writes records with retrieval timestamps. It assumes the page is publicly accessible and that its terms permit the activity.

Install dependencies

python -m pip install requests beautifulsoup4

Complete script

from datetime import datetime, timezone
from urllib.parse import urljoin
import json
import time

import requests
from bs4 import BeautifulSoup

START_URL = 'https://example.com/news'
MAX_PAGES = 5
HEADERS = {'User-Agent': 'ExampleResearchBot/1.0 ([email protected])'}

session = requests.Session()
session.headers.update(HEADERS)
records = []
seen = set()
url = START_URL

for _ in range(MAX_PAGES):
    if not url or url in seen:
        break
    seen.add(url)
    response = session.get(url, timeout=30)
    if response.status_code == 429:
        raise RuntimeError('The site is rate limiting the crawler; stop and retry later.')
    response.raise_for_status()

    soup = BeautifulSoup(response.text, 'html.parser')
    retrieved_at = datetime.now(timezone.utc).isoformat()
    for card in soup.select('article.card'):
        link = card.select_one('a')
        title = card.select_one('.title')
        if not link or not title:
            continue
        records.append({
            'title': title.get_text(' ', strip=True),
            'url': urljoin(response.url, link.get('href', '')),
            'retrieved_at': retrieved_at,
            'source_page': response.url,
        })

    next_link = soup.select_one('a[rel="next"]')
    url = urljoin(response.url, next_link['href']) if next_link else None
    time.sleep(1.0)  # keep request rate deliberately low

with open('records.json', 'w', encoding='utf-8') as output:
    json.dump(records, output, ensure_ascii=False, indent=2)

Replace the selectors with selectors verified against the target site’s HTML. Keep source_page and retrieved_at; they are essential when a value later needs to be checked. For a real job, add schema validation, persistent retries with backoff, caching and tests that fail when required selectors disappear.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

When to use Scrapy

Scrapy is a Python framework for larger asynchronous crawls. Its architecture centers on an engine, scheduler, downloader, spiders, items, pipelines and feed exports. It provides CSS/XPath selectors, concurrency controls, retries and robots.txt middleware. The official site listed version 2.19.0 in September 2026; check the current release information before pinning a dependency.

import scrapy

class ArticleSpider(scrapy.Spider):
    name = 'articles'
    start_urls = ['https://example.com/news']
    custom_settings = {'ROBOTSTXT_OBEY': True}

    def parse(self, response):
        for card in response.css('article.card'):
            yield {
                'title': card.css('.title::text').get(default='').strip(),
                'url': response.urljoin(card.css('a::attr(href)').get()),
            }
        next_url = response.css('a[rel="next"]::attr(href)').get()
        if next_url:
            yield response.follow(next_url, callback=self.parse)

Run a spider with a feed export such as scrapy crawl articles -O records.json. Set an explicit concurrency and delay instead of accepting aggressive defaults.

What robots.txt does—and does not do

Google Search Central defines robots.txt as a file that tells crawlers which URLs they may access and primarily uses it to manage crawler traffic. Digital.gov likewise describes it as instructions for internet bots that can help manage site performance. It is not an access-control system and does not remove a URL from search results by itself.

Read the file at the site’s root, apply the rules for your user-agent and honor applicable Disallow and Crawl-delay guidance. Scrapy’s ROBOTSTXT_OBEY setting enables its middleware. Scrapy also documents an explicit per-request override that ignores the file; using that is a deliberate governance decision requiring a documented reason and authorization, not a normal fallback.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Is web scraping legal?

There is no worldwide yes-or-no rule. Exposure depends on jurisdiction, authorization, the site’s terms, the kind of data, privacy and copyright interests, database rights, technical behavior and the purpose of the collection. Public visibility is not a universal permission to copy, republish or create profiles.

In its April 18, 2022 preliminary-injunction opinion in hiQ Labs v. LinkedIn, the U.S. Court of Appeals for the Ninth Circuit analyzed whether the Computer Fraud and Abuse Act’s “without authorization” language reaches publicly viewable LinkedIn profiles. The opinion also noted that other claims—such as copyright, breach of contract, trespass to chattels, unjust enrichment, conversion or privacy claims—may still be available. That decision is not a blanket license to scrape every public site, and it does not settle laws outside the circumstances and jurisdiction considered.

  • Review terms of use, API documentation, robots.txt and licensing instructions.
  • Obtain permission or use an official channel for sensitive, restricted or commercially important data.
  • Collect only fields necessary for a defined purpose and protect personal information and credentials.
  • Honor rate limits, stop on explicit operator contact and keep an audit trail of decisions.

Operate a scraper responsibly at scale

Rate, concurrency and caching

Use the lowest request rate and concurrency that meets the requirement. Cache successful responses, avoid duplicate URLs and schedule refreshes according to how often the source changes. Exponential backoff for transient errors is safer than immediate retries.

Data quality

Validate schemas, deduplicate by a stable key, retain raw values where practical and record every transformation. Separate parsing errors from genuinely missing source data so a redesign does not appear as a legitimate run of nulls.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Security and privacy

Keep API keys and session cookies out of source control and logs. Encrypt stored credentials, restrict access to raw personal data and set deletion periods. Do not collect passwords or defeat authentication controls.

Change management

Version selectors and parser code, maintain fixture pages for tests and alert on field-completeness or response-shape changes. A manual review queue is useful when a high-value field suddenly fails validation.

Common failures and fixes

  • 403 or a bot-check page: stop rather than trying to evade it; confirm permission and use an official API or contact the operator.
  • 429 responses: reduce concurrency, honor the stated limit, cache results and retry later with backoff.
  • Empty selectors: inspect the raw response. The content may be JavaScript-rendered, the selector may have changed or a consent page may have been returned.
  • Timeouts: set a finite connect/read timeout, reduce concurrency, retry only transient failures and record the URL for review.
  • Duplicate records: canonicalize URLs, persist a visited set and deduplicate on a source identifier plus normalized fields.
  • Broken characters or dates: respect the response encoding, normalize Unicode and parse dates with an explicit locale and timezone.
  • Pagination loops: track visited URLs and enforce a maximum page count or item count.
  • Silent data loss after a redesign: add required-field assertions and alerts on completeness, not just HTTP success.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

When a screenshot service helps a scraping workflow

A screenshot service is not a replacement for an API or an HTML parser when you need structured records. It is useful when you need a visual snapshot of a rendered page for QA, archival evidence, human review or an AI agent, especially when browser setup is the expensive part of the pipeline.

Or skip the browser setup:

ScreenshotNeo is a website screenshot API and MCP server. It accepts consent banners before capture and removes more than 60 known consent platforms, newsletter popups and chat widgets; each step can be disabled. Only clean shots are billed: bot checks, CAPTCHAs, blank pages, timeouts, failed loads and cache hits are not billed, and the response reports the result with X-Page-Verdict and X-Billed headers.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

See the ScreenshotNeo API documentation for options. The same service supports PNG, JPEG, WebP and PDF output; full-page capture with lazy images, CSS-selector element capture, dark mode, 12 device presets plus custom viewports, retina scale, PDF paper sizes and page ranges, custom CSS and JavaScript, clicks, hidden selectors, selector/delay/network-idle waits, blocking ads, trackers, requests or resource types, custom headers, cookies, user agents and Authorization, timezone and geolocation, transparent backgrounds, resizing, chosen cache TTLs, signed image links, asynchronous jobs with signed webhooks, bulk capture of up to 100 URLs per call, a usage API and an OpenAPI specification. Parameter names used by other screenshot APIs also work, which can simplify migration.

import requests
r = requests.get('https://api.screenshotneo.com/v1/shot', params={'access_key': 'YOUR_API_KEY', 'url': 'https://stripe.com'}, timeout=90)
r.raise_for_status()
open('shot.webp', 'wb').write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`HTTP ${res.status}`);
const fs = await import('node:fs/promises');
await fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer()));

An MCP server exposes take_screenshot, get_page_info and capture_pdf to Claude, Cursor and other MCP clients, so an AI agent can request captures without you building browser orchestration.

Plan Included shots per month Price
Free 1,000 $0, no card
Starter 3,000 $5
Growth 15,000 $15
Pro 60,000 $39
Scale 250,000 $99
Business 1,000,000 $249

Yearly billing gives two months free, and every feature is available on every plan. Create a free ScreenshotNeo account to get 1,000 screenshots a month with no card; paid plans start at $5 for 3,000.

FAQ

Does scraping require a browser?

No. If the response contains the data, an HTTP client and parser are simpler and faster. A browser is needed only for client-rendered content or interactions that an authorized workflow must perform.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What should I keep so another person can reproduce a record?

Keep the source URL, retrieval time, response status, raw response or an immutable snapshot where permitted, parser and selector versions, and a log of transformations.

How often should a scraper run?

Match the schedule to the source’s change rate and your business need. A slower schedule reduces load and cost; increase it only when fresher data has a defined value and the operator’s limits allow it.

Frequently Asked Questions

Does scraping require a browser?

No. If the response contains the data, an HTTP client and parser are simpler and faster. A browser is needed only for client-rendered content or authorized interactions.

What should I keep so another person can reproduce a record?

Keep the source URL, retrieval time, response status, raw response or permitted immutable snapshot, parser and selector versions, and transformation logs.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How often should a scraper run?

Match the schedule to the source’s change rate and your need for freshness, while staying within the operator’s limits.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Read next

Recommended PC Tool
Recommended PC Tool
Windows Errors? Fix Them Before They SpreadFree repair scan
Outdated Drivers Are Slowing You DownFree scan - exact matches

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.