Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

Web scraping is the automated process of fetching web pages and extracting selected fields into structured data such as JSON, CSV, or database rows. A scraper usually handles a defined set of pages; a crawler also discovers, schedules, and revisits links. For a small, stable target, an HTTP client and HTML parser may be enough. For a multi-page job, use a framework such as Scrapy, configure limits, validate the output, and treat robots.txt, terms, privacy, copyright, and other legal constraints as separate concerns.

What web scraping does

A scraper turns page content into records. For a product page, fields might be name, price, availability, and canonical URL. The basic loop is:

  1. Request a URL.
  2. Wait for the response and check its status and content type.
  3. Parse HTML (or another permitted representation).
  4. Select fields with CSS selectors or XPath.
  5. Normalize values, validate required fields, and save records.

A crawler adds URL discovery and scheduling. It follows pagination or internal links, avoids revisiting URLs, controls concurrency, and can resume a larger job. Scraping and crawling overlap, but they are not interchangeable terms: one page can be scraped without crawling, while a crawler normally performs extraction as it visits pages.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Choose the smallest approach that fits

One page or a short, known list

Use an HTTP client plus an HTML parser when URLs are known and the server returns the data in the initial response. This has little operational overhead and is easy to run as a scheduled script. It will not execute browser JavaScript, handle complex sessions, or discover a site for you.

Many pages, pagination, or recurring jobs

Scrapy is a documented framework for larger crawls. It schedules asynchronous requests, follows pagination, extracts with CSS or XPath, exports JSON, CSV, or XML feeds, and provides per-domain concurrency, download delays, and an auto-throttling extension. Those controls are configuration features, not a universal safe request rate; tune them for the site and your workload.

Rendered pages

If the required data appears only after client-side JavaScript runs, a plain HTTP request may return an incomplete document. Use the site’s supported API or feed when appropriate. If no suitable API exists, a browser-based collector may be required, with substantially higher CPU, memory, and failure-management costs. The right choice depends on the target’s behavior, not on a blanket claim that one browser package is best.

A minimal Python scraper

Install dependencies with python -m pip install requests beautifulsoup4. The example extracts article links from a page and writes records to standard output. Replace the URL and selectors only for a site you are allowed to access.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
import json
import requests
from bs4 import BeautifulSoup

url = "https://example.com/news"
headers = {"User-Agent": "my-research-bot/1.0 (contact: [email protected])"}
r = requests.get(url, headers=headers, timeout=30)
r.raise_for_status()

soup = BeautifulSoup(r.text, "html.parser")
rows = []
for card in soup.select("article"):
    link = card.select_one("a[href]")
    title = card.select_one("h2, h3")
    if not link or not title:
        continue
    rows.append({
        "title": title.get_text(" ", strip=True),
        "url": link.get("href"),
    })

print(json.dumps(rows, ensure_ascii=False, indent=2))

Production code should resolve relative links, handle redirects, check that the response is HTML, record failures, and validate that each required field is present. Save the response status and retrieval time with each record so a later run can be diagnosed.

Scrapy workflow for a multi-page crawl

Create a project with scrapy startproject catalog, then put a spider in catalog/spiders/. This example follows a “next” link and yields structured items:

import scrapy

class ArticlesSpider(scrapy.Spider):
    name = "articles"
    allowed_domains = ["example.com"]
    start_urls = ["https://example.com/articles"]

    def parse(self, response):
        for card in response.css("article"):
            yield {
                "title": card.css("h2::text, h3::text").get(),
                "url": response.urljoin(card.css("a::attr(href)").get()),
            }
        next_url = response.css("a[rel='next']::attr(href)").get()
        if next_url:
            yield response.follow(next_url, callback=self.parse)

Run it with scrapy crawl articles -O articles.json. Scrapy’s feed exports also support CSV and XML. Keep selectors narrow and test them against representative pages; a selector that silently returns an empty string can be more dangerous than a visible request error.

Define the data contract before requesting pages

  • Fields: specify names, types, units, and whether a field may be null.
  • Scope: list allowed domains, URL patterns, page count, and pagination rules.
  • Freshness: decide whether this is a one-time snapshot, daily update, or change monitor.
  • Destination: choose JSON, CSV, a queue, or a database and define a stable key.
  • Quality checks: reject malformed URLs, impossible values, duplicate records, and pages missing required markers.

Prefer an official API or feed when one exists and fits the use case. It can provide a more stable contract than parsing presentation HTML, but availability and terms differ by site.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Robots.txt, access, and responsible load

RFC 9309 defines robots.txt as a protocol for crawler requests. Its rules are “not a form of access authorization.” A parseable file should be followed after it is successfully retrieved; the standard also describes behavior when the file is unavailable or unreachable and says crawlers generally should not reuse cached content for more than 24 hours unless the file cannot be reached.

Google similarly describes robots.txt as traffic management, not a security control. A disallowed URL can still be discovered or indexed when another page links to it. Therefore, a robots rule neither grants permission nor settles whether collecting data is lawful.

Bound your job to the pages you need. Configure download delays, per-domain concurrency, retries, and (where suitable) auto-throttling. There is no universal “safe” request rate established by these standards; site capacity, response times, and the operator’s instructions matter. Cache responses when your use case permits, identify your client honestly, and stop when a site signals that requests are not welcome.

Browser-free screenshot and rendered-page option

When your goal is a visual capture rather than field extraction, ScreenshotNeo provides a website screenshot API and MCP server. A single GET request returns PNG, JPEG, WebP, or PDF. It accepts consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each cleanup step can be disabled. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and whether it was billed.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

It supports full-page captures with lazy images, CSS-selector element shots, dark mode, 12 device presets or custom viewports, retina scale, PDF paper size/margins/landscape/page ranges, HTML/CSS rendering, custom JavaScript and CSS, pre-capture clicks, hidden selectors, waits for selectors/delays/network idle, request and resource blocking, custom headers/cookies/user agents/Authorization, timezone and geolocation, transparent backgrounds, resizing, chosen cache TTLs, signed public-image links, asynchronous jobs with signed webhooks, bulk capture of up to 100 URLs per call, a usage API, an OpenAPI specification, and familiar parameter names for easier migration.

Or skip the browser setup

Use the API documented at https://screenshotneo.com/docs/:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Cookie banners, popups, and chat widgets are removed before the shot. Bot checks, blank pages, and failed loads are never billed. An MCP server lets AI agents take screenshots, inspect pages, and capture PDFs. The Free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Extraction, rendering, and reliability pitfalls

Selectors break

Markup changes can produce empty or shifted fields without an HTTP error. Track extraction counts, required-field rates, and sample records. Alert when they move outside expected ranges, and keep selectors for stable attributes where possible.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Pagination loops or duplicates

Normalize URLs, remove tracking parameters when appropriate, maintain a visited set, and impose a maximum page count. Treat “next” links as untrusted input; verify that they stay within the allowed domain and path.

Slow or partial responses

Use connect and read timeouts, bounded retries with backoff, and a durable queue for jobs that may resume. Do not retry every status blindly: authentication failures and deliberate blocks usually need a policy decision, not more traffic.

JavaScript data is missing

Inspect the initial HTML and network behavior. Look for a documented endpoint or embedded structured data before introducing a browser. If a browser is unavoidable, wait for a meaningful selector or network-idle condition rather than an arbitrary long sleep, and cap concurrent browser contexts.

Encoding and data quality

Honor the response charset, normalize whitespace, parse locale-specific numbers deliberately, and retain the original text when transformation could lose meaning. Store retrieval timestamps and source URLs alongside normalized values.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Cost and operational planning

HTTP crawlers mainly consume bandwidth, CPU, storage, and engineering time. Browser rendering adds memory and startup latency. Estimate pages per run, average response size, retry rate, retention period, and schedule frequency. A bounded crawl with validation is usually easier to operate than an unbounded “scrape the site” command.

For screenshot workloads, ScreenshotNeo pricing is: Free 1,000 shots/month; Starter $5 for 3,000; Growth $15 for 15,000; Pro $39 for 60,000; Scale $99 for 250,000; Business $249 for 1,000,000. Yearly billing gives two months free, and every feature is on every plan. Only clean shots are billed, while cache hits and the listed failure classes are not.

Legal and policy boundaries

Legal outcomes depend on jurisdiction and facts. Cornell Legal Information Institute’s Wex overview describes screen scraping and summarizes the US Ninth Circuit’s view in hiQ v. LinkedIn that data on a generally public network was likely not access without authorization under the Computer Fraud and Abuse Act. That narrow US summary does not resolve a site’s contract terms, copyright, privacy obligations, database rights, authentication barriers, or other claims. Public visibility is not a complete legal analysis, and robots.txt is not permission. Review applicable law and the target site’s terms with qualified counsel for consequential projects.

Troubleshooting checklist

  • 403 or 429: stop or slow the job, verify permission, reduce concurrency, and follow published guidance; do not try to evade an access control.
  • 200 response but no records: inspect saved HTML, confirm selectors and encoding, and determine whether content is rendered later.
  • Only the first page is collected: log the discovered next URL and test URL normalization and loop limits.
  • Duplicate rows: define a canonical key, deduplicate before storage, and account for URL variants.
  • Intermittent timeouts: separate connect/read timeouts, use bounded backoff, and measure response latency before changing concurrency.
  • Unexpected legal or policy concern: pause collection and obtain a site-specific review rather than assuming that a public page or robots rule answers it.

FAQ

Is scraping the same as crawling?

No. Scraping extracts data; crawling discovers and schedules pages. A crawler often includes a scraper, but a scraper can target one known URL.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Should I store raw HTML?

Keeping a limited, access-controlled copy can make debugging and reprocessing possible, but retention, privacy, and copyright obligations should shape what you keep and for how long.

Can robots.txt protect private information?

No. It is crawler guidance, not authentication or a security boundary. Protect private data with access controls and avoid collecting it without a lawful basis.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.