Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Web scraping is the automated collection of selected information from websites, followed by organizing it into usable records such as JSON, XML, or database rows. A scraper requests or renders a page, extracts the fields it needs, and validates and stores them. Crawling, by contrast, broadly discovers or downloads pages; scraping focuses on data within them. Before scraping, check whether an official API can provide the information and whether your intended collection is permitted.

What web scraping is—and what it is not

Web scraping turns website content into data that can be analyzed or used by another system. A scraper might collect product names and prices, article titles and dates, or links from pages, then normalize those values into consistent fields. The National Network of Libraries of Medicine describes scraping as collecting information systematically for structured analysis, distinguishing it from crawling and web archiving (NNLM: Web Scraping).

Scraping does not necessarily mean copying an entire site. A well-scoped job requests only the pages and fields needed for a defined purpose. The same idea can apply to data returned by an API, though people often use “scraping” to mean extracting data from web pages that do not offer a suitable API.

Scraping versus crawling

Crawling is the broad discovery or downloading of pages, often by following links across a site. Scraping selects specific information from a response and structures it. A crawler can supply pages to a scraper, and a single program may do both, but they describe different tasks. The distinction matters: finding every page is not the same problem as extracting reliable records from chosen pages.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Scraping versus web archiving

Web archiving aims to preserve pages or sites for later access, while scraping extracts selected data for analysis or another workflow. These activities can overlap in implementation, but their goals, scope, and retention needs differ.

How a web scraper works

A scraper is a sequence of decisions and transformations rather than one particular program. It may use a simple HTTP client and HTML parser, or a browser automation layer when the page depends on JavaScript. A responsible workflow is:

  1. Define the purpose and fields. Specify what information is needed, from which permitted pages, and how it will be used, retained, or shared.
  2. Look for an official API. Read its access terms, limits, and data schema. If it supplies the needed data under suitable conditions, using it is usually more stable than parsing page markup. The UK Food Standards Agency recommends assessing APIs and other collection methods before choosing scraping (Food Standards Agency web scraping policy).
  3. Review site instructions and applicable terms. Check the site’s terms and robots.txt, and review privacy and legal obligations relevant to the data and jurisdictions involved.
  4. Request the page or endpoint. A basic scraper sends an HTTP request and receives HTML, JSON, XML, or another response. A browser-based scraper loads and renders the page when relevant content is generated client-side.
  5. Parse and select fields. Extract only the needed values, such as a title, price, date, or link, using the response structure rather than indiscriminate text copying.
  6. Normalize and validate. Convert fields into consistent forms, identify missing or malformed values, and deduplicate records.
  7. Store and operate carefully. Save records in an appropriate format, use conservative request rates and caching, and monitor failures and changes to the page.

Not every scraper uses the same tools or exact order, but separating acquisition, extraction, validation, and storage makes errors easier to identify.

Should you use an API or scrape HTML?

Assess an official API first. An API is designed for machine access and typically provides a documented schema and clearer access conditions. HTML scraping can be useful when the required public information has no suitable API, but it depends on page structure and can break when the site changes.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Consideration Official API HTML scraping
Access method Requests a service’s documented endpoint, when available. Requests or renders pages, then parses their content.
Schema stability Usually more stable and documented, though providers can change APIs. Depends on page markup and may break when layouts or class names change.
Coverage Limited to the data and access the API offers. Can reach public page information without a suitable API, subject to site instructions and applicable rules.
Maintenance Often less parsing maintenance; clients still need to handle API changes and limits. Requires monitoring and updates for changed markup, pagination, and rendering behavior.
Operational concerns Follow documented quotas and access terms. Control request volume, cache where appropriate, and account for anti-bot measures and site load.
Output Often structured responses such as JSON or XML. Must be parsed and transformed into the desired records.

These are practical tendencies, not guarantees: an API may have restrictive coverage or limits, and a page may have consistent markup. Choose based on whether the method is authorized, fit for the data, and maintainable.

Static pages, JavaScript, and screenshots

Some pages contain the needed content in the initial HTML response. Others assemble it in the browser after scripts run, sometimes following additional network requests or user interaction. A scraper that parses only the initial HTML may therefore find a title but miss dynamically loaded listings or values.

When an HTTP client and parser are enough

For a static page, an HTTP client retrieves the response and an HTML parser selects elements. This is generally simpler and lighter than running a full browser. Confirm that the response actually contains the required fields; a successful status code does not prove that the page’s visible content is present.

When browser rendering is needed

Use browser automation when required content appears only after JavaScript runs or after a specific interaction. Wait for a meaningful selector or page state, rather than assuming a fixed delay always suffices. Browser rendering uses more resources and adds failure points, including scripts that hang, consent interfaces, and anti-bot checks. Do not use automation to bypass CAPTCHAs, authentication, paywalls, or other technical barriers.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

When a screenshot is the actual output

If the goal is a visual record rather than structured fields, a screenshot service may be a better fit than writing browser-capture infrastructure yourself. ScreenshotNeo is a website screenshot API and MCP server for developers; it returns PNG, JPEG, WebP, or PDF output. It is not a substitute for permission to collect data or a way to evade access controls.

What robots.txt means

Google Search Central states: “A robots.txt file tells search engine crawlers which URLs the crawler can access on your site.” (Google Search Central, Robots.txt Introduction and Guide.) A robots.txt file is normally located at a site root and gives crawler instructions for paths on a particular host, protocol, and port. Google documents that crawlers retrieve it with an HTTP GET request and parse its rules. MDN similarly describes it as specifying whether crawlers may access a site or selected resources (MDN robots-related documentation).

Robots.txt is not authentication, encryption, or a reliable way to keep private material secret. Google warns against using it to hide pages from search results; sensitive content should instead be protected with authentication or another access-control mechanism. Rules are crawler instructions, and support or interpretation can vary by crawler (Google Search Central; MDN). Read and respect applicable site instructions, but do not mistake a robots.txt allowance for legal permission or a robots.txt restriction for a technical security boundary.

Common technical problems and practical responses

  • The page structure changes. Selectors can stop matching after a redesign. Validate extracted fields, monitor for missing values, and update parsing logic when the structure changes.
  • Content is missing. The page may be JavaScript-rendered, require pagination, or load content after an interaction. Inspect the response and use an authorized rendering approach if necessary.
  • Pagination creates gaps or duplicates. Track which pages or records have been processed, define a stopping condition, and deduplicate using an appropriate record key.
  • Requests fail intermittently. Transient network or server errors need bounded retries and logging; avoid retry storms. Cache results and keep request rates conservative.
  • The site returns a CAPTCHA, blocks an IP, or flags traffic. These are access signals, not puzzles to defeat. Stop or reduce activity and seek an authorized API or permission. CNIL describes CAPTCHAs and IP-based detection among measures used to identify automated collection (CNIL guidance on web scraping).
  • Collection affects site performance. High request volume can burden a site. Google documents crawl-traffic considerations, and Digital.gov discusses performance concerns in its guidance (Google Search Central crawl-budget guidance; Digital.gov web scraping resource).
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Privacy, legality, and responsible collection

Whether scraping is lawful depends on the circumstances, data, site terms, jurisdiction, and how the collected information is used or shared. Publicly viewable does not automatically mean unrestricted to collect, retain, or republish. Personal data warrants particular care: review applicable privacy duties, document a valid purpose, and minimize collection and retention. CNIL’s guidance discusses controller obligations and publisher protections (CNIL). The UK Food Standards Agency policy calls for documented legal and ethical reasoning for scraping it commissions (FSA policy).

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Document the purpose, expected benefit, fields, retention period, and sharing plan.
  • Prefer a suitable official API and follow its terms and limits.
  • Read robots.txt and site terms; treat them as separate from legal advice or access control.
  • Use conservative request rates, identify the crawler where appropriate, and cache results when suitable.
  • Do not bypass authentication, paywalls, CAPTCHAs, or other technical barriers.
  • Stop if the site indicates that access is not wanted, and reassess the method and legal basis before continuing.

Rules vary across jurisdictions and use cases. This overview is not legal advice; consult qualified counsel when the data or intended use creates material legal risk.

Or skip the browser setup

For a visual capture rather than a structured data scraper, ScreenshotNeo can return a screenshot or PDF from one GET request. Its cleanup steps accept cookie or consent banners like a visitor and remove more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be turned off. Only clean shots are billed: bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and each response identifies the page verdict and billing status in headers. Its MCP server exposes screenshot tools to Claude, Cursor, and other MCP clients. The free plan includes 1,000 shots per month with no card; paid plans start at $5 for 3,000 shots. See the ScreenshotNeo API documentation.

cURL:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

Replace the example URL with the permitted page you want to capture and use your API key. A screenshot is a visual output; it does not extract page fields into records. Sign up for 1,000 free screenshots a month with no card.

FAQ

Does public information mean I can scrape it?

No. Public visibility alone does not settle whether collection, storage, or redistribution is allowed. Review applicable terms, privacy duties, and jurisdiction-specific rules for your use.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Can robots.txt keep a page private?

No. Robots.txt communicates crawler preferences; it is not an access-control mechanism. Use authentication or another security control for private information.

Do I need a browser to scrape a website?

Not always. An HTTP client and parser can handle pages whose needed content is present in the response. Browser rendering is relevant when the data appears only after client-side scripts or interactions.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.