Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

Data scraping is the automated collection of information from websites and its conversion into a structured or analyzable form. A scraper can request pages, find relevant content in HTML or rendered output, extract selected fields, transform them, and store or analyze the result. Whether that activity is appropriate or lawful depends on the data, purpose, jurisdiction, access method, site rules, and what happens to the collected information—not on the word “scraping” alone.

Data scraping, in plain terms

A scraper is a program that performs a repeatable collection task that a person could otherwise do manually. It may retrieve a page, identify a title, price, article body, profile field, table row, or other element, then save those values in a database, spreadsheet, JSON document, or analysis pipeline.

The National Network of Libraries of Medicine (NNLM) describes web scraping as systematic, programmatic collection and processing of online information. In practice, “scraping” and “crawling” overlap, but they emphasize different jobs:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Term Main emphasis Typical result
Scraping Extracting selected fields or content A structured dataset such as rows of titles, dates, or links
Crawling Systematically requesting many pages or discovering URLs A queue, index, or downloaded collection of pages
Archiving Preserving pages for future reference Copies retained for historical or evidentiary use
API access Using an interface the site deliberately documents for data access Responses in a documented format with stated conditions

One project can involve all four. For example, a crawler discovers article URLs, a scraper extracts metadata, an API supplies permitted records, and an archive retains source pages.

How web scraping works

Implementations differ. Some programs parse downloaded HTML; others execute JavaScript in a browser because the needed content appears only after page scripts run. The NNLM notes that researchers use specialized software and customized scripts, and that page HTML can help locate and collect information. HTML parsing is useful, but it is not the only technique.

1. Define the purpose and fields

Write down the question the dataset must answer and the minimum fields required. A precise schema prevents collecting an entire page when you need only a date and an identifier. Decide how often collection is necessary, how long records will be retained, and who can access them.

2. Choose an authorized access route

Check whether the publisher offers an official API, export, feed, or permitted download. A peer-reviewed 2025 study treats official APIs as distinct from scraping access methods. An API can make authentication, fields, limits, and change management clearer, but it does not automatically resolve privacy, copyright, or downstream-use obligations.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

3. Retrieve pages or responses

The program sends requests to the selected URLs or API endpoints, often recording the response status, collection time, and source address. Respect authentication requirements, published rate limits, and technical controls. Avoid bypassing access controls.

4. Locate the content

For HTML, selectors, labels, tables, links, or embedded structured data may identify fields. A browser-based process may wait for a selector, a delay, or network activity to finish before reading the rendered page. Because layouts change, selectors should be specific enough to avoid unrelated elements but not so fragile that a harmless redesign breaks the job.

5. Extract and transform

Convert text, dates, numbers, URLs, and identifiers into consistent formats. Normalize whitespace, preserve the original value when transformation could lose meaning, and record when a field is missing rather than silently substituting a guess.

6. Validate and store

Validation can check required fields, expected types, duplicate records, impossible dates, and sudden changes in record counts. Store provenance: the source URL, timestamp, method, and any version or query parameters that affect the response. Protect the dataset with access controls and encryption appropriate to its sensitivity.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

7. Retain, delete, and monitor

Define a retention period before collection begins. Monitor failures, layout changes, blocked requests, and changes in the source’s terms or API documentation. Delete records that are no longer necessary and honor applicable deletion or correction obligations.

What data scraping is used for

Research is a grounded example. The NNLM identifies researchers using specialized software and scripts to collect web information for analysis. Scraping can turn otherwise unstructured online material into records that can be compared, searched, or analyzed at scale.

The useful question is not simply “Can this page be scraped?” It is “What narrowly defined dataset is needed, from which source, for what purpose, under which permission, and with what safeguards?” That framing is especially important when records concern people.

Scraping versus an official API or download

Use this comparison before writing a collector:

Question Official API or permitted download Scraping
Does the source explicitly offer the route? Usually documented as an intended access method May not be offered or may be restricted by terms
Fields and format Defined by documentation and versioning Determined by page structure and selectors
Freshness and limits Stated quotas, pagination, or update schedules may exist Must be inferred and carefully controlled
Change management Changelogs or version policies may provide notice Layout or script changes can break extraction without warning
Privacy exposure Still depends on the fields and your use Easy to collect more personal information than intended
Validation work Schema and error rules may be documented More responsibility for parsing, provenance, and quality checks

An API is not a blanket legal safe harbor, and scraping is not automatically prohibited. The access route is one fact in a broader assessment.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Is data scraping legal?

There is no universal yes-or-no answer. Consider the data collected, whether individuals can be identified, your purpose, where the people and organization are located, how access was obtained, site terms, technical restrictions, intellectual-property rules, and downstream sharing or sale.

Public does not mean unrestricted

Data-protection authorities have emphasized that personal information can remain protected even when publicly accessible. A public profile, post, or directory entry may still relate to an identified or identifiable person. Reuse, sale, profiling, or intelligence gathering can create harms, and responsibility may apply to both the organization collecting information and the platform hosting it.

European Union and GDPR

The European Commission defines personal data as information relating to an identified or identifiable living person. Pseudonymised information remains personal data if it can be used to re-identify someone. GDPR “processing” includes collection, storage, retrieval, organization, and use, so scraping that handles personal data can fall within the GDPR.

On 8 July 2026, the European Data Protection Board (EDPB) announced adopted guidance on GDPR compliance in web scraping for generative-AI uses. The announcement states: “The GDPR applies to web scraping when it includes personal data processing operations, such as collection, storage, organisation and retrieval.” It highlights legal basis, special-category data, purpose limitation, transparency, reliable sources, timestamp recording, accuracy validation, and data minimisation. This is EU regulatory guidance focused on that context, not a single worldwide rule for every scraping project.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

CNIL guidance

France’s CNIL January 2026 guidance says personal-data collection through scraping is often considered under legitimate interest, but that approach requires additional measures to reduce effects on people’s rights and freedoms. Its focus sheet discusses large-scale collection, difficulty exercising deletion rights, and the risk of collecting private or sensitive information without safeguards. CNIL also notes that site terms, database-producer rights, copyright, robots.txt, and CAPTCHAs may matter. Its guidance should not be converted into one legal test for every country.

United States consumer-data concerns

The U.S. Federal Trade Commission’s 2024 commentary warns that companies may risk enforcement when they fail to honor privacy commitments or use consumer data for other purposes without clear and conspicuous notice and affirmative express consent in the circumstances described by the FTC. This is regulator commentary about consumer-data practices, not a universal scraping statute or a ruling on every scraping dispute.

robots.txt and technical restrictions

robots.txt is a technical crawler convention that communicates paths a site asks crawlers to access or avoid. Google’s documentation explains Google’s interpretation of the robots.txt specification. Treat the file as an important signal to review, not as legal authorization or a substitute for site terms, applicable law, authentication rules, CAPTCHAs, or other access controls.

A responsible scraping checklist

  • Prefer an official API, feed, or permitted download when one meets the requirement.
  • Read the site’s terms, privacy notices, API conditions, and applicable restrictions.
  • Do not bypass authentication, CAPTCHAs, paywalls, bot checks, or other access controls.
  • Collect the minimum fields and frequency necessary for the stated purpose.
  • Classify personal, sensitive, confidential, and publicly available data separately.
  • Record source URLs, collection timestamps, method, and transformation history.
  • Validate accuracy, deduplicate records, and keep the original value where practical.
  • Set retention and deletion rules before collection; restrict internal access.
  • Provide a process for correction, deletion, or other rights where applicable.
  • Recheck terms, permissions, and legal requirements when the project, source, or jurisdiction changes.

This checklist reflects recommendations and concerns described by the European Commission, EDPB, CNIL, other data-protection authorities, the FTC, and NNLM. Following it does not guarantee that a particular project is lawful; consequential uses warrant jurisdiction-specific legal advice.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Common failure modes and safer fixes

The page contains no expected fields

The content may be rendered by JavaScript, moved behind an interaction, or changed by a layout update. Compare the raw response with the rendered page, wait for a specific condition, and add a schema check that stops the job when required fields disappear.

Requests are blocked or challenged

A bot check, CAPTCHA, authentication wall, or rate limit is an access control. Do not attempt to defeat it. Use the site’s documented API or request permission; reduce unnecessary frequency only where the site’s rules allow continued access.

Records are inaccurate or duplicated

Save provenance and timestamps, normalize identifiers, validate types and ranges, and use deterministic deduplication keys. Keep an error queue for records that need human review instead of silently dropping them.

The collector breaks after a redesign

Centralize selectors, monitor extraction completeness, test against representative pages, and version changes. A sudden zero-result run should fail visibly rather than publish an empty dataset.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The dataset contains more personal data than planned

Stop collection, review the purpose and legal basis, remove unnecessary fields, restrict access, and apply retention and deletion rules. Do not assume that because a field was visible it can be retained or reused indefinitely.

Performance, reliability, and cost decisions

Collection frequency should follow the source’s update rate and your purpose, not a default polling habit. Caching previously processed URLs, using backoff after transient errors, and limiting concurrency can reduce load and improve reliability where permitted. Record response status, elapsed time, and failure reason so you can distinguish a source change from a network problem.

Browser rendering generally consumes more resources than retrieving a static response because it starts a browser, executes scripts, and may load images and third-party resources. An API or permitted bulk export can be more predictable when its fields satisfy the requirement. In every approach, budget for validation, monitoring, storage, privacy controls, and deletion—not only request volume.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup: ScreenshotNeo

If your immediate need is a reliable visual capture rather than a structured personal-data dataset, ScreenshotNeo is a website screenshot API and MCP server. It accepts one GET request and returns a PNG, JPEG, WebP, or PDF. Before capture, it can accept cookie or consent banners and remove more than 60 known consent platforms, newsletter popups, and chat widgets; each step can be turned off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers identify the page verdict and whether the request was billed.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use the ScreenshotNeo API documentation for parameters and authentication. Replace the example URL with the page you are authorized to capture.

cURL

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

ScreenshotNeo also supports full-page captures with lazy images loaded, CSS-selector element capture, dark mode, 12 device presets plus custom viewports, retina scale, PDF paper sizes and page ranges, custom CSS and JavaScript, pre-capture clicks, hidden selectors, selector/delay/network-idle waits, request and resource blocking, custom headers, cookies, user agents, Authorization, timezone and geolocation, transparent backgrounds, resizing, configurable-TTL caching, signed image links, asynchronous jobs with signed webhooks, up to 100 URLs per bulk call, a usage API, and an OpenAPI specification. Parameter names used by other screenshot APIs also work for easier migration.

For AI workflows, its MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients.

Every plan includes every feature: Free provides 1,000 shots per month with no card; Starter is $5 for 3,000; Growth $15 for 15,000; Pro $39 for 60,000; Scale $99 for 250,000; and Business $249 for 1,000,000. Yearly billing gives two months free. Create a free ScreenshotNeo account to start with 1,000 screenshots a month and no card.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

FAQ

Frequently Asked Questions

Can scraping collect only public information?

Yes, a program can target public pages, but public visibility does not remove privacy, copyright, database-rights, contractual, or purpose-limitation concerns.

Is an API always safer than scraping?

An official API can clarify the permitted access route, fields, limits, and conditions. You still must assess personal-data processing, retention, security, and downstream use.

Should a scraper store the entire page?

Only when preservation is necessary for the defined purpose and permitted. Otherwise, extract the minimum fields, retain provenance, and set a deletion date.

What should be logged for each collected record?

At minimum, keep the source URL or endpoint, collection timestamp, access method, transformation or parser version, and validation status.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.