What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Ethical web scraping is a controlled, documented way to collect only the information a project needs while respecting permission, privacy, site rules and server capacity. A page being publicly visible does not make its contents free of privacy obligations, and robots.txt is a crawler protocol—not legal authorization. A defensible project checks the target’s rules and applicable law, prefers an API or written permission, identifies itself, limits load, protects personal data, and stops when access is restricted or harm appears.

Is ethical web scraping legal?

There is no universal yes-or-no answer. The result depends on the country or countries involved, the target site’s terms and access controls, the type of data, your purpose, and what you do with the output. Contract, copyright, database rights, confidentiality and computer-access laws may all matter. The sources available for this guide do not establish a single rule that makes every scrape legal or illegal.

Privacy is a separate issue from visibility. A joint statement signed by 16 international privacy regulators says: “Personal information that is publicly accessible is subject to data protection and privacy laws in most jurisdictions.” That means a public profile, directory or article can still contain regulated personal data. Research, journalism or commercial intent does not automatically create an exception.

Think of ethics as a project practice rather than a legal status. You should be able to explain why each page and field is needed, why your access route is appropriate, how you limit disruption, how people can raise concerns, and when you will delete the data. For a high-risk project, obtain advice from a qualified lawyer or privacy professional in every relevant jurisdiction before collecting.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What separates ethical scraping from reckless scraping?

Evaluate the project across the same dimensions you would use for any data-collection review:

Dimension Questions to answer
Authorization Is there an official API, written permission, applicable contract, or crawler instruction for this host? Does the proposed use match it?
Purpose and sensitivity What decision will the dataset support? Could names, contact details, account information, location, health, political views or other sensitive details appear?
Load and operations How many pages are necessary? What request concurrency, caching and back-off behavior will keep the service stable?
Scope and retention Which fields are essential, who receives them, how long are they kept, and what gets deleted immediately?
Transparency Can the crawler be identified, and can affected people or the site operator understand and challenge the collection where required?
Jurisdiction and legal basis Which laws apply, and what documented basis permits processing each category of personal data?
Auditability Can you reproduce the crawl’s date, source URL, rule checks, response status, transformations and deletion decisions?

A practical ethical web-scraping workflow

1. Define the purpose and minimum scope

Write a short collection specification before writing code. State the question the dataset must answer, the exact hosts and paths, the fields required, the people who may be affected, the recipients and the retention period. Exclude credentials, private areas and identifying or sensitive fields unless a specific permission and lawful basis support them.

Use a sampling plan when a complete crawl is unnecessary. A bounded list of URLs, a date range or one page per entity can answer many questions with far less exposure and load than an unrestricted spider. Record excluded fields in the specification so a later code change does not silently expand the project.

2. Check the target and the access route

Read the current terms, API documentation and usage limits for the exact host, subdomain, protocol and port you will contact. If an official API or a written permission route exists, prefer it: privacy regulators note that an API can give the host better control through credentials, logs and monitoring. An API contract still does not make otherwise unlawful personal-data processing lawful.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Download the host’s top-level /robots.txt and identify the rules that match your crawler. Google documents its own interpretation as scoped to the host, protocol and port of the robots URL; do not assume every crawler implements Google-specific behavior. A rule may disallow a path, but it is not a grant of permission to use everything else.

3. Understand what robots.txt does—and does not do

RFC 9309, the Internet Engineering Task Force’s Robots Exclusion Protocol standard, is explicit: “These rules are not a form of access authorization.” It also says: “If the crawler successfully downloads the robots.txt file, the crawler MUST follow the parseable rules.” Treat a successfully retrieved, parseable file as an operational constraint, then separately verify permission, terms and law.

RFC 9309 says crawlers SHOULD NOT use a cached robots file for more than 24 hours unless the file is unreachable. That is a robots-file cache recommendation, not a universal request interval or a statement that 24 hours of crawling is safe. The standard also specifies a parser limit of at least 500 kibibytes; that technical floor is not an ethical allowance to collect 500 KiB of data.

For implementation details that are specific to Google’s crawler, see Google’s robots.txt specification documentation. Never treat a successful fetch, an absent robots file or a permissive rule as authorization to defeat authentication, paywalls, CAPTCHAs or other access controls.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

4. Identify the crawler and reduce impact

Use a clear user-agent that identifies the crawler and, where practical, a contact address or project page. Fetch only the URLs and fields in scope. Avoid parallel bursts, cache responses, and choose conservative limits that reflect the target’s published instructions and observed capacity; there is no universally safe requests-per-second number.

Monitor status codes, latency, response sizes and error rates. Stop or back off on repeated 429, 403, 401 or 5xx responses, connection failures, explicit objections or signs that the service is under strain. Do not rotate identities, disguise traffic, bypass authentication or defeat a CAPTCHA and call the result ethical.

5. Protect people and personal data

Map every field to a purpose before collection. Names, email addresses, phone numbers, account identifiers, precise locations, health information, political opinions and similar attributes may be personal or special-category data depending on the law and context. Minimize fields, restrict access, encrypt stored copies where appropriate, set a deletion date and document the lawful basis.

The European Data Protection Board’s 8 July 2026 announcement on web scraping for generative AI discusses purpose limitation, transparency, accuracy and data minimization. It says processing special-category data requires both a lawful basis under Article 6 and an applicable exception under Article 9(2) of the GDPR. Public visibility and a research label do not automatically supply either condition.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

If people may not reasonably expect the collection, plan the required notice and a way to handle access, correction, objection or deletion requests. Keep a record of the decision, the source and collection timestamp, the transformation steps and the person responsible for responding.

6. Validate, secure and delete

Preserve the source URL, retrieval time, relevant response metadata and a content hash or version marker so later users can tell when a value was observed. Validate important facts against reliable sources before using them. Restrict dataset access to the people who need it, log exports, and delete raw pages and derived records when the stated retention period ends.

The EDPB material is specifically about web scraping in generative-AI contexts; applying its accuracy and governance recommendations to another project is a prudent practice, not a claim that one checklist resolves every legal question.

7. Reassess and stop when conditions change

Recheck terms, API conditions and robots rules before each recurring crawl and after a material site change. Stop if access is revoked, restrictions are added, unexpected sensitive information begins appearing, or the service shows distress. A one-time review cannot guarantee ongoing permission.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A runnable, low-impact Python pattern

The following example is deliberately bounded. It checks robots.txt, identifies the crawler, requests one page at a time, waits between requests, caches successful responses in memory for this run and stops on common refusal or overload signals. The two-second delay is only a conservative starting example; it is not a universal safe rate. Replace the example URLs, user-agent contact and extraction logic after reviewing the target’s rules.

import time
from urllib.parse import urljoin, urlparse
import requests
from urllib.robotparser import RobotFileParser

START_URL = 'https://example.com/'
URLS = [START_URL]
USER_AGENT = 'ExampleResearchBot/1.0 (+mailto:[email protected])'
DELAY_SECONDS = 2.0

parsed = urlparse(START_URL)
robots_url = f'{parsed.scheme}://{parsed.netloc}/robots.txt'
robots = RobotFileParser(robots_url)
try:
    robots.read()
except Exception as exc:
    raise SystemExit(f'Could not retrieve robots.txt: {exc}')

session = requests.Session()
session.headers.update({'User-Agent': USER_AGENT})
cache = {}

for url in URLS:
    if not robots.can_fetch(USER_AGENT, url):
        print(f'Skipping disallowed URL: {url}')
        continue
    if url in cache:
        html = cache[url]
        print(f'Using cached response: {url}')
        continue
    try:
        response = session.get(url, timeout=20)
    except requests.RequestException as exc:
        print(f'Network error; stopping: {exc}')
        break
    if response.status_code in (401, 403, 429):
        print(f'Access or rate limit response {response.status_code}; stopping.')
        break
    if 500 <= response.status_code < 600:
        print(f'Server error {response.status_code}; stopping for back-off.')
        break
    response.raise_for_status()
    cache[url] = response.text
    html = response.text
    print(f'Fetched {url}: {len(html)} characters')
    time.sleep(DELAY_SECONDS)

For a production job, add a persistent cache with an explicit expiry, a maximum page and byte budget, a queue that prevents duplicate URLs, structured logs, exponential back-off for transient failures, and an operator-controlled kill switch. Store only the fields in your collection specification; do not turn this example into an unrestricted crawler.

API, permission and direct crawling: choosing a route

Route Control advantages Risks and checks
Official API Documented fields, credentials, quotas, logs and a channel for the host to revoke or adjust access. Terms may limit uses; personal-data law, purpose limitation and retention still apply.
Written permission Can define paths, fields, timing, support contacts and deletion procedures for a specific project. A contract alone cannot make otherwise unlawful processing lawful; verify privacy and other applicable laws.
Public pages with crawler rules May be appropriate for genuinely public, low-sensitivity information when the host’s instructions and law allow it. Robots.txt is not authorization. You still need a purpose, minimization, low-impact operation and a response plan for objections.

When two routes are technically possible, prefer the one that gives the site operator more visibility and control. Document why the chosen route is proportionate to the purpose.

Performance, reliability and cost controls

  • Bound the job: Set maximum URLs, response bytes, runtime and storage before launch.
  • Cache deliberately: Cache pages and assets you are allowed to reuse; keep the robots file current and do not treat its 24-hour guidance as a crawl-rate promise.
  • Use conditional requests where supported: ETags and Last-Modified headers can avoid downloading unchanged content, subject to the host’s rules.
  • Back off: Increase delays after 429 or 5xx responses and stop on repeated failures instead of retrying indefinitely.
  • Separate extraction from collection: Save a minimal, auditable record first, then parse or enrich it in a controlled job.
  • Plan failure recovery: Checkpoint completed URLs, make writes idempotent, and record skipped or refused pages so a restart does not repeat requests unnecessarily.
  • Budget all costs: Account for bandwidth, storage, API usage, legal review, privacy assessment, monitoring and staff time—not only compute.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Common failure modes and fixes

Robots.txt cannot be downloaded

Do not treat an unavailable file as permission to proceed. Pause, retry according to a bounded policy, contact the operator or use an authorized API. If the file is unreachable, RFC 9309’s 24-hour cache exception concerns using an existing cached file; it does not authorize a new high-volume crawl.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The server returns 403 or 401

Assume access is restricted. Stop and request permission or credentials through the documented channel. Do not rotate IP addresses, spoof identities or attempt to bypass the control.

The server returns 429 or repeated 5xx errors

Reduce concurrency, increase the delay, honor any Retry-After value, and stop if errors persist. Notify the operator if you have a relationship with the site. A successful retry later does not erase the load your earlier requests caused.

Pages contain unexpected personal or sensitive data

Pause the crawl, quarantine the affected records, remove fields that are not necessary, reassess lawful basis and notice obligations, and decide whether the project should continue. Do not distribute the unexpected data while the assessment is unresolved.

Results are stale or inaccurate

Record retrieval timestamps, identify the source version, compare important values with reliable sources and define a correction process. Do not present an old snapshot as current simply because the URL still works.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A site operator asks you to stop or delete data

Stop the relevant collection immediately, preserve only the records needed to investigate your obligations, acknowledge the request, and follow your documented deletion and escalation process. If the request raises a legal dispute, obtain jurisdiction-specific advice rather than continuing by another route.

Or skip the browser setup

If your project needs visual snapshots rather than a DOM data extract, ScreenshotNeo provides a website screenshot API and MCP server. It can accept a cookie or consent banner as a visitor and remove more than 60 known consent platforms, newsletter popups and chat widgets before capture; each cleanup step can be turned off. Use it only for pages you are authorized to access—its controls do not override a site’s restrictions.

A single request returns PNG, JPEG, WebP or PDF. The API reports X-Page-Verdict and X-Billed headers: bot checks or CAPTCHAs, blank pages, timeouts, failed loads and cache hits are not billed, while only clean shots are billed. An MCP server exposes take_screenshot, get_page_info and capture_pdf to Claude, Cursor and other MCP clients.

Relevant controls include full-page capture with lazy images loaded, CSS-selector element capture, dark mode, device presets and custom viewports, retina scale, PDF paper size and page ranges, custom CSS or JavaScript, pre-capture clicks, hidden selectors, waits for a selector, delay or network idle, request and resource blocking, custom headers, cookies, user agent, Authorization, timezone and geolocation, transparent backgrounds, resizing, chosen cache TTLs, signed image links, asynchronous jobs with signed webhooks, bulk capture for up to 100 URLs per call, a usage API and an OpenAPI specification. Parameter names used by other screenshot APIs also work, which can ease migration.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Example cURL (see the ScreenshotNeo documentation):

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

Python:

import requests
r = requests.get('https://api.screenshotneo.com/v1/shot', params={'access_key': 'YOUR_API_KEY', 'url': 'https://stripe.com'}, timeout=90)
open('shot.webp', 'wb').write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

The Free plan includes 1,000 screenshots each month with no card. Paid plans start at $5 for 3,000 shots; yearly billing provides two months free. Create a free ScreenshotNeo account and begin with the 1,000-shot allowance.

Pre-launch checklist

  • Purpose, scope, necessary fields, recipients and retention are documented.
  • Terms, API conditions, robots rules and applicable jurisdictions were checked for the exact host and route.
  • An API or written permission was chosen where practical.
  • The crawler has an identifiable user-agent, bounded URL and byte budgets, caching and back-off.
  • Authentication, CAPTCHA and other access controls will not be bypassed.
  • Personal-data lawful basis, transparency, security, accuracy and deletion procedures are assigned to an owner.
  • Logs capture timestamps, source URLs, decisions, errors, refusals and deletion events.
  • A stop condition and a contact path for objections are tested before the full run.

Frequently Asked Questions

Can a one-time approval cover every future crawl?

Usually not. A recurring job should be reviewed when its purpose, fields, recipient list, target site, terms or access conditions change; keep the approval tied to the scope that was actually authorized.

What should I do if my dataset will be shared with another organization?

Treat the recipient as part of the original purpose and legal analysis. Document what it may receive, why it needs each field, its security and deletion duties, and whether a new notice or agreement is required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Is ethical scraping suitable for training a generative-AI system?

It can require additional analysis. The EDPB’s 8 July 2026 material addresses purpose limitation, transparency, accuracy, minimization and GDPR lawful bases in generative-AI scraping; check the current consultation status and obtain advice for your jurisdictions before collecting.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.