Hardware FixRecommendedDevice not working? Your driver may be the problemCheck updates for common hardware issues.Fix DriversOctober DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsPC HealthRecommendedCrashes, freezes, slowdowns? Check your PC nowSpot repairable issues before they interrupt work.Check PC×
Skip to content
APIs

How to Collect Big Data from Online Sources: APIs, Scraping, Privacy and Quality

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Short answer: Start with a defined business question, source inventory, target schema and retention policy. Prefer an official API, bulk download or licensed feed. Scrape only when those channels cannot provide the required data, and then identify your crawler, limit requests, follow robots.txt and site terms, and preserve a complete audit trail. Validate, deduplicate, timestamp and document every transformation before the data reaches analysis.

1. Define the dataset before collecting anything

Large volume does not make a dataset useful. Write a collection specification that another engineer can implement without guessing.

State the purpose and decision

Describe the decision the data will support, the population you need to represent, the fields required, the acceptable freshness, and the period you will retain records. A market-monitoring feed, for example, may need daily prices but not customer names. A historical research archive may need immutable snapshots instead of only the latest value.

Set scope and stopping rules

  • List domains, URL patterns, geographic editions, languages and date ranges.
  • Define inclusion and exclusion rules before the first request.
  • Set a maximum request rate, concurrency limit, storage budget and collection end date.
  • Specify how updates, deletions and corrections will be handled.

Design a target schema

Choose stable field names and types for identifiers, dates, numbers, currency, units, categories and source URLs. Record whether a value is missing, not applicable or deliberately redacted; do not collapse those states into an empty string. Version the schema so a later parser can be compared with the one that produced earlier records.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

2. Choose the least risky source channel

Evaluate each source against permission, authority, coverage, freshness, cost, engineering effort, server impact, privacy exposure, reproducibility and the ability to request corrections or deletion.

Channel Best use Strengths Risks and limits
Official API Structured, recurring collection Documented fields, authentication, predictable pagination and clearer contractual terms Quotas, paid tiers, incomplete history or fields that are absent from the public API
Bulk download or licensed feed Large historical loads and repeatable refreshes Efficient transfer, stable files and an explicit licence Refresh schedule, redistribution limits and potentially large files to validate
Web scraping Information exposed only in rendered pages Can capture page-specific content and long-tail sources Layout changes, bot controls, higher server load, uncertain permission and greater privacy risk

Eurostat’s ESS guidance describes both APIs and web scraping as ways to obtain current information, while advising organisations to seek agreements and alternative channels such as APIs and file transfer. Ask the publisher for a feed when the data is important enough to justify recurring collection.

3. Check permission, privacy and intellectual-property constraints

Public does not mean unrestricted

A page that anyone can view is not automatically available for unlimited copying, resale or republication. Check the site’s terms, robots.txt, copyright notices, database rights where applicable, API licence, rate limits and any contract attached to an account. Robots.txt communicates crawler preferences; it is not a substitute for a licence or a legal determination.

Personal data changes the project

The European Data Protection Board states: “The GDPR applies to web scraping when it includes personal data processing operations, such as collection, storage, organisation and retrieval.” If names, emails, user IDs, precise locations, photographs or inferred profiles can identify people, document a lawful basis, purpose, minimisation, transparency notice, security controls, retention period and procedure for access or deletion requests.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Canadian privacy commissioners likewise state that publicly accessible personal information remains subject to privacy laws. Lawful access requires a lawful basis, transparency, appropriate contractual controls and monitoring. Do not assume that removing a login removes privacy obligations.

Apply privacy by design

The European Commission describes privacy by design and default as collecting only what the defined purpose needs, retaining it for the shortest necessary period and restricting access to people who require it. Hash or separate identifiers where practical, encrypt raw data and credentials, and make deletion a tested operation rather than a manual promise.

4. Build an auditable collection pipeline

Separate acquisition from cleaning and analysis so you can reproduce a result without re-downloading a changing website.

  1. Acquire: Save the response, file or page snapshot in immutable raw storage with source and retrieval metadata.
  2. Normalize: Parse into the versioned schema, preserving the original value when a conversion is lossy.
  3. Validate: Run type, range, uniqueness, relationship and completeness checks.
  4. Curate: Deduplicate, correct documented errors, remove irrelevant fields and quarantine records that fail checks.
  5. Publish: Expose only the approved dataset, with a data dictionary and quality report.

For every record or batch, retain the source URL or endpoint, retrieval timestamp in UTC, HTTP status, content hash, parser version, schema version, transformation history, licence or consent reference, and the reason for any exclusion. Keep a manifest that maps each output file to the exact raw inputs.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

5. Implement an API collector first

Use the provider’s documented pagination, authentication and quota headers. The following Python example uses a cursor, bounded retries and checkpointing; replace the endpoint and field names with those in your provider’s documentation.

import json, time
from pathlib import Path
import requests

API_URL = 'https://api.example.org/v1/records'
TOKEN = 'YOUR_TOKEN'
state_file = Path('checkpoint.json')
rows = []
cursor = None
if state_file.exists():
    cursor = json.loads(state_file.read_text()).get('cursor')

session = requests.Session()
session.headers.update({'Authorization': f'Bearer {TOKEN}', 'User-Agent': 'ExampleResearchBot/1.0'})
while True:
    params = {'limit': 100, 'cursor': cursor} if cursor else {'limit': 100}
    for attempt in range(5):
        response = session.get(API_URL, params=params, timeout=30)
        if response.status_code in (429, 500, 502, 503, 504):
            delay = min(60, 2 ** attempt)
            time.sleep(delay)
            continue
        response.raise_for_status()
        break
    payload = response.json()
    batch = payload.get('data', [])
    rows.extend(batch)
    cursor = payload.get('next_cursor')
    state_file.write_text(json.dumps({'cursor': cursor}))
    if not cursor:
        break

Path('raw_records.json').write_text(json.dumps(rows, ensure_ascii=False))

In production, write each page to durable storage before advancing the checkpoint, record response headers that describe quotas, and avoid putting tokens in source control or logs. If the API offers incremental fields such as updated-since or entity tags, use them instead of downloading the full history each time.

Equivalent cURL check

curl -G 'https://api.example.org/v1/records' -H 'Authorization: Bearer YOUR_TOKEN' --data-urlencode 'limit=100' -o page.json

6. Scrape only when it is necessary

First inspect the page for an export, RSS feed, embedded JSON, downloadable file or documented endpoint. If rendering is required, use a browser only for the pages that need it and keep a low, steady request rate.

Responsible request loop

import time, urllib.robotparser
from datetime import datetime, timezone
import requests
from bs4 import BeautifulSoup

root = 'https://example.org'
agent = 'ExampleResearchBot/1.0 ([email protected])'
robots = urllib.robotparser.RobotFileParser(f'{root}/robots.txt')
robots.read()
url = f'{root}/catalog/item-1'
if not robots.can_fetch(agent, url):
    raise RuntimeError('robots.txt disallows this URL')

response = requests.get(url, headers={'User-Agent': agent}, timeout=30)
response.raise_for_status()
soup = BeautifulSoup(response.text, 'html.parser')
record = {
    'source_url': url,
    'retrieved_at': datetime.now(timezone.utc).isoformat(),
    'title': soup.select_one('h1').get_text(' ', strip=True),
}
time.sleep(2.0)
print(record)

Use the site’s contact address in your user agent when appropriate, identify your organisation, minimise downloaded elements, honour disallow rules and stop when the server signals overload. Do not bypass CAPTCHAs, access controls or technical restrictions. For JavaScript-heavy pages, wait for a specific selector or network-idle condition rather than sleeping for an arbitrary long period; save the rendered result and the browser version in your manifest.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

7. Add quality gates before analysis

CNIL describes cleaning as including correction of empty values, detection of outliers, correction of errors, removal of duplicates and deletion of unnecessary fields. Turn those ideas into automated tests with thresholds appropriate to your dataset.

  • Schema: Reject malformed dates, impossible numbers, invalid encodings and unexpected columns.
  • Completeness: Report missingness by source, field and collection date instead of hiding it with defaults.
  • Identity: Deduplicate on a stable source ID where available; otherwise use a documented composite key and retain collision cases.
  • Ranges and relationships: Flag values outside domain limits, negative quantities that cannot be negative, and child records without a parent.
  • Drift: Alert when selectors, column names, category values, page counts or response sizes change sharply.
  • Sampling: Manually inspect a fixed sample from every batch and preserve the inspection result.

Keep rejected records in a quarantine area with the failed rule and parser version. Silent deletion makes later correction impossible.

8. Scale collection without losing control

Partition the workload

Split by source, date range, geographic edition or stable ID ranges. Put jobs on a queue, cap concurrency per host, and checkpoint after each partition. A bounded worker pool is safer than launching one process per URL.

Use caching and conditional requests

Store response hashes and use ETag or Last-Modified headers when offered. Cache immutable historical pages indefinitely; assign a shorter refresh interval to volatile pages. Never treat a cache hit as new evidence without recording when the original response was retrieved.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Make failures recoverable

Retry transient network and server errors with exponential backoff and jitter. Do not retry authentication failures, permanent 404 responses or robots exclusions indefinitely. Send exhausted jobs to a dead-letter queue containing the URL, status, error and last attempt time.

Protect data and credentials

Use secret storage, least-privilege API keys, encrypted transport and encrypted raw storage. Separate production credentials from development data, rotate keys, and redact tokens from exception traces. Restrict raw personal data more tightly than aggregated outputs.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

9. Collect visual page data with ScreenshotNeo

When the required evidence is the rendered appearance of a page rather than its text API, ScreenshotNeo is the first screenshot service to try: it removes consent banners, newsletter popups and chat widgets before capture, bills only clean shots, and has the lowest paid plan described here.

Its website screenshot API accepts one GET request and returns PNG, JPEG, WebP or PDF. You can capture full pages with lazy images loaded, a single CSS-selected element, dark mode, 12 device presets or a custom viewport, retina scale, paper-size PDFs with margins and page ranges, HTML/CSS, custom JavaScript, pre-capture clicks, hidden selectors, waits for selectors, delays or network idle, blocked ads, trackers, requests or resource types, custom headers, cookies, user agents and Authorization, timezone and geolocation, transparent backgrounds, resizing, a chosen cache TTL, signed links, asynchronous jobs with signed webhooks, up to 100 URLs per bulk call, usage data and an OpenAPI specification. Parameter names used by other screenshot APIs also work.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Or skip the browser setup

Use the API directly; the complete options are documented at https://screenshotneo.com/docs/.

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
import requests
r = requests.get('https://api.screenshotneo.com/v1/shot', params={'access_key': 'YOUR_API_KEY', 'url': 'https://stripe.com'}, timeout=90)
r.raise_for_status()
open('shot.webp', 'wb').write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`HTTP ${res.status}`);
const fs = await import('node:fs/promises');
await fs.writeFile('shot.webp', Buffer.from(await res.arrayBuffer()));

Cookie banners, popups and chat widgets are removed before the shot. Bot checks, blank pages and failed loads are never billed, and response headers identify the page verdict and whether the request was billed. Its MCP server lets Claude, Cursor and other MCP clients call take_screenshot, get_page_info and capture_pdf. The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 shots.

Plan Included screenshots Price
Free 1,000 per month $0, no card
Starter 3,000 $5
Growth 15,000 $15
Pro 60,000 $39
Scale 250,000 $99
Business 1,000,000 $249

Yearly billing gives two months free, and every feature is available on every plan. Sign up for the free 1,000-screenshot plan with no card.

10. Troubleshoot common collection failures

HTTP 429 or repeated throttling

Reduce concurrency, obey the provider’s Retry-After value, increase backoff and request a higher quota. Do not rotate identities to evade a limit.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Empty or partial pages

Check whether content is loaded by JavaScript, whether a consent layer blocks the DOM, and whether a selector changed. Capture a diagnostic response, update the parser only after reviewing the rendered page, and quarantine the affected batch.

Duplicates after a restart

Use idempotent keys and commit the checkpoint only after durable writes. Reconcile source IDs and retrieval timestamps before publishing.

Encoding or locale errors

Record the response charset, normalize Unicode deliberately, and set an explicit locale, timezone and decimal convention. Keep the original text beside parsed values.

ScreenshotNeo reports an unbilled failure

Inspect X-Page-Verdict and X-Billed. For bot checks, blank pages or timeouts, adjust waits, headers, cookies, viewport or blocked resources; the failed capture is not billed.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

11. Budget the whole lifecycle

Estimate API or feed fees, bandwidth, proxy or browser infrastructure, storage for raw and curated copies, queue and monitoring costs, engineering time for parser changes, and privacy review. A cheap download can become expensive if it must be reprocessed manually after every layout change. Keep raw retention no longer than your purpose requires, and price deletion, correction handling and audit storage as real work.

The durable result is not merely a large table. It is a permitted, documented and reproducible chain from source response to validated record, with enough provenance to explain where every value came from and when it was observed.

Frequently Asked Questions

Can I merge API records with scraped records from the same publisher?

Yes, but keep separate source-type and retrieval fields, map both into the same versioned schema, and resolve conflicts with a documented precedence rule rather than silently overwriting one channel.

What should I do when a publisher changes or removes historical pages?

Preserve the lawfully collected raw snapshot and its manifest, mark the record’s availability state, and record the publisher’s correction or deletion request. Do not continue retaining material after a valid legal or contractual deletion duty.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How do I prove which parser produced a number in a report?

Store the raw-object hash, retrieval timestamp, parser and schema versions, transformation log and batch identifier, then link the report query to that batch manifest.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Read next

Recommended PC Tool
Recommended PC Tool
Crashes, No Sound, or Screen Glitches?Free driver scan
Windows Errors? Fix Them Before They SpreadFree repair scan

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.