Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
Build a stock scraper as a small, restartable data pipeline—not a script that only prints prices. Define the symbols, interval, timezone, adjustment policy, freshness target and redistribution rights first. Then place an authorized provider behind a fetch adapter, save every raw response, normalize validated OHLCV records, upsert them into durable storage, and run the job with checkpoints, rate awareness and alerts.
This guide shows a Python implementation using Alpha Vantage for price time series, an SQLite store, and deployment patterns that also work with Postgres or object storage. SEC EDGAR is covered separately for filing and XBRL data, because filings are a different workload from market-price extraction.
1. Define the data contract before writing code
A scraper is easier to replace and audit when its output is specified independently of any provider. Write these decisions in a versioned configuration file or document:
- Universe: ticker symbols for market data, or company identifiers such as CIKs and filing types for SEC data.
- Interval and lookback: daily, weekly, monthly or intraday bars; start and end dates; and whether a backfill is required.
- Time semantics: the canonical timezone for stored timestamps and the market session represented by each bar.
- Adjustment policy: raw prices or adjusted values, and how splits and dividends are represented.
- Freshness: the maximum acceptable delay after a trading session or filing.
- Retention and redistribution: how long raw and cleaned data are kept and whether your product may redistribute it.
Keeping this contract outside provider code means changing vendors does not require rewriting storage, validation or scheduling.
#1 Best Overall
2. Choose an authorized source
Use a documented API or file system whose terms match your use case. Do not build a crawler against an interactive quote page simply because it is visible in a browser.
Alpha Vantage for symbol-based time series
Alpha Vantage documents daily, weekly, monthly and intraday stock endpoints. Its daily response contains open, high, low, close and volume fields, and the documented full option covers more than 25 years of history. The service also documents adjusted-close data, split and dividend information, symbol parameters, API-key authentication, and JSON or CSV output.
The default quote endpoint is updated at the end of each trading day. Real-time or 15-minute-delayed U.S. quotes may require a premium membership. Alpha Vantage notes that real-time and delayed U.S. market data is regulated by exchanges, FINRA and the SEC; commercial users should contact its sales team before relying on the feed.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
SEC EDGAR for filings and XBRL
SEC developer resources provide REST APIs on data.sec.gov for company submissions and extracted XBRL data, returned as JSON. The EDGAR HTTPS file system and RSS feeds support filing searches. Choose this adapter when your product needs reported fundamentals, filing events or XBRL facts rather than exchange price bars. Keep CIK and filing-type keys separate from ticker-symbol keys.
Compare providers on the dimensions that affect your product
| Dimension | Questions to answer |
|---|---|
| Coverage | Does it include your equities, ETFs, funds, regions, filings or XBRL concepts? |
| Interval and latency | Is the feed end-of-day, delayed or real time, and does that match your freshness contract? |
| History and adjustments | How far back does it go, and are split/dividend adjustments explicit? |
| Limits and authentication | What request limits, API-key rules and commercial tiers apply? |
| Reliability and operations | How are errors, maintenance, schema changes and retries handled? |
| Rights and cost | May you store and redistribute the data, and what recurring provider and infrastructure costs result? |
3. Build a provider adapter in Python
Keep extraction behind one function. The rest of the pipeline should receive a predictable response envelope, not provider-specific field names.
Install dependencies and set the key
python -m venv .venv
. .venv/bin/activate
pip install requests
export ALPHA_VANTAGE_KEY='YOUR_API_KEY'
Do not commit the key. Use your host, container platform or secret manager to inject it at runtime.
Fetch, preserve and parse a daily response
import hashlib
import json
import os
import time
from datetime import datetime, timezone
from pathlib import Path
import requests
API_URL = os.getenv('ALPHA_VANTAGE_URL', 'https://www.alphavantage.co/query')
API_KEY = os.environ['ALPHA_VANTAGE_KEY']
RAW_DIR = Path(os.getenv('RAW_DIR', 'raw'))
def fetch_daily(symbol, outputsize='full', session=None, attempts=4):
client = session or requests.Session()
params = {
'function': 'TIME_SERIES_DAILY_ADJUSTED',
'symbol': symbol,
'outputsize': outputsize,
'apikey': API_KEY,
}
for attempt in range(attempts):
retrieved_at = datetime.now(timezone.utc).isoformat()
try:
response = client.get(API_URL, params=params, timeout=30)
response.raise_for_status()
payload = response.json()
if 'Note' in payload:
raise RuntimeError(f'provider limit: {payload["Note"]}')
if 'Error Message' in payload:
raise RuntimeError(payload['Error Message'])
if not any(key.startswith('Time Series') for key in payload):
raise RuntimeError('response has no time-series field')
raw = json.dumps(payload, sort_keys=True).encode('utf-8')
RAW_DIR.mkdir(parents=True, exist_ok=True)
digest = hashlib.sha256(raw).hexdigest()
path = RAW_DIR / f'{symbol}-{digest}.json'
if not path.exists():
path.write_bytes(raw)
return {
'provider': 'alpha_vantage',
'symbol': symbol,
'retrieved_at': retrieved_at,
'request': params,
'sha256': digest,
'payload': payload,
}
except (requests.RequestException, ValueError, RuntimeError):
if attempt == attempts - 1:
raise
time.sleep(2 ** attempt)
def parse_daily(envelope):
payload = envelope['payload']
series_key = next(k for k in payload if k.startswith('Time Series'))
rows = []
for day, values in payload[series_key].items():
row = {
'provider': envelope['provider'],
'symbol': envelope['symbol'],
'interval': '1d',
'timestamp': f'{day}T00:00:00+00:00',
'open': float(values['1. open']),
'high': float(values['2. high']),
'low': float(values['3. low']),
'close': float(values['4. close']),
'adjusted_close': float(values.get('5. adjusted close', values['4. close'])),
'volume': int(values['6. volume']),
'adjustment_state': 'adjusted_close_present',
'provider_retrieved_at': envelope['retrieved_at'],
'raw_sha256': envelope['sha256'],
}
validate_row(row)
rows.append(row)
return rows
def validate_row(row):
if row['high'] < row['low']:
raise ValueError('high is below low')
if row['volume'] < 0:
raise ValueError('volume is negative')
for field in ('open', 'high', 'low', 'close', 'adjusted_close'):
if row[field] != row[field]:
raise ValueError(f'{field} is NaN')
if __name__ == '__main__':
envelope = fetch_daily('IBM')
rows = parse_daily(envelope)
print(f'parsed {len(rows)} rows for {envelope["symbol"]}')
The example writes an immutable raw payload before parsing. It records retrieval time, request parameters and a checksum, so a parser change can replay the original response instead of requesting history again.
Rank #2
- Comes with secure packaging
- Easy to read text
- It can be a gift option
4. Normalize into stable records
Provider labels change; your internal schema should not. A practical price record contains:
provider,symbol,intervaland a UTC timestamp;- open, high, low, close and volume, with adjusted close when supplied;
- an explicit adjustment state such as raw, adjusted or adjusted-close-present;
- provider retrieval time, request identifier if available, and the raw-response checksum.
Convert timestamps to one documented timezone. Treat missing values deliberately: reject or quarantine malformed rows rather than silently converting them to zero. Enforce high greater than or equal to low, nonnegative volume and uniqueness on (provider, symbol, interval, timestamp, adjustment_state). Keep raw and adjusted series distinct so downstream returns are not accidentally mixed.
5. Store both raw and query-ready data
For a small deployment, SQLite is enough. Postgres is a good shared-worker choice; larger histories can be partitioned by provider and date in object storage or an analytical database. The important design is two workloads: immutable payloads for replay and a cleaned table for queries.
CREATE TABLE IF NOT EXISTS prices (
provider TEXT NOT NULL,
symbol TEXT NOT NULL,
interval TEXT NOT NULL,
timestamp TEXT NOT NULL,
open REAL NOT NULL,
high REAL NOT NULL,
low REAL NOT NULL,
close REAL NOT NULL,
adjusted_close REAL,
volume INTEGER NOT NULL,
adjustment_state TEXT NOT NULL,
provider_retrieved_at TEXT NOT NULL,
raw_sha256 TEXT NOT NULL,
PRIMARY KEY (provider, symbol, interval, timestamp, adjustment_state)
);
Insert with an upsert keyed by that primary key. A rerun then repairs a missing row instead of creating a duplicate. Store the code version and dependency lockfile alongside each run or batch manifest.
Free tools Windows power users keep installed
One-click scans. No signup required.
6. Add checkpoints, backfills and rate awareness
Incremental mode
After each successful symbol, write the last accepted timestamp and the provider checksum to a checkpoint table. On restart, read the checkpoint, request a bounded window that overlaps the last bar, validate it, and upsert. The overlap protects against a late correction while the unique key prevents duplication.
Backfill mode
Use a separate command for historical loading. Process symbols in small batches, lower concurrency than the live job, and record each completed range. A failed batch can resume from its last range without replaying the entire universe.
Respect limits and freshness
Throttle requests according to the provider plan, use exponential backoff for transient HTTP failures, and stop when a response signals a quota limit. Do not interpret an empty response as proof that a symbol has no data; classify it as an extraction result requiring review. Schedule end-of-day jobs after the relevant market session if your contract accepts end-of-day data. Intraday requirements need a source and entitlement that explicitly support them.
Rank #3
- Ideal for Gifting
- Ideal for a bookworm
- Comes with Proper Binding
7. Schedule and deploy the scraper
Containerize the reproducible job
FROM python:3.12-slim
WORKDIR /app
COPY requirements.txt .
RUN pip install --no-cache-dir -r requirements.txt
COPY scraper.py .
CMD ["python", "scraper.py"]
Pin the requests version in requirements.txt, inject ALPHA_VANTAGE_KEY as a secret, and mount durable storage for raw payloads and the database. Never bake credentials into the image.
Use a scheduler with an explicit command
A host cron entry can run a daily job after the session:
15 22 * * 1-5 cd /srv/stock-scraper && /srv/stock-scraper/.venv/bin/python scraper.py --mode incremental
For production, a managed scheduler and worker can provide retries and isolated logs. Evaluate each target on scheduler guarantees, secret handling, observability, persistent storage and recovery behavior. Whichever target you choose, make the process exit nonzero on a failed batch so alerts are triggered.
8. Make operation observable
Emit structured events for every run, symbol and request. At minimum monitor:
- HTTP failures, timeouts and provider-limit responses;
- empty or schema-changing payloads;
- stale newest timestamps versus the freshness contract;
- rows fetched, accepted, rejected and quarantined;
- duplicate-key rates and checkpoint age;
- run duration, retry count and raw-payload checksums.
Alert on a missing run, a sudden row-count drop or a timestamp that is older than your contract. After upgrading code, reconcile a sample of symbols against the provider and retain the comparison result.
Quick wins for a faster PC:
Scan for outdated or missing drivers - takes under a minuteDriver Scan →Clear out junk files and repair common Windows errorsFree Scan →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →9. Troubleshoot common failures
HTTP 401, 403 or an authentication message
Check that the key is present in the process environment, belongs to the intended account and has permission for the endpoint. Confirm that a secret manager did not inject a blank value.
A response contains a quota note instead of prices
Slow the worker, reduce concurrency and wait for the provider’s quota window. Keep the failed request in logs, but do not write the response as market data.
The parser cannot find a time-series field
Save the raw payload, inspect its top-level keys and compare them with the provider’s current schema. Route unknown shapes to quarantine and update the adapter with a fixture test before deploying.
Prices appear duplicated or jump after a split
Check the adjustment state and provider adjustment fields. Do not combine raw close values with adjusted close values in one return calculation; rerun the affected range after documenting the chosen policy.
The Tool Desk
Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →The job times out or loses progress
Use per-request timeouts, bounded batches and checkpoints. Rerun with the last successful checkpoint; the idempotent upsert makes recovery safe.
Data is fresh but cannot be published
Freshness does not grant redistribution rights. Re-check the provider agreement, exchange entitlements and whether your plan permits commercial or public use before exposing the series to customers.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.10. Cost, freshness and rights are product requirements
Provider price is only one part of operating cost. Include storage for immutable responses, scheduled-worker time, logs, retries and backfills. Alpha Vantage’s documented end-of-day default may be sufficient for research dashboards but not for an intraday trading interface; real-time or delayed U.S. data can require premium access. SEC filing data has different timing and usage considerations from price feeds. Record the entitlement and redistribution decision in the same contract that defines the interval.
Or skip the browser setup
If you also need a clean image of a rendered market dashboard, ScreenshotNeo can capture the page without maintaining a headless-browser worker. It accepts consent banners as a visitor and removes more than 60 known consent platforms, newsletter popups and chat widgets before capture; each cleanup step can be disabled. Only clean shots are billed: bot checks or CAPTCHAs, blank pages, timeouts, failed loads and cache hits cost nothing, and the response identifies the result with X-Page-Verdict and X-Billed headers.
Windows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallCrashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minuteOne GET request is enough (see the ScreenshotNeo API documentation):
curl -G 'https://api.screenshotneo.com/v1/shot' -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com/market-dashboard -o shot.webp
ScreenshotNeo also provides an MCP server for AI agents such as Claude, Cursor and other MCP clients, with tools for screenshots, page information and PDF capture. The free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 shots. See ScreenshotNeo for the available options and create a free account.
FAQ
Can I use SEC EDGAR for stock prices?
EDGAR is suited to company submissions and extracted XBRL facts. Use a market-data provider for OHLCV bars, and keep the two adapters separate.
Should I store adjusted or unadjusted prices?
Store the provider value you need for the product, but preserve an explicit adjustment state and retain the raw response so you can recompute or audit the choice.
Recommended Free Tools
Is SQLite suitable for a production scraper?
It is suitable for a small, single-worker pipeline. Move to Postgres or partitioned object storage when concurrent writers, larger history or analytical workloads require it.
How do I handle a provider switch?
Implement the new source behind the same fetch contract, map it into the existing normalized schema, replay retained raw data through validation, and compare overlapping symbols before changing the production source.
Frequently Asked Questions
Can I use SEC EDGAR for stock prices?
EDGAR is suited to company submissions and extracted XBRL facts. Use a market-data provider for OHLCV bars, and keep the two adapters separate.
Should I store adjusted or unadjusted prices?
Store the provider value you need for the product, but preserve an explicit adjustment state and retain the raw response so you can recompute or audit the choice.
Is SQLite suitable for a production scraper?
It is suitable for a small, single-worker pipeline. Move to Postgres or partitioned object storage when concurrent writers, larger history or analytical workloads require it.
How do I handle a provider switch?
Implement the new source behind the same fetch contract, map it into the existing normalized schema, replay retained raw data through validation, and compare overlapping symbols before changing the production source.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

