Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
Process a scraping dataset as an evidence pipeline, not as a one-off cleanup script: preserve every raw response and its provenance, profile the data, ingest in bounded batches, normalize without destroying originals, deduplicate with a declared identity key, validate a written contract, quarantine failures, and publish a curated Parquet layer with lineage. This approach keeps bad rows explainable and lets you rerun transformations when your parser or rules change.
1. Preserve a raw layer before cleaning
The raw layer is part of the dataset. Save the exact downloaded file or response body before parsing or rewriting it. Alongside each object, record:
- Source URL and the final URL after redirects, if available.
- Retrieval timestamp with an explicit timezone.
- HTTP status and relevant response headers.
- Parser and scraper code versions.
- A content hash, such as SHA-256, for byte-level identity.
Use immutable paths, for example raw/source=shop/date=2026-09-29/run=001/. Write a separate manifest (JSON, CSV, or a database row) that maps each raw object to its metadata. Never overwrite a raw response when a parser discovers an error; create a new staged or curated output instead.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Why this matters
Pages change, selectors drift, and a later normalization rule may be wrong. With the raw body, you can reparse without crawling again, compare two retrievals, and explain exactly which input produced a published row.
#1 Best Overall
2. Profile the dataset before transforming it
Run a profile on a representative sample first, then repeat the same checks over every batch. At minimum inspect row count, column names, null rates, duplicate rates, text encoding, inferred types, and representative values. Look specifically for:
- Columns whose names vary by crawl (for example,
product_nameversusname). - Numbers containing currency symbols, thousands separators, or locale-specific decimal marks.
- Dates mixing formats or time zones.
- URLs with fragments, tracking parameters, or inconsistent host casing.
- HTML error pages that were saved with a nominally successful status.
A small sample is useful for discovering selectors and type problems, but it is not a quality result. Sampling can miss a malformed page or a rare category, so aggregate counts and validation results over the complete run.
3. Ingest large files in bounded batches
For small and medium files, pandas is convenient. For large CSV exports, use usecols, explicit dtypes, and chunksize so memory use is bounded. Compression can be inferred from the filename. Parse non-standard dates with to_datetime() after loading, using an explicit format or timezone policy.
Recommended Free Tools
import pandas as pd
columns = ['url', 'title', 'price', 'retrieved_at', 'status_code']
dtypes = {
'url': 'string',
'title': 'string',
'price': 'string',
'status_code': 'Int64'
}
for batch_no, frame in enumerate(pd.read_csv(
'raw/pages.csv.gz',
usecols=columns,
dtype=dtypes,
chunksize=50_000,
compression='infer'
)):
frame['retrieved_at'] = pd.to_datetime(
frame['retrieved_at'], errors='coerce', utc=True
)
# profile_batch(frame, batch_no) and transform_batch(frame) go here
frame.to_parquet(f'staged/batch={batch_no:05d}.parquet', index=False)
Choose a chunk size that leaves headroom for parsing and validation. Do not concatenate every chunk into one giant DataFrame just to write the result; aggregate counters and write staged batches independently.
4. Normalize without losing the original value
Normalization should make equivalent values comparable while retaining the source representation whenever conversion can be lossy. Keep pairs such as price_raw/price_numeric, date_raw/retrieved_at, and url_raw/canonical_url.
Field names and text
Map source columns to a versioned internal schema, trim surrounding whitespace, standardize Unicode where appropriate, and define how empty strings differ from nulls. Do not silently replace an unparseable value with zero or an empty string.
URLs
Canonicalization is a policy decision. A conservative policy lowercases the scheme and host, removes a fragment, and preserves the query string unless your site-specific rules prove that certain parameters are tracking-only. Keep the raw URL beside the canonical value so the policy can be revised.
from urllib.parse import urlsplit, urlunsplit
def canonicalize_url(value):
if value is None:
return None
raw = str(value).strip()
if not raw:
return None
parts = urlsplit(raw)
scheme = parts.scheme.lower()
host = (parts.hostname or '').lower()
port = parts.port
netloc = host
if port and not ((scheme == 'http' and port == 80) or (scheme == 'https' and port == 443)):
netloc = f'{host}:{port}'
return urlunsplit((scheme, netloc, parts.path or '/', parts.query, ''))
frame['url_raw'] = frame['url']
frame['canonical_url'] = frame['url'].map(canonicalize_url)
Dates, numbers, units, and booleans
Declare a timezone policy and parse with an explicit format when the source is known. For prices and measurements, store the unit and currency separately; a value of 10 is not comparable until you know whether it means dollars, euros, kilograms, or something else. For booleans, map only documented source values and send unknown tokens to quarantine.
Rank #2
5. Deduplicate with an identity key that matches the meaning
A URL alone is often insufficient: the same page can change between crawls. Choose the key according to the question your dataset answers:
| Use case | Identity key | What it preserves |
|---|---|---|
| Current catalog snapshot | canonical_url |
One current record per page. |
| Historical page states | canonical_url + retrieval_date |
One record per page per observation date. |
| Entity-centric data | Stable product or article ID | Records even when URLs change. |
| Exact response inventory | Content hash | Byte-identical bodies, regardless of URL. |
Declare which duplicate you keep and why. pandas supports drop_duplicates(subset=..., keep=...); keep='first' and keep='last' are different business decisions, while keep=False marks every member of a duplicate group for review.
key = ['canonical_url', 'retrieval_date']
frame['retrieval_date'] = frame['retrieved_at'].dt.date
before = len(frame)
frame = frame.sort_values('retrieved_at').drop_duplicates(key, keep='last')
dropped = before - len(frame)
Count dropped rows and retain that count in the run manifest. If records from multiple chunks can collide, deduplicate after all staged batches are combined logically (for example, with a distributed job or an external sort), not only inside each chunk.
6. Define and enforce a data contract
Write the contract before publishing. It should state required columns, data types, nullability, allowed ranges, category sets, and uniqueness rules.
| Expectation | Example rule | Failure action |
|---|---|---|
| Schema | canonical_url is a string and retrieved_at is UTC datetime. |
Quarantine the row and report the expectation name. |
| Required fields | URL and retrieval timestamp are non-null. | Reject; do not invent values. |
| Range | HTTP status is between 100 and 599; price is non-negative when present. | Quarantine for source or parser review. |
| Categories | Status is one of the documented values. | Quarantine unknown tokens. |
| Uniqueness | The declared identity key is unique in the curated partition. | Apply the documented deduplication policy, then fail if collisions remain. |
Great Expectations can attach repeatable expectations to filesystem data assets and batches and can work with pandas or Spark connections. Whether you use it, another validator, or custom code, run the same contract on every batch and store the result with the run metadata. Validate representative CSV or Parquet batches before promoting an entire run.
7. Quarantine failures instead of hiding them
Write invalid rows to a quarantine location with the failed expectation name, a human-readable reason, and the run identifier. Keep the original fields so an engineer can reproduce the decision. Track counts for each failure class: malformed dates, missing required fields, out-of-range numbers, unknown categories, and duplicate-key collisions.
Do not use a blanket coercion such as errors='coerce' without measuring its effect. If malformed dates become null, count them, inspect samples, and decide whether to fix the parser, accept nullability, or reject the records. Promotion should fail when rejection exceeds a threshold defined for the dataset, rather than silently producing a smaller table.
Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchWindows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstall8. Publish a curated analytical layer
Apache Parquet is an open source, column-oriented data file format designed for efficient data storage and retrieval. Use it for the cleaned layer when downstream work is analytical, while retaining raw files for interoperability and forensic review.
| Layer | Recommended format | Purpose |
|---|---|---|
| Raw evidence | Original response, CSV, or compressed files | Exact replay and audit. |
| Staged | Batch Parquet or typed tables | Bounded transformations and restartable work. |
| Curated | Partitioned Parquet | Stable, typed analytical access. |
| Shared serving | Warehouse or lakehouse table | Recurring jobs, access control, and team queries. |
Partition by a stable date or source key only when it matches query patterns. Over-partitioning creates many tiny files; under-partitioning makes selective reads expensive. Record the schema version and transformation version in table metadata or the run manifest.
9. Track lineage and make reruns safe
For each run, record source URLs, crawl timestamp, scraper code version, parser version, schema version, transformation version, input and output row counts, rejection counts, duplicate counts, validation results, and the locations of raw, staged, curated, and quarantine data. Give each run an immutable identifier.
Make transformations deterministic where possible. A rerun should read the same raw objects, produce comparable counts, and explain any differences caused by a changed code or schema version. Promote a curated partition only after validation passes; otherwise leave the previous published partition intact.
Do these 3 things before closing this tab:
1Scan for outdated or missing drivers - takes under a minute2Repair Windows errors before they cause bigger problems3Fix the driver behind crashes, sound loss and screen glitches10. Respect crawl controls before collecting data
Read the target site’s robots.txt for the actual user agent before fetching, and apply its directives together with rate limits, authentication rules, terms, and applicable law. Python’s urllib.robotparser.RobotFileParser answers whether a user agent may fetch a URL under the published robots file; it is a parser, not a legal-permission engine.
from urllib.robotparser import RobotFileParser
rp = RobotFileParser('https://example.com/robots.txt')
rp.read()
if not rp.can_fetch('my-scraper/1.0', 'https://example.com/catalog'):
raise RuntimeError('robots.txt disallows this fetch for the declared user agent')
Revisit crawl policies when a target changes. A previously permitted path can become disallowed, and an authenticated or rate-limited endpoint may have additional requirements.
Choosing tools and scale
| Situation | Approach | Controls to emphasize |
|---|---|---|
| Exploration or small-to-medium files | pandas | usecols, explicit dtypes, and chunked reads. |
| Volume or concurrency exceeds one machine | Spark or another distributed engine | Partitioning, shuffle size, and distributed validation. |
| Repeatable quality gates | Great Expectations or equivalent | Versioned expectations attached to batches. |
| Recurring shared analytics | Warehouse or lakehouse | Access control, scheduled runs, and raw-data retention. |
Move to a distributed engine because the workload requires it, not simply because the file is large. A well-designed chunked pandas job can process a file that does not fit in memory, while a distributed job adds operational complexity and shuffle costs.
Performance, reliability, and cost practices
- Read only needed columns during each stage and use typed columns to reduce memory pressure.
- Write restartable batches so a failed job resumes from the last completed batch instead of recrawling.
- Keep network fetching separate from parsing and publishing; this lets you reprocess raw data without spending crawl time again.
- Measure rows in, rows out, bytes read, bytes written, rejection counts, and processing time per batch.
- Retain raw data according to your legal, privacy, and storage policies; remove or restrict sensitive fields deliberately rather than during an opaque cleanup step.
Troubleshooting common failures
Out-of-memory errors
Reduce chunksize, select fewer columns, set explicit dtypes, and avoid keeping prior chunks in a list. If a global operation such as deduplication still exceeds one machine’s memory, use an external sort or distributed engine.
Unexpected duplicate counts
Inspect the identity key and canonicalization policy. If the same URL changes over time, add retrieval date or a content hash. If duplicates cross chunk boundaries, run a global deduplication step.
Rank #4
Many null dates or numbers
Sample the raw strings, identify every format and locale, and parse with explicit rules. Preserve the raw field, count conversion failures, and quarantine values that do not match a supported format.
Schema validation fails after a source redesign
Compare the raw response and parser version with the last successful run. Update the mapping and schema version deliberately, then backfill from retained raw files; do not weaken the expectation merely to make the run pass.
Curated output is smaller than expected
Check quarantine and rejection counts first, then duplicate counts and required-field failures. A smaller output is acceptable only when the loss is measured, reviewed, and consistent with the contract.
Free tools Windows power users keep installed
One-click scans. No signup required.
Or skip the browser setup
If your workflow needs screenshots as evidence for a scrape, ScreenshotNeo provides a website screenshot API and MCP server. One GET request returns PNG, JPEG, WebP, or PDF. It accepts cookie and consent banners as a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each cleanup step can be turned off. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and the response identifies the page verdict and billing status in X-Page-Verdict and X-Billed headers.
cURL (see the ScreenshotNeo API documentation):
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com -o shot.webp
Python:
import requests
r = requests.get(
'https://api.screenshotneo.com/v1/shot',
params={'access_key': 'YOUR_API_KEY', 'url': 'https://example.com'},
timeout=90
)
r.raise_for_status()
open('shot.webp', 'wb').write(r.content)
Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://example.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`HTTP ${res.status}`);
await Bun.write('shot.webp', res);
The service also supports full-page captures with lazy images loaded, CSS-selector element capture, dark mode, 12 device presets and custom viewports, retina scale, PDF paper size/margins/orientation/page ranges, HTML/CSS-to-image, custom JavaScript and CSS, clicks before capture, hidden selectors, waits for selectors, delays or network idle, request and resource blocking, custom headers/cookies/user agents/Authorization, timezone and geolocation, transparent backgrounds, resizing, configurable-TTL caching, signed public-image links, asynchronous jobs with signed webhooks, bulk capture of 100 URLs per call, a usage API, and an OpenAPI specification. Existing parameter names used by other screenshot APIs are accepted to ease migration. An MCP server exposes take_screenshot, get_page_info, and capture_pdf to Claude, Cursor, and other MCP clients.
Every plan includes every feature. The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account to begin.
Frequently Asked Questions
How should I version a schema change?
Assign a new schema version, keep the prior contract available for historical partitions, and rerun retained raw inputs through the new mapping so differences are attributable to the version change.
When is CSV still the right published format?
Keep CSV when interoperability or direct human inspection is the main requirement; use Parquet for the curated analytical layer and retain the original CSV or response files for replay.
Should one crawl run mix pages fetched at different times?
Only if the dataset explicitly models observation time. Otherwise, partition or label retrieval dates so consumers do not mistake a mixed-time run for a single snapshot.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

