Reliable scraped data comes from treating extraction and data processing as separate stages. First parse a response into structured records; then normalize values, validate fields and domain rules, handle duplicates deliberately, and export or store only records that meet your criteria. In Scrapy, item pipelines provide a natural place for this post-extraction work. This guide builds that workflow in Python and explains how to control crawls and interpret robots.txt along the way.
Why processing scraped data needs its own stage
A selector returning text does not prove that the text is complete, correctly typed, or meaningful. A page redesign can make a selector return an empty string; a price may contain a currency symbol and thousands separator; a date may be ambiguous; and the same record may appear on more than one page. If those problems reach a database or analysis job unnoticed, downstream code has to guess what the values mean.
Keep site-specific extraction in the spider and reusable cleanup, validation, duplicate handling, and persistence in processing code. Scrapy spiders parse responses and yield key-value items; item pipelines receive those items sequentially and can pass them along or drop them. See the Scrapy building blocks and the Scrapy item pipeline documentation.
1. Specify the record before scraping
Write down the shape of an accepted record before selecting page elements. Required fields, optional fields, types, canonical formats, and a stable identity key are dataset decisions; there is no universal schema. For a product catalogue, for example, a record might include a source URL, product name, price amount, currency, and scrape timestamp. Keep a raw value as well when it is useful for auditing or reprocessing.
PC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minute#1 Best Overall
| Design question | Example decision |
|---|---|
| Which fields are mandatory? | A product name and canonical product URL are required; a description may be optional. |
| What type and format should each value have? | Store price as a decimal amount plus a separate currency code, not a formatted display string. |
| What makes two records the same? | Use a source-provided product ID or a canonical URL, if stable and unique for the dataset. |
| What context helps diagnose a record? | Retain the source URL and crawl-run identifier or timestamp where operationally useful. |
Do not choose a key casually: a title can change or be shared by multiple records. Decide what a collision means before the pipeline runs—discard a later copy, update an existing record, or flag it for review.
2. Extract structured items from responses
Scrapy supports CSS and XPath selection for HTML and XML. A spider callback should turn response content into a record with predictable keys, rather than leaving later stages to understand page markup. Extraction rules remain specific to the site; the example below uses illustrative selectors that you must adapt to the pages you are permitted to crawl.
Create a project with scrapy startproject catalog, then add a spider such as catalog/spiders/products.py:
import scrapy
class ProductsSpider(scrapy.Spider):
name = "products"
allowed_domains = ["example.com"]
start_urls = ["https://example.com/products"]
def parse(self, response):
for card in response.css(".product-card"):
yield {
"source_url": response.url,
"product_url": response.urljoin(
card.css("a.product-link::attr(href)").get()
),
"name_raw": card.css(".product-name::text").get(),
"price_raw": card.css(".price::text").get(),
}
next_page = response.css("a.next::attr(href)").get()
if next_page:
yield response.follow(next_page, callback=self.parse)
The selectors and domain above are examples, not verified selectors for a real site. A selector can silently stop matching after a layout change, so treat missing values as a data-quality signal rather than assuming the page has no record. Spider callbacks can yield both items and follow-up requests; the pipeline sees the items after extraction.
3. How do I clean data after web scraping?
Normalize consistently and deterministically. Typical field-level rules include trimming and collapsing whitespace, parsing numeric values into numeric types, standardizing dates, and using a declared unit or currency convention. Keep meaning intact: stripping punctuation from every field, for example, may corrupt identifiers or legitimate text. Preserve original values when they help explain a conversion or let you reprocess older data.
Here is a small, explicit normalizer for the example record. The price parser intentionally accepts only a simple format; adapt it to the site’s locale and currency conventions rather than guessing how ambiguous separators should be interpreted.
import re
from decimal import Decimal, InvalidOperation
def clean_text(value):
if value is None:
return None
cleaned = " ".join(value.split())
return cleaned or None
def parse_price(value):
text = clean_text(value)
if text is None:
return None
# Example only: supports a leading currency symbol and comma grouping.
normalized = re.sub(r"[^0-9.]", "", text.replace(",", ""))
try:
return Decimal(normalized)
except InvalidOperation:
return None
A production parser should define what inputs are allowed and how it handles signs, decimal separators, ranges, and currency. Do not silently convert an unparseable value to zero: zero is a real value, while parse failure is a different state.
4. How do I validate scraped data?
Validation should check presence and type first, then domain-specific rules. Examples include a parseable date, a nonnegative amount where the domain requires it, a URL with an expected scheme, or a status drawn from an allowed set. A failed record can be rejected, repaired by a documented transformation, or routed for review. Avoid quietly inventing missing values.
Quick wins for a faster PC:
Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Repair Windows errors before they cause bigger problemsFix Now →Scrapy’s item pipeline documentation describes checking required fields and dropping items that should not continue. The following pipeline demonstrates normalization, required-field validation, and duplicate rejection. Add the helper functions from the preceding section above the class in catalog/pipelines.py:
from scrapy.exceptions import DropItem
class CleanValidatePipeline:
def process_item(self, item, spider):
item["name"] = clean_text(item.pop("name_raw", None))
item["price"] = parse_price(item.pop("price_raw", None))
required = ("product_url", "name")
missing = [field for field in required if not item.get(field)]
if missing:
raise DropItem(f"missing required fields: {', '.join(missing)}")
if item["price"] is not None and item["price"] < 0:
raise DropItem("price must not be negative")
return item
class DuplicatePipeline:
def open_spider(self, spider):
self.seen_urls = set()
def process_item(self, item, spider):
key = item["product_url"]
if key in self.seen_urls:
raise DropItem(f"duplicate product URL: {key}")
self.seen_urls.add(key)
return item
In a Python source file, write the comparison operator as < in the code rather than the HTML-escaped form shown above. If placing code in HTML, encode angle brackets as entities for valid markup; if copying the logic into Python, use if item["price"] is not None and item["price"] < 0: with a literal less-than sign.
Rank #3
Enable the stages in catalog/settings.py; the order matters because the duplicate stage expects a normalized product_url:
ITEM_PIPELINES = {
"catalog.pipelines.CleanValidatePipeline": 100,
"catalog.pipelines.DuplicatePipeline": 200,
}
This in-memory duplicate set is suitable only for one spider process and one run: it is lost when the process exits and may use substantial memory for large crawls. For cross-run uniqueness, use a persistent store with a unique constraint or another deliberate persistence strategy.
5. How do I remove duplicates from scraped data?
Choose a stable key that reflects record identity, then define collision behavior. Scrapy’s documented duplicate-pipeline example uses an ID set and drops an item when that ID has already appeared. The URL-based example above follows the same principle, but URLs may need canonicalization if query parameters, trailing slashes, or redirects cause one record to have multiple addresses.
- Normalize the key before comparison using the same rule on every run.
- Decide whether the first or latest observed version wins, or whether a collision should be reviewed.
- For persistent datasets, enforce uniqueness where records are stored; an in-memory set cannot prevent duplicates across separate runs.
- Do not deduplicate by comparing every field unless that is genuinely the identity rule. A changed description should not necessarily create a new entity.
6. How do I store scraped data?
For straightforward output, Scrapy feed exports support JSON, CSV, and XML. Run the example spider and write accepted items to JSON Lines with scrapy crawl products -O products.jsonl; use an extension and format suited to your downstream workflow. Feed export is convenient when the output is a file. A pipeline is the more flexible point for custom transformations or database persistence.
For a database, map the validated item to a schema and use parameterized inserts or the database driver’s equivalent. Make writes idempotent where practical: a stable key plus an upsert or unique constraint can support reruns without accumulating copies. Include source context when it helps diagnose malformed or stale records, but avoid collecting fields that the task does not need.
Scrapy also documents using pipelines for database storage. Its overview covers feed exports and the framework’s broader crawling components. Ryan Mitchell’s Web Scraping with Python, 3rd Edition, listed by O’Reilly as a February 2024 publication, includes Scrapy, storage, normalized text, and cleaning dirty data.
Do these 3 things before closing this tab:
1Repair Windows errors before they cause bigger problems2Fix the driver behind crashes, sound loss and screen glitches3Clear out junk files and repair common Windows errors7. Monitor data quality by crawl run
Track counts that help distinguish a site change from a normal fluctuation: items extracted, records missing required fields, validation rejections, duplicate drops, and records successfully written. Keep those counts associated with a run identifier and, where useful, the source URL or crawl timestamp. A sudden shift can then prompt inspection of selectors or changed page content rather than silently degrading the dataset.
There is no universal rejection or duplicate threshold established by the framework documentation. Set alerts based on your own site’s history and the cost of bad records. Scrapy’s building blocks and extension mechanisms provide places to implement project-specific monitoring.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.8. Robots.txt and crawl controls
RFC 9309, the IETF Standards Track specification for the Robots Exclusion Protocol published in September 2022, says: “These rules are not a form of access authorization.” Robots.txt is a crawler coordination mechanism, not authentication or a security boundary. RFC 9309 specifies how crawlers handle retrieved, unavailable, and unreachable robots files; implementations should consult the RFC text rather than collapsing those cases into one universal rule.
Scrapy documents download delays, per-domain concurrency controls, and AutoThrottle. They are mechanisms for controlling request behavior, not a guarantee that any particular rate is acceptable to every site. Check the site’s applicable terms and policies, respect applicable robots rules, and choose conservative settings appropriate to the site and your workload. The Scrapy overview describes these controls.
The Tool Desk
Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Best Value
Or skip the browser setup
For a visual capture of a page rather than structured records, ScreenshotNeo can return a screenshot or PDF through one request. It does not replace selectors, validation, or a data pipeline, but can be useful when a workflow also needs a rendered-page image. For example, this cURL request saves a WebP capture:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
See the ScreenshotNeo API documentation for request options. Cookie banners, popups, and chat widgets are removed before the shot; bot checks, blank pages, and failed loads are never billed. An MCP server lets AI agents take screenshots, and 1,000 screenshots a month are free with no card; paid plans start at $5 for 3,000. Learn about ScreenshotNeo, then sign up for 1,000 free screenshots a month with no card.
Troubleshooting common pipeline failures
- Required fields suddenly go missing: inspect a recent response and confirm the CSS or XPath selectors still match; check whether content is present in the response Scrapy receives before changing validation to accept blanks.
- Prices become
None: compare raw source values with the parser’s accepted format. Handle locale-specific decimal and grouping rules explicitly instead of stripping characters until a number appears. - Valid items are dropped as duplicates: inspect the chosen key and canonicalization. Confirm that distinct records do not legitimately share it and that URL variants are handled consistently.
- Duplicates return on the next run: the sample set exists only in process memory. Move uniqueness enforcement to persistent storage or another cross-run mechanism.
- Output contains unexpected types or formats: inspect the item after each pipeline stage and ensure transformations have a single documented canonical representation before export.
- The crawl puts too much load on a site: reduce concurrency, set a delay, or use AutoThrottle as appropriate; no one request rate is established as safe for all sites.
Frequently Asked Questions
Does robots.txt grant permission to scrape a site?
No. RFC 9309 explicitly distinguishes crawler instructions from access authorization; it is not a substitute for authentication or applicable site terms.
Should I keep the original scraped value after normalization?
Keep it when auditability or future reprocessing matters, and define which field is raw versus canonical so downstream users do not confuse them.
Can an item pipeline fix every extraction problem?
No. Pipelines can clean, validate, reject, deduplicate, or persist items, but site-specific selectors and response handling still belong to extraction logic.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




