Turn scraper output into an RSS feed by normalizing every result into a record, mapping those records to RSS 2.0 <item> elements, validating the XML, and serving the latest valid document from a stable HTTPS URL. The durable design is more important than the XML itself: use canonical links and immutable identifiers, escape untrusted text, publish atomically, and keep the previous good feed when a run fails.
The data flow: page to feed
An RSS pipeline has five stages:
- Fetch: request target pages and handle retries, timeouts, robots rules, and rate limits.
- Extract: locate title, canonical URL, summary, publication time, and a source identifier.
- Normalize: convert dates and text to consistent values, remove malformed characters, and reject unusable records.
- Serialize: create one RSS 2.0 channel containing repeated items.
- Publish: validate the document, write it atomically, and serve it at a URL readers can keep.
The channel normally has a title, description, and link. Each item commonly has a title, link, description, publication date, and guid. RSS readers use the identifier to decide whether an item is new, so a stable guid matters even when a title or summary changes.
Design the normalized record first
Do not generate XML directly from arbitrary scraper dictionaries. Define a small internal record and make the feed generator consume only that record.
Recommended fields
| Field | Purpose | Required? |
|---|---|---|
title |
Human-readable item heading | Yes |
link |
Canonical HTTPS page URL | Yes |
description |
Short plain-text or sanitized HTML summary | Yes for a useful feed |
published |
Timezone-aware publication timestamp | Strongly recommended |
guid |
Immutable source key, usually the canonical URL | Yes for reliable deduplication |
Normalize and reject early
- Trim whitespace and collapse accidental repeated spaces in titles and summaries.
- Resolve relative links, then prefer the page’s canonical link over a tracking URL.
- Convert every timestamp to a timezone-aware value and serialize it in RFC 822/RSS form, such as
Tue, 29 Sep 2026 12:00:00 +0000. - Remove XML-illegal control characters. Treat scraped HTML as untrusted input; either convert it to text or sanitize an allow-list of tags.
- Discard records without a stable link, title, or identifier. Logging the discarded URL and reason makes extraction failures diagnosable.
- Deduplicate by
guidbefore sorting. Sort newest first, using the source timestamp; use crawl time only when the source has no publication date and label that policy in your documentation.
Generate RSS 2.0 in a custom Python pipeline
The following complete example accepts normalized dictionaries, escapes XML safely with Python’s standard library, keeps the newest 50 records, and writes a feed. It uses the canonical URL as the identifier and preserves HTML-free descriptions.
Quick wins for a faster PC:
Clear out junk files and repair common Windows errorsFree Scan →Scan for outdated or missing drivers - takes under a minuteDriver Scan →#1 Best Overall
- Includes Raspberry Pi 5 with 2.4Ghz 64-bit quad-core CPU (8GB RAM)
- Includes 128GB Micro SD Card pre-loaded with 64-bit Raspberry Pi OS, USB MicroSD Card Reader
- CanaKit Turbine Black Case for the Raspberry Pi 5
- CanaKit Low Noise Bearing System Fan
- Mega Heat Sink - Black Anodized
from datetime import datetime, timezone
from email.utils import format_datetime
from pathlib import Path
import html
import re
import xml.etree.ElementTree as ET
OUTPUT = Path("public/feed.xml")
MAX_ITEMS = 50
# Replace this function with your scraper's output.
def scrape_and_normalize():
return [
{
"title": "Example announcement",
"link": "https://example.com/news/announcement",
"description": "A short summary from the source page.",
"published": datetime(2026, 9, 29, 12, 0, tzinfo=timezone.utc),
"guid": "https://example.com/news/announcement",
}
]
def clean_text(value):
value = re.sub(r"[\x00-\x08\x0b\x0c\x0e-\x1f]", "", str(value or ""))
return " ".join(value.split())
def normalize(rows):
seen, result = set(), []
for row in rows:
title = clean_text(row.get("title"))
link = clean_text(row.get("link"))
description = clean_text(row.get("description"))
guid = clean_text(row.get("guid")) or link
published = row.get("published")
if not title or not link or not guid or guid in seen:
continue
if published and published.tzinfo is None:
published = published.replace(tzinfo=timezone.utc)
seen.add(guid)
result.append({"title": title, "link": link,
"description": description, "guid": guid,
"published": published})
result.sort(key=lambda x: x["published"] or datetime.min.replace(tzinfo=timezone.utc),
reverse=True)
return result[:MAX_ITEMS]
def build_feed(items):
rss = ET.Element("rss", {"version": "2.0"})
channel = ET.SubElement(rss, "channel")
for tag, value in {
"title": "Example site updates",
"description": "Latest items collected from Example site.",
"link": "https://example.com/",
}.items():
ET.SubElement(channel, tag).text = value
for item in items:
node = ET.SubElement(channel, "item")
for tag in ("title", "link", "description", "guid"):
child = ET.SubElement(node, tag)
child.text = item[tag]
ET.SubElement(node, "guid", {"isPermaLink": "true"}).text = item["guid"]
if item["published"]:
ET.SubElement(node, "pubDate").text = format_datetime(item["published"])
return ET.ElementTree(rss)
items = normalize(scrape_and_normalize())
tree = build_feed(items)
ET.indent(tree, space=" ")
OUTPUT.parent.mkdir(parents=True, exist_ok=True)
tmp = OUTPUT.with_suffix(".xml.tmp")
tree.write(tmp, encoding="utf-8", xml_declaration=True)
tmp.replace(OUTPUT) # atomic replacement on the same filesystem
print(f"wrote {len(items)} items to {OUTPUT}")
There is one deliberate correction to make in production: the loop above writes guid in the common element loop and then again with isPermaLink. Remove "guid" from that tuple when using the script, leaving the explicitly attributed guid element:
for tag in ("title", "link", "description"):
child = ET.SubElement(node, tag)
child.text = item[tag]
ET.SubElement(node, "guid", {"isPermaLink": "true"}).text = item["guid"]
ElementTree escapes ampersands, angle brackets, and quotes in text and attributes. If you intentionally include HTML in description, sanitize it first and use a CDATA strategy only after confirming your downstream readers handle it.
Use Scrapy Feed Exports when the scraper is already Scrapy
Scrapy’s Feed Exports feature is the shortest path when items already flow through a Scrapy spider. Its documented serializers include JSON, JSON Lines, CSV, XML, Pickle, and Marshal; storage backends include the local filesystem, FTP, S3, and standard output. Define item fields in your spider, then configure an XML feed destination:
Rank #2
- Includes Raspberry Pi 4 4GB Model B with 1.5GHz 64-bit quad-core CPU (4GB RAM)
- Includes Pre-Loaded 32GB EVO+ Micro SD Card (Class 10), USB MicroSD Card Reader
- CanaKit Premium High-Gloss Raspberry Pi 4 Case with Integrated Fan Mount, CanaKit Low Noise Bearing System Fan
- CanaKit 3.5A USB-C Raspberry Pi 4 Power Supply (US Plug) with Noise Filter, Set of Heat Sinks, Display Cable - 6 foot (Supports up to 4K60p)
- CanaKit USB-C PiSwitch (On/Off Power Switch for Raspberry Pi 4)
# settings.py
FEEDS = {
"public/feed.xml": {
"format": "xml",
"encoding": "utf8",
"overwrite": True,
}
}
Feed Exports handles serialization and storage, but your spider still owns canonical URLs, stable identifiers, date normalization, deduplication, and content cleaning. If you need a strict RSS 2.0 shape or custom extensions, a pipeline that builds the XML explicitly gives more control. Scrapy’s documentation describes storing scraped data as one of the most frequently required scraper features.
Do these 3 things before closing this tab:
1Scan for outdated or missing drivers - takes under a minute2Clear out junk files and repair common Windows errors3Fix the driver behind crashes, sound loss and screen glitchesValidate before a reader sees the feed
Use a parser-based check in every scheduled run. Universal Feed Parser can parse a remote URL, local filename, or raw feed string, making it suitable both for deployment checks and tests.
import feedparser
feed = feedparser.parse("public/feed.xml")
if getattr(feed, "bozo", 0):
raise ValueError(f"invalid XML: {feed.bozo_exception}")
if not feed.feed.get("title") or not feed.feed.get("link"):
raise ValueError("channel title and link are required")
seen = set()
for entry in feed.entries:
if not entry.get("title") or not entry.get("link"):
raise ValueError("every item needs title and link")
identifier = entry.get("id") or entry["link"]
if identifier in seen:
raise ValueError(f"duplicate identifier: {identifier}")
seen.add(identifier)
if "published_parsed" in entry and entry.published_parsed is None:
raise ValueError("unparseable publication date")
print(f"validated {len(feed.entries)} items")
Checks worth automating
- The document is well-formed XML and declares UTF-8.
- The channel has title, description, and link.
- Each item has title, link, description, date when available, and a unique identifier.
- Links are absolute HTTPS URLs and return an expected status.
- No item exceeds your chosen summary length, and no malformed control character remains.
- The item count is within an operational limit, preventing an extraction bug from publishing thousands of entries.
Publish safely and refresh predictably
Stable URL and content type
Serve the file at a permanent HTTPS address such as https://example.com/feed.xml with an XML content type (commonly application/rss+xml or application/xml). Do not rotate the URL when you change hosts; readers poll the address they originally subscribed to.
Atomic replacement
Write a temporary file, validate that file, then rename it over the previous feed on the same filesystem. If scraping or validation fails, leave the previous valid document untouched. This prevents readers from receiving a truncated file during a deployment.
Scheduling and retention
Run at an interval appropriate to the source’s update rate and terms of access. Retain logs containing run time, fetched pages, rejected records, item count, and validation errors. Keep enough history to diagnose a disappearing item, but do not republish an old item with a new identifier.
Common failures and fixes
| Symptom | Likely cause | Fix |
|---|---|---|
| Reader reports malformed XML | Unescaped text or an illegal control character | Use an XML library, sanitize scraped text, and run parser validation before replacement. |
| Every refresh shows duplicates | guid changes with crawl time or title edits |
Derive it from the canonical URL or another immutable source key. |
| Dates appear wrong or disappear | Naive datetimes or inconsistent formats | Attach the source timezone, convert consistently, and serialize with format_datetime. |
| Feed is empty after a failed run | Output was truncated before scraping completed | Write a temporary file and atomically replace only after validation. |
| Descriptions show markup or scripts | Raw scraped HTML was copied into XML | Convert to text or sanitize an explicit HTML allow-list; never trust source markup. |
| Scrapy exports the wrong fields | Item schema and feed settings do not match | Declare stable item fields and inspect one generated XML file before scheduling. |
| Source blocks requests | Excessive rate, missing headers, or access controls | Respect the site’s rules, throttle requests, cache where appropriate, and handle retries with backoff. |
Performance, reliability, and cost decisions
- Fetch cost: avoid refetching unchanged pages when the source supports conditional requests; cache parsed results and only regenerate the feed when records change.
- Memory: stream or batch large crawls, then sort a bounded set of recent records. A feed with a practical item limit is faster for readers than an unbounded archive.
- Reliability: separate extraction from publication. A successful HTTP crawl is not proof of valid XML; validation is a release gate.
- Ordering: use source publication time for editorial order. If it is absent, document your fallback rather than pretending crawl time is publication time.
- Security: keep credentials out of feed content and logs, enforce URL allow-lists for internal crawlers, and treat all scraped fields as hostile input.
Or skip the browser setup
If your scraper first needs clean page images or PDFs—for example, to archive a visual version alongside each feed item—ScreenshotNeo provides a website screenshot API and MCP server. A single request returns PNG, JPEG, WebP, or PDF; it accepts cookie banners before capture and removes more than 60 known consent platforms, newsletter popups, and chat widgets. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing result.
Use the API call shown in the ScreenshotNeo documentation:
Rank #4
- Broadcom BCM2711, quad-core Cortex-A72 (ARM v8) 64-bit SoC @ 1. 5GHz
- 2. 4 GHz and 5. 0 GHz IEEE 802. 11b/g/n/ac wireless LAN, Bluetooth 5. 0, BLE
- 2 × USB 3. 0 ports, 2 x USB 2. 0 Ports
- 2 × micro HDMI ports supproting up to 4Kp60 video resolution
- Micro SD card slot for loading operating system and data storage
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Equivalent Python and Node.js requests:
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
ScreenshotNeo also exposes an MCP server for AI agents, with take_screenshot, get_page_info, and capture_pdf tools. Every plan includes its features; 1,000 screenshots per month are free with no card, and paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.
Frequently asked questions
Should a guid always be a URL?
No. A canonical URL is convenient when it is immutable; otherwise use another permanent source key and set isPermaLink accordingly.
Can RSS contain full articles?
It can, but a short sanitized description that links to the source is simpler and reduces unsafe markup. Follow the source site’s terms and your readers’ expectations.
Best Value
- Includes Raspberry Pi 5 16GB with 2.4Ghz 64-bit quad-core CPU (16GB RAM)
- Includes 128GB Micro SD Card pre-loaded with 64-bit Raspberry Pi OS, USB MicroSD Card Reader
- CanaKit Turbine Black Case for the Raspberry Pi 5
- CanaKit Low Noise Bearing System Fan
- Mega Heat Sink - Black Anodized
How many items should the feed contain?
There is no universal limit. Choose a bounded window that covers the reader’s polling interval and your source’s update frequency, then enforce it consistently.
Is Atom required instead?
No. RSS 2.0 is sufficient for the structure described here. A parser such as Universal Feed Parser can validate RSS and Atom if you later support both.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




