Do these 3 things before closing this tab:
1Clear out junk files and repair common Windows errors2Fix the driver behind crashes, sound loss and screen glitches3Repair Windows errors before they cause bigger problemsShort answer: Start with permission, not code. A property website’s terms, an MLS agreement, or an official licensed API determines whether you may collect, store, display, or resell its listings. If collection is authorized, use a documented feed where possible; otherwise build a conservative collector that preserves raw responses, normalizes fields, deduplicates records, and records every status change.
This guide shows a repeatable workflow for authorized collection of prices, addresses, bedrooms, bathrooms, area, status, and related metadata. It also explains why visible-in-a-browser does not mean reusable, and how to avoid brittle parsers and accidental redistribution.
Permission is the first technical requirement
Real-estate data has several overlapping rights layers. A page can be publicly readable while its terms prohibit automated collection or reuse. Realtor.com’s Move Network Terms of Use prohibit scraping, screen scraping, database scraping, and automated collection without express written permission. Zillow’s terms restrict reproducing or publicly displaying listing data and images on another service except where explicitly permitted. Treat those rules as contractual requirements, not suggestions.
Photos, descriptions, agent details, logos, virtual tours, and video can each have separate copyright, trademark, privacy, or contractual restrictions. NAR Policy Statement 7.85 says listing brokers should own, or have authority to license, photographs, images, graphics, audio/video, descriptions, remarks, pricing, and other listing details submitted to an MLS. Public visibility alone does not grant a copying or resale license.
Windows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallCrashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minute#1 Best Overall
What to verify before writing a collector
- Current Terms of Use and any data-license or partner agreement.
- Official API, MLS, broker-feed, or syndication documentation.
robots.txtand published crawler preferences. These communicate operational preferences but do not replace contractual permission.- Permitted fields, geographic coverage, refresh frequency, rate limits, attribution, branding, storage duration, and redistribution rights.
- Whether your use is internal research, a public search product, lead generation, resale, or display on behalf of a client.
Keep a written authorization record containing the account or contract, API key owner, allowed fields, request limits, attribution wording, retention period, and takedown process. Recheck the live terms and API documentation before launch and whenever your use changes.
Choose a source before choosing a scraper
An official API or licensed MLS/broker feed is usually the cleanest production source because its schema, limits, and rights are defined in a contract. HTML extraction can be useful for an authorized source that has no suitable feed, but it creates ongoing parser and policy-maintenance work.
| Source option | Strengths | Risks and questions |
|---|---|---|
| Official API | Documented fields, stable authentication, explicit update mechanism | License scope, branding, quotas, cost, and redistribution terms still apply |
| MLS or broker feed | Structured listing data and defined professional-use rules | Membership, geography, field restrictions, attribution, and retention requirements |
| Authorized HTML collection | Can capture fields not exposed by a feed | Layout changes, slower requests, anti-bot controls, and higher maintenance |
| Unapproved copying | None that makes it suitable for production | Contract, copyright, privacy, security, and access-enforcement exposure |
Zillow documents API products for home valuation, property details, homes posted for sale, mortgage, and professional reviews or directory data. Those products are subject to API terms, licensing, and branding requirements; an API key is not blanket permission to republish everything returned.
Define the dataset and permitted use
Write a one-page data specification before making requests. State the geography, sale or rental scope, fields, refresh interval, retention period, intended audience, and whether records or images will be resold or publicly displayed. Mark every field as “licensed,” “derived,” or “internal” so a developer cannot accidentally expose restricted content.
Useful listing fields
- Source name and source listing ID.
- Canonical listing URL and observed-at timestamp in UTC.
- Address components: street, locality, region, postal code, and country.
- Price, currency, billing period for rentals, and any displayed qualifiers.
- Bedrooms, bathrooms, floor area, lot area, and their units.
- Property type, listing status, first-seen and last-seen timestamps.
- Broker or agent fields only when the license permits storage and display.
- Image or video references only when reuse is expressly allowed.
Keep the original text for each parsed value alongside a normalized value. For example, preserve “$2,450/mo” while storing numeric rent 2450, currency USD, and period month. This makes later parser corrections auditable.
Inspect access paths and obtain authorization
- Read the target’s current terms, API documentation, robots instructions, and partner or MLS rules.
- Ask the owner for written permission when automated HTML collection is not clearly allowed. Specify domains, paths, fields, request rate, retention, and display or resale plans.
- Prefer the official endpoint or feed. Record authentication method, pagination, update timestamps, error semantics, and quota headers.
- Define a stop condition: repeated authorization failures, a changed policy, a CAPTCHA, or a material schema change must pause collection for review.
Never bypass authentication, CAPTCHAs, paywalls, access controls, or technical restrictions. Do not rotate identities or proxies to evade a block. If a source denies automated access, use its licensed channel or stop.
Build a respectful collector for an authorized HTML source
Request behavior
- Identify your application with a clear user agent and a contact address where the agreement allows it.
- Use low, bounded concurrency and exponential backoff for transient failures.
- Cache successful responses and use conditional requests such as
If-None-MatchorIf-Modified-Sincewhen supported. - Set connection and total timeouts. A hung request should not consume every worker.
- Store the response body, status code, headers needed for auditing, source URL, and fetch time for the retention period allowed by the contract.
Runnable Python example
The following example is intentionally conservative. It expects an authorized URL, looks first for Schema.org JSON-LD, and writes a normalized record. Install dependencies with python -m pip install requests beautifulsoup4.
import json, re, sys, time
from datetime import datetime, timezone
from urllib.parse import urlparse
import requests
from bs4 import BeautifulSoup
URL = sys.argv[1]
USER_AGENT = "AuthorizedListingCollector/1.0 ([email protected])"
session = requests.Session()
session.headers.update({"User-Agent": USER_AGENT, "Accept": "text/html,application/xhtml+xml"})
for attempt in range(4):
try:
response = session.get(URL, timeout=(10, 45))
if response.status_code in (429, 500, 502, 503, 504):
if attempt == 3:
response.raise_for_status()
time.sleep(2 ** attempt)
continue
response.raise_for_status()
break
except requests.RequestException:
if attempt == 3:
raise
time.sleep(2 ** attempt)
soup = BeautifulSoup(response.text, "html.parser")
items = []
for tag in soup.select('script[type="application/ld+json"]'):
try:
value = json.loads(tag.string or tag.get_text())
items.extend(value if isinstance(value, list) else [value])
except (json.JSONDecodeError, TypeError):
continue
listing = next((x for x in items if isinstance(x, dict) and
x.get("@type") in ("Residence", "SingleFamilyResidence", "Apartment", "House")), {})
offer = listing.get("offers") if isinstance(listing.get("offers"), dict) else {}
address = listing.get("address") if isinstance(listing.get("address"), dict) else {}
def number(value):
if value is None:
return None
match = re.search(r"[0-9]+(?:[.,][0-9]+)?", str(value))
return float(match.group(0).replace(",", "")) if match else None
record = {
"source_url": response.url,
"observed_at": datetime.now(timezone.utc).isoformat(),
"listing_id": listing.get("identifier"),
"address_raw": address.get("streetAddress"),
"city": address.get("addressLocality"),
"region": address.get("addressRegion"),
"postal_code": address.get("postalCode"),
"price_raw": offer.get("price"),
"price": number(offer.get("price")),
"currency": offer.get("priceCurrency"),
"beds": number(listing.get("numberOfBedrooms")),
"baths": number(listing.get("numberOfBathroomsTotal")),
"area_raw": listing.get("floorSize"),
"property_type": listing.get("@type"),
}
print(json.dumps(record, ensure_ascii=False, indent=2))
JSON-LD is not guaranteed to be present or complete. If you must use visible HTML, select stable semantic attributes rather than presentation classes, and retain the source value when a normalized conversion fails. Do not assume that a missing field means zero.
The Tool Desk
Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Equivalent request patterns
For an authorized JSON endpoint, keep pagination and authentication supplied by its documentation. The shell pattern below preserves the response for inspection:
curl --fail-with-body --retry 3 --retry-delay 2
-H "User-Agent: AuthorizedListingCollector/1.0"
-H "Authorization: Bearer $API_TOKEN"
"$API_ENDPOINT/listings?updated_since=2026-09-01"
-o listings.json
In Node.js 18 or newer, use the built-in fetch API and abort slow requests:
const controller = new AbortController();
const timer = setTimeout(() => controller.abort(), 45000);
try {
const res = await fetch(`${process.env.API_ENDPOINT}/listings`, {
headers: {
Authorization: `Bearer ${process.env.API_TOKEN}`,
"User-Agent": "AuthorizedListingCollector/1.0"
},
signal: controller.signal
});
if (!res.ok) throw new Error(`${res.status} ${await res.text()}`);
const data = await res.json();
console.log(JSON.stringify(data));
} finally {
clearTimeout(timer);
}
Normalize, deduplicate, and preserve history
Normalization rules
- Store currency as an ISO-style code supplied by the source; never infer it solely from a symbol when the market is ambiguous.
- Convert area to a chosen internal unit only after recording the displayed unit and original text.
- Normalize whitespace and address casing for matching, but retain the original address for audit and display where licensed.
- Represent unknown, not applicable, and zero as different states.
- Keep parser version and source response hash on every observation.
Deduplication and relistings
Use a source listing ID as the primary key when available. Without one, combine the canonical URL with a cautiously normalized address; never treat that combination as permanent because a property can be relisted or its URL can change. Keep an alias table for old URLs and IDs. Store field-level history or immutable snapshots so a price or status change can be explained rather than overwritten.
Validate and monitor the pipeline
Validation should run before a record reaches a public database. Require the fields your product actually needs, check numeric ranges, verify currency and units, measure geocoding confidence, and reject impossible transitions such as a negative price or a status change that skips required states in your business rules.
Compare a sample of normalized records with the source page on each release. Track parser errors, empty-result rates, HTTP status distribution, latency, and the age of the newest observation. Alert on a sudden increase in missing prices or a changed JSON-LD shape. Keep raw evidence only for the period allowed by the source agreement, then delete it on schedule.
Publish only what the license allows
Before launch, map every output field to a permission in your contract. Preserve required source attribution, broker or agent notices, branding language, correction procedures, and takedown contacts. Do not republish photos, descriptions, logos, phone numbers, or email addresses merely because a browser displayed them. If public display is not licensed, expose derived analytics or internal identifiers instead of the underlying content.
Performance, reliability, and cost decisions
- Freshness: choose a refresh interval that matches the agreement and the listing lifecycle; event or delta feeds are more efficient than repeatedly downloading every page.
- Throughput: bounded concurrency protects the source and makes your own failure recovery predictable. Queue work and cap retries.
- Cache: cache unchanged responses and use validators. This lowers bandwidth and helps stay within quotas.
- Storage: raw pages improve auditability but may carry additional retention and copyright obligations. Encrypt them and delete them on schedule.
- Cost: compare API calls, feed fees, geocoding, storage, and engineering time. A seemingly free HTML source can be expensive to maintain when layouts or policies change.
- Reliability: design for partial failure. A timeout on one listing should not discard a complete batch, and a source outage should be visible rather than converted into “listing removed.”
Troubleshooting common failures
403, 401, or repeated access denials
Cause: missing authorization, an expired token, or a source policy that disallows your method. Fix: verify the contract and credentials, reduce request rate, and contact the owner. Do not attempt to evade the control.
429 rate-limit responses
Cause: your request rate or quota is too high. Fix: honor Retry-After, add exponential backoff, reduce concurrency, cache results, and request a documented quota increase if available.
HTML loads but fields are empty
Cause: data is rendered by JavaScript, moved into JSON-LD, or the layout changed. Fix: use the documented API or embedded structured data first; add parser tests and alert on missing-field rates. Do not silently publish empty values.
Prices or areas are wrong
Cause: locale-specific separators, rental periods, currency symbols, or unit conversions. Fix: retain the raw value, parse with an explicit locale and unit map, and require a known currency before conversion.
Duplicate properties appear
Cause: multiple source IDs, changed URLs, or a relisting. Fix: reconcile by source ID first, then use cautious address matching with manual review for conflicts. Keep both source records when ownership or status differs.
A listing disappears
Cause: delisting, temporary outage, pagination drift, or a parser failure. Fix: mark the record as unobserved only after repeated successful collection cycles and a healthy parser check; preserve the last observation and reason code.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Best Value
Or skip the browser setup
If your authorized workflow needs a rendered page rather than an API, ScreenshotNeo provides a website screenshot API and MCP server. One GET request can return PNG, JPEG, WebP, or PDF. It can accept cookie or consent banners as a visitor and remove more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be disabled. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing result. Its MCP tools—take_screenshot, get_page_info, and capture_pdf—work with Claude, Cursor, and other MCP clients.
Use the documented parameters and confirm that your source agreement permits screenshots. The complete options include full-page capture with lazy-image loading, CSS-selector element capture, dark mode, device presets or custom viewports, retina scale, PDF paper size and margins, custom CSS or JavaScript, pre-capture clicks, hidden selectors, selector or network-idle waits, request and resource blocking, custom headers and cookies, user-agent, authorization, timezone, geolocation, transparent backgrounds, resizing, chosen cache TTL, signed public-image links, asynchronous jobs with signed webhooks, bulk capture of up to 100 URLs per call, usage API, and OpenAPI specification. Parameter names used by other screenshot APIs are accepted to ease migration.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
See the ScreenshotNeo documentation for authentication, output formats, and options. The same call from Python is:
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
And in Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`${res.status} ${await res.text()}`);
Replace the example URL with a property page you are authorized to capture. ScreenshotNeo includes 1,000 screenshots per month free with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.
Free tools Windows power users keep installed
One-click scans. No signup required.
FAQ
Frequently Asked Questions
Can I combine records from several licensed feeds?
Yes, but keep source provenance on every field, document precedence when values conflict, and honor the strictest attribution, retention, and redistribution rule that applies to the combined output.
What should happen when two sources report different prices?
Do not overwrite one value silently. Store both observations with timestamps and source IDs, flag the conflict for review, and select a display value only under a documented business rule.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

