October DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsWindows FixRecommendedWindows errors stealing your time? Find the fix fastScan stability, cleanup and performance issues.Fix NowOctober DealsAmazon USDeal season is back - check today's better picksAmazon US: current deals, useful picks and tech finds.See Picks×
Skip to content
data automation

How to Automate Real Estate Data Extraction: Scheduled Property Listings

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

To keep property listings current, get authorized access to the relevant MLS or listing provider, use its RESO Web API where available, and run a scheduled incremental sync keyed to each record’s modification timestamp. Save the provider’s original payload, map approved fields into your own schema, and advance the sync watermark only after every page has been processed successfully. RESO defines standards; it does not provide listing data, credentials, or universal API access.

Choose an authorized listing source first

The source determines which records you may access, how often you may request them, and what you may store or display. Start with the MLS or provider serving the relevant market, not a consumer-facing property page. A page scraper can break when a site changes its layout and may conflict with access or licensing terms; an authorized feed gives you a defined interface and applicable terms to work from.

  • United States: Ask the local MLS about access to its RESO Web API or an approved provider. RESO says it does not provide MLS data, property records, or access to other organizations’ APIs. The MLS or provider controls access and terms.
  • Canada: REALTOR.ca DDF provides authenticated RESO/OData access to Property and related listing resources. Access and permissions are controlled by brokerage owners.
  • Other commercial sources: Zillow documents APIs for property details, postings, valuation, mortgage, reviews, and directories. Approval and branding or display conditions apply; confirm the current terms before designing around an endpoint.

Coverage and availability vary by market and provider. RESO’s certification page, updated September 28, 2026, reports 484 functioning MLS systems in the United States and says at least 90% of MLSs in the industry have RESO-certified Web API services. Those figures describe the certification page’s reported U.S. systems and industry share; they do not mean every MLS offers access to every applicant.

Confirm the rights and operating limits

Before implementation, ask the provider for API documentation, credentials, permitted use, rate limits, supported resources, field metadata, pagination rules, and any change or deletion mechanism. Get explicit answers about internal analytics, public display, attribution, photo and media use, retention, derivative fields, and redistribution. A field being present in a response does not by itself grant the right to retain or republish it.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

RESO Web API is the modern standards-based transport; RETS is deprecated and no longer supported by RESO. If a provider still offers only RETS, ask about its migration path rather than making a new integration depend on a deprecated transport.

Use RESO Web API and Data Dictionary concepts

The RESO Web API uses RESTful design, OData V4, JSON, OAuth, metadata, and live queries. RESO’s Data Dictionary standardizes concepts and field names for resources such as Property, Member, Office, and Media. That shared vocabulary can reduce custom mapping, but it does not guarantee that every MLS exposes the same fields or values. Inspect the specific endpoint’s metadata and contract before relying on a field.

A practical integration has four layers:

  1. Provider adapter: handles OAuth, endpoint paths, filters, paging, and provider-specific field names.
  2. Raw archive: retains each permitted source payload with retrieval time and source context for debugging and replay.
  3. Normalized store: maps the approved subset of provider fields into your stable application schema.
  4. Consumer interface: serves only fields, images, and history your agreement allows the intended audience to use.

Keep raw and normalized data separate. This lets you revise a mapping without losing the original response, while preserving a traceable record of what the provider returned.

Design a safe incremental sync

Do an initial bounded backfill, then request only records changed since the last fully successful run. The key is a provider-supported modification timestamp, commonly ModificationTimestamp, plus the provider’s stable record key. RESO describes replication as ongoing requests for recent changes and supports replication and queueing-oriented capabilities; the exact filters and deletion behavior remain provider-specific.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Rank #2
Sale
The Millionaire Real Estate Investor
  • Business & Economics
  • Real Estate

Backfill, watermark, and overlap

  1. Choose a limited first scope, such as a permitted geography, status set, or date interval. Do not start by downloading every historical record unless the provider has approved it and your storage rights allow it.
  2. Persist the provider’s stable listing key and the most recent successfully completed modification timestamp. Store timestamps in UTC and preserve the original provider value where practical.
  3. For each incremental run, begin slightly before the saved watermark. An overlap of a few minutes can help recover updates affected by clock skew or timing boundaries. Deduplicate by stable listing key and modification timestamp.
  4. Advance the watermark only after every page has been read and all accepted rows are committed. If a page fails, retry the run from its previous watermark rather than skipping ahead.

Pagination and idempotent writes

Follow the provider’s pagination mechanism, including any returned continuation link or cursor; do not infer that a short page means the result set is complete unless the API says so. Upsert each record using the provider’s stable key so replaying a page does not create duplicates. If multiple updates to a listing arrive, retain the version with the latest valid modification timestamp according to provider semantics.

Status changes and removals

Model status as a changing property of a listing, not as a reason to discard the record. Track transitions such as active, pending, sold, or withdrawn when the feed supplies them and your use permits retention. Deletions and removals differ among feeds: process provider-declared tombstones, queue messages, or reconciliation results as documented. Do not assume that a listing absent from one response has been deleted.

Python example: pull pages and upsert records

This example shows the sync mechanics for an OData-style endpoint. It requires Python 3 and requests (python -m pip install requests). Set the endpoint, bearer token, and field names to the values documented by your provider. The example assumes the endpoint accepts an OData timestamp filter, returns an @odata.nextLink when another page exists, and includes a stable key plus a modification timestamp. Confirm those details against your actual feed; resource paths, OAuth flows, filters, timestamps, and deletion handling are not universal.

Save as sync_listings.py, set the environment variables, then run python sync_listings.py. The local SQLite database stores a raw JSON payload and a normalized key/timestamp pair. It advances its watermark only after all pages are processed.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
import json
import os
import sqlite3
import time
from datetime import datetime, timedelta, timezone
from urllib.parse import urlparse

import requests

ENDPOINT = os.environ["LISTINGS_ENDPOINT"]  # Provider's Property resource URL
TOKEN = os.environ["LISTINGS_TOKEN"]
KEY_FIELD = os.getenv("LISTING_KEY_FIELD", "ListingKey")
MODIFIED_FIELD = os.getenv("LISTING_MODIFIED_FIELD", "ModificationTimestamp")
DB_PATH = os.getenv("LISTINGS_DB", "listings.sqlite3")
OVERLAP_MINUTES = int(os.getenv("OVERLAP_MINUTES", "5"))


def utc_now():
    return datetime.now(timezone.utc)


def parse_time(value):
    if not value:
        raise ValueError("Missing modification timestamp")
    parsed = datetime.fromisoformat(value.replace("Z", "+00:00"))
    if parsed.tzinfo is None:
        raise ValueError(f"Timestamp lacks timezone: {value}")
    return parsed.astimezone(timezone.utc)


def connect():
    db = sqlite3.connect(DB_PATH)
    db.execute("CREATE TABLE IF NOT EXISTS state (name TEXT PRIMARY KEY, value TEXT NOT NULL)")
    db.execute("""CREATE TABLE IF NOT EXISTS listings (
        listing_key TEXT PRIMARY KEY,
        modified_utc TEXT NOT NULL,
        raw_json TEXT NOT NULL,
        retrieved_utc TEXT NOT NULL
    )""")
    db.commit()
    return db


def read_watermark(db):
    row = db.execute("SELECT value FROM state WHERE name='watermark'").fetchone()
    return parse_time(row[0]) if row else None


def save_record(db, record, retrieved):
    key = record.get(KEY_FIELD)
    modified = record.get(MODIFIED_FIELD)
    if not key or not modified:
        raise ValueError(f"Record lacks {KEY_FIELD} or {MODIFIED_FIELD}")
    modified_utc = parse_time(modified)
    db.execute("""INSERT INTO listings(listing_key, modified_utc, raw_json, retrieved_utc)
        VALUES (?, ?, ?, ?)
        ON CONFLICT(listing_key) DO UPDATE SET
        modified_utc=excluded.modified_utc,
        raw_json=excluded.raw_json,
        retrieved_utc=excluded.retrieved_utc
        WHERE excluded.modified_utc >= listings.modified_utc""",
        (str(key), modified_utc.isoformat(), json.dumps(record), retrieved.isoformat()))
    return modified_utc


def fetch_page(session, url, params=None):
    for attempt in range(5):
        response = session.get(url, params=params, timeout=(10, 60))
        if response.status_code in (401, 403):
            response.raise_for_status()  # Stop: credentials or access need operator action.
        if response.status_code == 429 or response.status_code >= 500:
            if attempt == 4:
                response.raise_for_status()
            delay = min(2 ** attempt, 30)
            retry_after = response.headers.get("Retry-After")
            time.sleep(float(retry_after) if retry_after and retry_after.isdigit() else delay)
            continue
        response.raise_for_status()
        return response.json()
    raise RuntimeError("Retry limit reached")


def main():
    db = connect()
    previous = read_watermark(db)
    run_started = utc_now()
    # First run needs a provider-approved bounded backfill filter instead.
    if previous is None:
        raise SystemExit("Set up an approved initial backfill before the first incremental run.")
    start = previous - timedelta(minutes=OVERLAP_MINUTES)
    # Confirm the timestamp literal syntax and filterable field with your provider.
    params = {"$filter": f"{MODIFIED_FIELD} gt {start.isoformat().replace('+00:00', 'Z')}"}
    session = requests.Session()
    session.headers.update({"Authorization": f"Bearer {TOKEN}", "Accept": "application/json"})
    url, page_params = ENDPOINT, params
    page_count = rows = 0
    newest = previous
    try:
        while url:
            payload = fetch_page(session, url, page_params)
            records = payload.get("value")
            if not isinstance(records, list):
                raise ValueError("Expected OData JSON response with a value array")
            retrieved = utc_now()
            with db:
                for record in records:
                    changed = save_record(db, record, retrieved)
                    newest = max(newest, changed)
                    rows += 1
            page_count += 1
            next_link = payload.get("@odata.nextLink")
            if next_link:
                # Do not forward credentials to an unexpected host in a provider-supplied link.
                if urlparse(next_link).netloc != urlparse(ENDPOINT).netloc:
                    raise ValueError("Pagination link changed host; verify provider response")
            url, page_params = next_link, None
        with db:
            db.execute("INSERT INTO state(name,value) VALUES('watermark',?) "
                       "ON CONFLICT(name) DO UPDATE SET value=excluded.value",
                       (newest.isoformat(),))
        print(json.dumps({"started_utc": run_started.isoformat(), "pages": page_count,
                          "rows_received": rows, "watermark_utc": newest.isoformat()}))
    except Exception:
        # The old watermark remains available for a safe replay.
        raise
    finally:
        db.close()


if __name__ == "__main__":
    main()

Before production, add provider-specific field mapping and validation, a bounded first-run backfill, structured logs, alerting, and the feed’s documented removal or queue mechanism. The sample deliberately does not invent a universal OAuth token exchange or deletion rule. Do not increase request frequency beyond the provider’s limits.

Schedule, observe, and recover the job

Choose a cadence from the feed’s rate limits and the freshness your application actually needs. A 5–15 minute interval is an example only when the provider permits it; it is not a universal recommended or guaranteed interval. For a simple Linux deployment, a cron entry such as */10 * * * * /path/to/venv/bin/python /path/to/sync_listings.py attempts a run every ten minutes. Use a process lock or scheduler concurrency control so a slow run cannot overlap the next one.

Record at least the following for every attempt:

  • Start and completion time, endpoint/resource, filter window, and watermark before and after.
  • Pages requested, rows received, rows upserted, rejected records, and whether the run completed.
  • HTTP status and retry count, without logging access tokens or sensitive headers.
  • Provider-specific warnings, malformed fields, and removal events.

Alert on a late or failed run, repeated identical watermarks, unexpected zero-row runs, rising rejection counts, or a sudden change in page volume. Zero changes may be normal in a quiet market, so make alerts conditional on expected activity and use a reconciliation process rather than treating every empty result as proof of failure.

Retry and replay policy

Retry transient network errors, rate limits, and server errors with exponential backoff, respecting a provider’s Retry-After instruction where present. Stop and alert on authentication or authorization failures rather than retrying indefinitely. Keep failed or malformed records in a dead-letter queue with enough context to replay them after correcting the mapper. Preserve raw payloads only as long and in the manner your agreement allows.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Reconcile missed updates

Incremental syncs should be paired with periodic wider-window reconciliation. Re-read an approved historical interval, compare stable keys and modification timestamps, and process removals only according to the provider’s documented method. This helps recover from outages, imperfect timestamp boundaries, or a job that was offline longer than its normal overlap window.

Handle data changes and permissions deliberately

Build a field allowlist from the provider agreement and the endpoint metadata. Map to your internal schema explicitly instead of passing every source field through to users. Record when a source value is missing or rejected; do not silently substitute a different field with similar wording. The same RESO field name does not guarantee identical local availability or semantics.

Media needs special care. A listing payload may contain media references, but access to a URL does not establish permission to store the image, create derivative images, or redistribute it. Apply the provider’s photo, attribution, display, and retention terms to both the URL and any downloaded copy. Make status and removal handling consistent across search indexes, caches, and downstream exports so withdrawn records do not linger where they are no longer permitted.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Performance, reliability, and cost trade-offs

Incremental replication is generally more efficient than repeatedly downloading the full listing set: it limits requests and writes to recent changes. Its reliability depends on completing all pages, preserving a replayable watermark, deduplicating overlap, and reconciling a wider window. Full refreshes may be simpler for a small approved dataset, but can increase request volume and do not remove the need to follow rate limits and data rights.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Plan capacity around the provider’s page limits, response sizes, permitted cadence, and peak update volume. Keep database writes idempotent, bound network timeouts, and make the job safe to resume. Total cost is provider-specific: confirm access fees, usage or request limits, storage costs, support arrangements, and the engineering cost of operating the pipeline. No universal MLS API price or freshness SLA applies.

Troubleshooting common sync failures

  • 401 or 403: The token may be expired, mis-scoped, or unauthorized for the resource. Verify the provider’s OAuth flow, permissions, and account status; stop repeated retries until corrected.
  • 400 on the filter: The field may not be filterable, the timestamp literal may be wrong for that implementation, or the resource may use different metadata. Check its endpoint documentation and metadata rather than guessing another field.
  • Duplicate records: The upsert key may not be stable or unique in the provider’s scope. Confirm the correct key and include any provider-required scope identifier; deduplicate overlapped windows by key and modification time.
  • Missing updates: Check whether the watermark advanced before all pages completed, whether the clock/timezone was handled correctly, and whether the overlap is sufficient for the documented feed behavior. Replay from the last successful watermark and reconcile an approved wider window.
  • Pagination stops early: Follow the returned cursor or next link exactly and check for a failed later page. Do not construct page offsets unless the provider documents them.
  • Unexpected zero rows or stale data: Inspect the filter window, endpoint/resource, auth scope, scheduler logs, and most recent successful watermark. Compare with a provider-approved test query before treating an empty response as a genuine absence of changes.
  • Records disappear: Determine whether the feed provides tombstones, status changes, or a separate reconciliation path. Absence from an incremental result is not by itself a deletion signal.
  • 429 or 5xx responses: Back off, follow Retry-After when supplied, and reduce frequency or concurrency if required by the provider. Alert if retries exhaust rather than silently marking the run complete.

Or skip the browser setup

For the actual listing sync, use the authorized feed above; a screenshot is not structured listing data and does not replace MLS or provider permission. If you also need a visual page capture as an audit artifact, [ScreenshotNeo](/a) is a website screenshot API and MCP server. It can capture a page image or PDF, but it is not a substitute for an authorized listing API.

One-call cURL example (replace the URL with a page you are authorized to capture):

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

See the ScreenshotNeo API documentation. Cookie/consent banners are accepted like a visitor and 60+ known consent platforms, newsletter popups, and chat widgets can be removed before capture; each step can be turned off. Bot checks/CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, with verdict and billing information in response headers. Its MCP server gives AI agents tools for screenshots, page information, and PDFs. The free plan includes 1,000 shots a month with no card; paid plans start at $5 for 3,000. Sign up for 1,000 free screenshots a month, with no card required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Implementation checklist

  • Get credentials and written use terms from the MLS or approved provider.
  • Verify resource metadata, stable keys, timestamp filters, pagination, OAuth, and removal behavior.
  • Bound the backfill; persist raw permitted payloads and normalized approved fields separately.
  • Use overlapping incremental pulls, idempotent upserts, retries, and a watermark committed after full completion.
  • Log run health, alert on failures and anomalies, and reconcile a wider window periodically.
  • Apply retention, display, attribution, media, and redistribution rules to every downstream store and consumer.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Leave a Reply

Your email address will not be published. Required fields are marked *

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Read next

Recommended PC Tool
Recommended PC Tool
PC Slower Than It Used to Be?Free scan - under a minute
Crashes, No Sound, or Screen Glitches?Free driver scan

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.