Driver FixRecommendedSound, Wi-Fi or graphics acting up? Check drivers firstFind missing or outdated drivers fast.Check DriversOctober DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsSlow PC?RecommendedPC slow today? Run a repair scan before it gets worseResolve common Windows issues and optimize system performance.Scan Now×
Skip to content
Beautiful Soup

How to Build a Real Estate Web Scraper—Legally and Reliably

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Start with permission, not code. Before collecting real estate listings, define which data you need, where it will come from, how often you will update it, and whether you will use it privately or publish it. Check the source’s terms and licensing rules; if you need listing data for an ongoing product, investigate an authorized local MLS feed or RESO Web API access. Only scrape HTML when the site’s rules and applicable permissions allow it.

For an authorized static page, Python’s Requests library can retrieve the page and Beautiful Soup can parse its HTML. Use Playwright only when permitted data requires browser rendering. Neither a browser nor a scraper grants permission to collect, retain, or republish listing data.

Define what your scraper is allowed to collect

A useful scraper begins with a collection contract: a short record of the source, allowed data, intended use, and operating limits. Decide these before choosing selectors or writing a schedule. Listing pages can show information that is not licensed for copying, storage, or redistribution.

  • Source: Record the exact domains and, where relevant, paths you are authorized to access.
  • Geography: Identify the market and jurisdiction. MLS access and data-use rules are local and provider-specific.
  • Fields: List only the fields your use case needs, such as a permitted listing identifier, asking price and currency, location fields, property type, bedroom and bathroom counts, area and units, and status.
  • Cadence: Set an update frequency consistent with the provider’s documented limits. There is no universal safe request rate.
  • Audience and use: Distinguish private analysis from a public listing search, commercial product, or redistribution.
  • Retention: Specify what you may store, for how long, who can access it, and what happens when a listing is removed or your permission ends.

Start with the smallest dataset that answers the actual question. Each field adds parsing, validation, licensing, and deletion obligations.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Check rights and access before making requests

Read the current terms for the particular site, API, or feed, along with any license or data-use agreement. Check the site’s robots.txt and honor crawler rules that apply to your process, but do not treat that file as permission. RFC 9309, the Robots Exclusion Protocol standard, expressly says that its rules are not a form of access authorization. A robots file and a site’s terms answer different questions.

Zillow is a concrete example, not a rule for every real estate site: its consumer terms prohibit automated queries, including scraping, spiders, robots, and crawlers, and prohibit bypassing access restrictions. Do not treat a publicly visible listing page as permission to automate requests or reuse its contents.

If you need authorized listing data for a continuing application or analysis, investigate the licensed route first. The Real Estate Standards Organization (RESO) says access to its Web API data is gained through local MLSs after agreeing to their data-use and licensing policies. RESO’s Web API uses OData V4 and can return JSON, but the local MLS controls access, credentials, available data, and permitted uses. A standardized interface does not mean universal or unrestricted access.

Some providers also offer their own approved API. Zillow’s developer API, for example, is for approved licensees and comes with limits on use, display, calls, and retention. Zillow says listings are published from MLS IDX feeds; it describes rental listings as coming through Zillow Feed Connect or Zillow Rental Manager. Its consumer pages, developer API, and feed arrangements are distinct routes with distinct conditions. Confirm which agreement applies to your use rather than assuming one route authorizes another.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Choose the acquisition route that fits your authorization

Route Access basis Best fit Main trade-off
Licensed MLS or RESO API/feed Local MLS approval, credentials, and data-use agreement Ongoing analysis or applications that need authorized listing data Access, fields, and permitted uses vary by MLS and license.
Provider-approved API Provider approval and its API terms A use case and data set covered by that API Scope, display, retention, and call limits may shape the design.
HTML parsing Site terms and other applicable permissions must allow collection Narrow collection from permitted, stable pages Page changes can break parsing, and visible content is not automatically reusable.
Browser automation The same permission required for any other collection method Permitted content that appears only after browser rendering More operational complexity; it does not authorize access or bypass restrictions.

For a licensed API, follow the provider’s authentication, query, pagination, refresh, and field rules. The example below is for a permitted HTML page only; it does not make a prohibited source permissible.

Build a small Python scraper for an authorized static page

This example fetches one permitted page, checks the HTTP result, parses listing cards, normalizes a few fields, and writes JSON Lines. It intentionally uses placeholder selectors: inspect the authorized page and replace them with its actual structure. If you have a licensed API or feed, use that interface instead of scraping its pages.

Rank #2
Sale
The Millionaire Real Estate Investor
  • Business & Economics
  • Real Estate

Install the dependencies

Use Python 3 and install Requests and Beautiful Soup in your project environment:

python -m pip install requests beautifulsoup4

Save the scraper

Replace LISTINGS_URL with a page you are permitted to collect from, and change the CSS selectors to match that page. The script does not follow listing links, circumvent access controls, or retry denied requests.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
import json
import logging
from datetime import datetime, timezone
from decimal import Decimal, InvalidOperation
from pathlib import Path

import requests
from bs4 import BeautifulSoup

LISTINGS_URL = "https://example.com/authorized-listings"
OUTPUT_PATH = Path("listings.jsonl")
TIMEOUT_SECONDS = 20

logging.basicConfig(level=logging.INFO, format="%(levelname)s %(message)s")


def clean_text(node):
    return node.get_text(" ", strip=True) if node else None


def parse_price(text):
    """Parse a simple displayed price; preserve unknown formats as None."""
    if not text:
        return None
    digits = "".join(ch for ch in text if ch.isdigit() or ch in ".,")
    # This deliberately handles common US-style values such as $425,000.
    # Adapt it to the source's locale and currency instead of guessing.
    digits = digits.replace(",", "")
    try:
        return str(Decimal(digits))
    except (InvalidOperation, ValueError):
        return None


def fetch_html(url):
    response = requests.get(
        url,
        headers={"User-Agent": "AuthorizedListingResearch/1.0"},
        timeout=TIMEOUT_SECONDS,
    )
    response.raise_for_status()
    return response.text


def extract_records(html, source_url):
    soup = BeautifulSoup(html, "html.parser")
    observed_at = datetime.now(timezone.utc).isoformat()
    records = []

    # These are example selectors, not selectors for any named property site.
    for card in soup.select(".listing-card"):
        title = clean_text(card.select_one(".listing-title"))
        price_text = clean_text(card.select_one(".listing-price"))
        address = clean_text(card.select_one(".listing-address"))
        listing_id = card.get("data-listing-id")

        if not listing_id:
            logging.warning("Skipping card without the permitted listing identifier")
            continue
        if not title and not address and not price_text:
            logging.warning("Skipping apparently empty listing card: %s", listing_id)
            continue

        records.append({
            "source_url": source_url,
            "listing_id": listing_id,
            "observed_at": observed_at,
            "title": title,
            "asking_price": parse_price(price_text),
            "price_source_text": price_text,
            "currency": None,  # Set only when the source makes it clear.
            "address_source_text": address,
        })
    return records


def append_jsonl(records, output_path):
    with output_path.open("a", encoding="utf-8") as output:
        for record in records:
            output.write(json.dumps(record, ensure_ascii=False) + "\n")


def main():
    try:
        html = fetch_html(LISTINGS_URL)
    except requests.Timeout:
        logging.error("Request timed out; check the source and permitted timeout policy")
        return 1
    except requests.HTTPError as exc:
        logging.error("Source returned an HTTP error: %s", exc)
        return 1
    except requests.RequestException as exc:
        logging.error("Request failed: %s", exc)
        return 1

    records = extract_records(html, LISTINGS_URL)
    if not records:
        logging.warning("No records parsed; verify page structure and permission")
        return 0

    append_jsonl(records, OUTPUT_PATH)
    logging.info("Wrote %d records to %s", len(records), OUTPUT_PATH)
    return 0


if __name__ == "__main__":
    raise SystemExit(main())

The parser keeps source text alongside a cautiously parsed price because a displayed value such as “Contact agent,” a range, or a non-US format cannot safely be converted by stripping punctuation. The example leaves currency unknown rather than assuming it. Extend the schema only with fields your grant permits and your source actually supplies.

Run it and inspect the output

  1. Save the file as scrape_listings.py and replace the example URL and selectors.
  2. Run python scrape_listings.py from the environment where the dependencies are installed.
  3. Open listings.jsonl. Each line should be a separate JSON object with an observation timestamp and source URL.
  4. Test against a small sample, compare extracted values with the authorized source, and confirm the output fields before scheduling any collection.

Normalize and validate records without inventing data

Keep a stable internal schema even when source pages differ. A useful starting point, if allowed by the source agreement, includes source and listing identifiers, observed time, asking price and currency, location fields, property type, bedroom and bathroom counts, area and its units, status, and source URL. Not every provider exposes or licenses all of these fields.

  • Store a provider’s identifier when permitted and use it for deduplication; do not assume addresses are unique identifiers.
  • Keep the source value or original unit when conversion could lose meaning. Do not infer missing bedrooms, currency, area units, or status.
  • Represent missing or unparseable values as unknown, not as zero or an estimated value.
  • Validate required fields before writing records, and log retrieval failures separately from parse failures.
  • Record when you observed a value. An observation timestamp is not proof that the source’s listing status is still current.

For a licensed API, reconcile additions, changes, and removals according to its documented update mechanism. For HTML, a changed or missing card may mean the layout changed, the record disappeared, or the page returned incomplete content; do not silently interpret every parse miss as a confirmed delisting.

Use Playwright only when permitted content needs a browser

Requests and Beautiful Soup work when the permitted page returns the data in its HTML. If that data appears only after ordinary browser rendering, Playwright can automate Chromium, WebKit, or Firefox with Python. Browser automation adds a browser installation and more moving parts, and does not change the access rights that apply to the source.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

For an authorized page, install Playwright and its Chromium browser:

python -m pip install playwright
python -m playwright install chromium

A minimal rendering check can save the resulting HTML for the same parser to process:

import asyncio
from pathlib import Path
from playwright.async_api import async_playwright

URL = "https://example.com/authorized-listings"

async def main():
    async with async_playwright() as p:
        browser = await p.chromium.launch()
        page = await browser.new_page()
        response = await page.goto(URL, wait_until="domcontentloaded", timeout=30000)
        if response is None or not response.ok:
            status = response.status if response else "no response"
            raise RuntimeError(f"Page did not load successfully: {status}")
        await page.locator(".listing-card").first.wait_for(timeout=10000)
        html = await page.content()
        Path("rendered.html").write_text(html, encoding="utf-8")
        await browser.close()

asyncio.run(main())

Use a selector that represents the authorized content you need; if it does not appear, investigate whether the page changed or failed to load. Do not add steps intended to evade CAPTCHA, bot checks, authentication restrictions, or other access controls. If access is denied or permission changes, stop collection and contact the provider.

Or skip the browser setup

ScreenshotNeo is a website screenshot API and MCP server, not a structured real estate listing feed. For a permitted page, it can return a screenshot or PDF without your managing a browser in this workflow. One GET request:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://example.com/authorized-listings -o shot.webp

See the ScreenshotNeo API documentation for request options. Cookie and consent banners are accepted and removed along with more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each of those steps can be turned off. Bot checks and CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and the response reports the page verdict and billing status in headers. Its MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients. The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 shots.

Those captures are visual outputs, not parsed property records; you still need an authorized data source and a separate extraction method for structured fields. Try the free ScreenshotNeo sign-up for 1,000 screenshots a month with no card.

Operate the pipeline conservatively

Put retrieval, parsing, validation, and storage in separate steps so failures are visible. Set explicit timeouts, check response status, log enough detail to diagnose a broken selector, and keep an eye on the provider’s documented usage limits. Avoid aggressive polling; no universal request rate is established here. A successful HTTP response is not proof that the response contains a complete page or that the collection is authorized.

  • Before each run: Confirm the source, permission, allowed fields, and any provider limits remain applicable.
  • During retrieval: Use timeouts and handle status and transport errors explicitly. Do not automatically hammer a source after denial, throttling, or repeated failures.
  • After parsing: Check required fields and record counts. Alert on sudden empty results or major schema changes instead of writing misleading data.
  • At storage: Limit access to collected data, keep provenance, and apply the agreement’s attribution, retention, and deletion requirements.
  • At publication: Verify that your license permits the exact display and audience. Do not assume listing content can be copied to another site.

These are source-specific obligations. Zillow’s developer API terms, for example, require immediate end-user delivery and prohibit retaining API data copies; Zillow’s consumer terms also restrict displaying its data elsewhere. Those conditions do not establish the rules for a different provider. Review the agreement that actually covers your source and use.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Troubleshoot common failures

The request times out or fails

Check the URL, network connection, and provider status, then use a bounded timeout consistent with the permitted workflow. A timeout is a failed retrieval, not an empty listing result. Do not turn repeated failures into rapid retries.

The server returns an error or denies access

Inspect the status code and stop if the response indicates denial or the provider’s limits. Confirm that your credentials, route, and authorization cover the requested access. Do not attempt to disguise traffic or work around an access restriction.

The script runs but extracts zero cards

The example selectors are placeholders. Inspect the HTML you are authorized to receive and verify the card and field selectors. If the content is absent from the returned HTML but appears in a permitted browser, consider Playwright; if the source denies automation, do not switch tools to bypass that decision.

Prices or areas parse incorrectly

Do not strip punctuation and assume a single locale. Preserve the displayed source value, determine the source’s number format and currency, and add a parser for that known format. Keep unrecognized values unknown until they can be validated.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Records duplicate or appear to vanish

Use the provider’s permitted listing identifier where available and retain observation times. A missing card can reflect a layout or retrieval problem as well as a changed listing; compare with the authorized update mechanism before treating it as a confirmed removal.

The page changes or returns incomplete content

Separate HTTP success from extraction success: check that expected elements and required fields exist before persisting data. Log a parse failure and review the authorized source. Do not fill gaps by guessing or silently reuse stale values as current ones.

Plan for reliability, performance, and cost

Static HTTP requests are generally simpler operationally than launching a browser for every page, while browser rendering can be necessary when permitted content is produced client-side. The best choice is the least complex method that reliably obtains the authorized data. No general speed, success-rate, or request-rate figure applies across sites and MLS providers.

Keep collection narrow, follow documented quotas, and use the provider’s supported change or refresh mechanism where available. Browser automation adds installation, execution, and page-wait management. A licensed feed or API may involve approval and agreement constraints but gives you a defined access route. Plan storage and refresh behavior around those terms rather than around how much data a scraper can technically fetch.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Cost cannot be estimated universally from the methods alone: access fees, usage limits, infrastructure, and retention conditions depend on the provider and deployment. Price those items from the actual MLS/API agreement and hosting plan before committing to a cadence or product design.

Frequently Asked Questions

Does a screenshot count as a listing-data record?

No. A screenshot is an image of a page, not a normalized record with fields such as price, address, and listing identifier. Use it for visual capture; use an authorized structured feed or a permitted extraction workflow for data.

Can I use the sample code on any listing site if I reduce the request rate?

No. Lower request frequency does not replace authorization. Check the specific source’s terms and licensing conditions before making automated requests.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Read next

Recommended PC Tool
Recommended PC Tool
Crashes, No Sound, or Screen Glitches?Free driver scan
Windows Errors? Fix Them Before They SpreadFree repair scan

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.