Start with an official API or licensed feed. If you have permission to collect from HTML pages, build a conservative pipeline that respects the publisher’s terms and robots.txt, extracts structured data where possible, preserves the source and retrieval time, and validates changing prices and availability before using them. Stop if the publisher prohibits automated collection or blocks your requests; do not try to defeat a CAPTCHA or other access control.
Choose the data source before writing a scraper
First look for an official API, partner feed, or other licensed source. Read its documentation for permitted uses, available fields, authentication, quotas, geographic coverage, update frequency, and licensing terms. An API is not automatically unrestricted: its terms and limits still govern what you may collect and how you may use it. APIs can also give data owners more control over third-party collection and help them identify unauthorized scraping, as Canada’s privacy regulator notes.
| # | Preview | Product | Price | |
|---|---|---|---|---|
| 1 |
|
The Proxy Playbook: The Complete Guide to Proxy Servers: How to Source, Test, and Scale Residential,... | $29.95 | Buy on Amazon |
| 2 |
|
How to Host your own Web Server | $15.60 | Buy on Amazon |
If you cannot use an API or feed, inspect the site’s terms and robots.txt before making requests. Digital.gov describes robots.txt as a file that instructs crawlers about which parts of a site they should or should not access. Treat a disallow rule, an explicit prohibition in the terms, a CAPTCHA, or a comparable technical barrier as a reason not to proceed with that source. Robots.txt is not a substitute for permission, and an allow rule does not override terms or privacy obligations.
GDPR applies when scraping involves processing personal data, including collection, storage, organization, or retrieval. CNIL says web scraping is not, by itself, prohibited under the GDPR; that does not mean every collection or subsequent use is lawful. Minimize personal data, document the source and purpose, use reliable information, and validate it. For legal questions about a particular source or use, get advice appropriate to the relevant jurisdiction.
Recommended Free Tools
#1 Best Overall
Design a shared listing record without losing source meaning
Travel, event, and property listings have overlapping fields, but they are not interchangeable. Keep a source-specific record as well as normalized fields, so that transforming a price, date, or status never erases what the publisher actually stated.
| Listing type | Fields to capture | Important modeling detail |
|---|---|---|
| Travel and lodging | Property identity, address, amenities, coordinates; room or accommodation type; stay dates, occupancy, price, currency, and terms | Schema.org distinguishes the lodging business, the accommodation or room, and the offer. Associate dates, occupancy, price, currency, and booking terms with the offer, not as timeless facts about the hotel. |
| Events | Unique event URL, name, start and end times, location, organizer, ticket URL, price, currency, availability, and sale timing | Keep ticket price and availability tied to a particular event and capture time. Google’s event guidance calls for a unique URL and accurate name, start date, and location, and says ticket prices should include service charges and fees. |
| Real estate | Listing URL and ID, property type, sale or lease status, price and currency, bedrooms, bathrooms, floor and lot size where supplied, year built, address, coordinates, broker or agent, and listing/update timestamps | Preserve the source’s property facts and listing identity. Schema.org Accommodation examples include bedrooms, bathrooms, floor size, year built, address, latitude, and longitude; not every listing publishes every field. |
Schema.org data may appear as JSON-LD, Microdata, or RDFa. Prefer it to brittle visual selectors when it is present and relevant, but treat it as publisher-supplied data to validate—not as a guarantee of accuracy, completeness, or permission to collect.
Build the collection pipeline in controlled stages
- Discover permitted pages. Use the API or feed first. Otherwise enumerate only pages you are permitted to access, using documented links or sitemaps where appropriate. Avoid guessing hidden endpoints or expanding into unrelated areas of a site.
- Check access rules. Review the terms, robots.txt, authentication requirements, and documented quotas. Identify yourself with a descriptive user agent and contact information where appropriate. If the site signals that automation is not allowed, exclude it.
- Fetch conservatively. Set connection and read timeouts. Keep concurrency low, cache responses, and use conditional requests such as ETag or Last-Modified when supported. Follow documented quotas. If none are published, begin slowly; use exponential backoff for temporary failures and stop on access blocks rather than rotating identities or trying to evade defenses.
- Extract and retain provenance. Parse documented API fields or structured markup first. Keep the canonical source URL, source ID, retrieval timestamp, and published update timestamp if available. If retaining a raw response or hash is lawful and permitted, it can help diagnose parser changes; protect retained data and limit its retention.
- Normalize without erasing originals. Store normalized UTC dates alongside the source timezone and original date text. Store numeric prices and ISO currency codes alongside the original displayed value, fee wording, and source URL. Do not silently convert currencies or assume that a nightly or base ticket price is the full price.
- Validate before use. Require a stable source ID or canonical URL, check date ordering, valid currency codes, nonnegative prices, plausible locations, and whether mandatory-fee information is present. Route missing or contradictory values to review instead of filling them with guesses.
- Deduplicate and refresh. Match on canonical URL or source ID first; use normalized title, location, and date only as supporting signals. Keep source-specific identifiers. Set refresh intervals by source and volatility: event availability and lodging prices may need frequent rechecks, while less volatile fields may not.
- Monitor the pipeline. Track HTTP errors, blocks, empty-result rates, schema changes, and field-level drift. Alert when a parser suddenly produces missing prices, dates, or locations, and prevent stale availability or prices from being presented as current.
A conservative Python example for permitted pages
The example below is a starting point for a page you are authorized to fetch. It checks robots.txt for the supplied user agent, fetches one HTML page with a timeout, and extracts JSON-LD blocks for inspection. It deliberately does not attempt to evade blocks, solve CAPTCHAs, crawl the site, or assume a universal listing schema. Install the dependency with python -m pip install requests, save as inspect_listing.py, then run python inspect_listing.py https://example.com/permitted-listing after replacing the URL with a permitted page.
import json
import sys
from urllib.parse import urlparse
from urllib.robotparser import RobotFileParser
import requests
from bs4 import BeautifulSoup
USER_AGENT = "ListingResearchBot/1.0 (contact: [email protected])"
TIMEOUT = (5, 20)
def allowed_by_robots(url):
parts = urlparse(url)
robots_url = f"{parts.scheme}://{parts.netloc}/robots.txt"
parser = RobotFileParser()
parser.set_url(robots_url)
try:
parser.read()
except OSError as exc:
raise RuntimeError(f"Could not read {robots_url}: {exc}") from exc
return parser.can_fetch(USER_AGENT, url)
def main(url):
if urlparse(url).scheme not in ("http", "https"):
raise ValueError("Provide an http or https URL")
if not allowed_by_robots(url):
raise RuntimeError("robots.txt disallows this URL for the chosen user agent")
response = requests.get(
url,
headers={"User-Agent": USER_AGENT, "Accept": "text/html"},
timeout=TIMEOUT,
)
if response.status_code in (403, 429):
raise RuntimeError(f"Access refused or rate limited: HTTP {response.status_code}; stop and review site rules")
response.raise_for_status()
soup = BeautifulSoup(response.text, "html.parser")
records = []
for script in soup.select('script[type="application/ld+json"]'):
text = script.string or script.get_text()
if not text.strip():
continue
try:
records.append(json.loads(text))
except json.JSONDecodeError:
print("Skipping malformed JSON-LD block", file=sys.stderr)
print(json.dumps({
"source_url": response.url,
"retrieved_at": response.headers.get("Date", "record a UTC timestamp in your pipeline"),
"http_status": response.status_code,
"json_ld_blocks": records,
}, indent=2, ensure_ascii=False))
if __name__ == "__main__":
if len(sys.argv) != 2:
raise SystemExit("Usage: python inspect_listing.py https://example.com/listing")
main(sys.argv[1])
For production, record your own UTC retrieval timestamp rather than relying on the server’s Date header, which may be missing or reflect response time rather than the exact moment your application stored the record. Also check the site’s terms separately: this sample’s robots check is only one access check. JSON-LD may contain a single object, a graph, or nested offers, and markup names vary by vertical. Map fields deliberately and retain the original structured block when permitted so that a mapping error can be diagnosed.
Keep prices and availability honest
Travel rates and ticket prices can change between collection and display. Store when you retrieved each value, when the source says it was updated, and the source timezone. Keep base amount, currency, mandatory fees, optional charges, taxes, and government charges distinct when the source provides that breakdown. If mandatory charges are known, do not present a base nightly rate or ticket price as the final total.
The U.S. Federal Trade Commission’s Unfair or Deceptive Fees rule took effect May 12, 2025, and covers businesses offering, displaying, or advertising live-event tickets and short-term lodging. Covered advertised prices must disclose the total mandatory price upfront; optional charges, taxes, government charges, and shipping have separate treatment. This is a pricing disclosure rule for covered businesses, not a permission to scrape a site and not a replacement for checking the rules that apply to your own service. When you cannot establish the complete mandatory price, label the amount accurately and disclose that the total is unavailable rather than implying it is final.
Choose collection frequency, reliability, and cost deliberately
APIs often make quotas and supported fields easier to understand, but may have access charges, licensing constraints, latency, or gaps in geographic coverage. HTML collection can expose fields that a feed omits, but adds parser maintenance, site-by-site permission review, and greater risk of blocks or schema changes. Compare candidate sources on permission and licensing, field coverage, freshness, fee accuracy, rate limits and cost, bot defenses, geographic scope, privacy exposure, and the operational work needed to keep records current.
Keep request volume proportional to your need. Cache unchanged pages, honor quotas, avoid fetching the same listing repeatedly, and use conditional requests where supported. A 429 response, CAPTCHA, or explicit denial is not a cue to increase concurrency or disguise the client; pause or stop and resolve access through the publisher’s documented channel. For reliability, separate fetching from parsing so temporary network errors do not corrupt stored records, and quarantine results when validation fails.
Rank #2
There is no universal refresh interval: a hotel offer may change rapidly, while a property’s year-built field is comparatively stable. Choose a per-source schedule based on the publisher’s update behavior, your product’s freshness promise, and the cost and permission limits of collection. Display the timestamp and source context when stale data could mislead a user.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Troubleshoot common failures
- HTTP 403, CAPTCHA, or an explicit block: Stop requests to that source. Recheck terms and access rules; use a permitted API or ask the publisher about authorization. Do not attempt to bypass the barrier.
- HTTP 429 or repeated timeouts: Reduce request rate, respect any published quota, cache results, and apply backoff to transient errors. If limits are undocumented or the issue persists, pause and seek the publisher’s guidance.
- No JSON-LD appears: The page may use Microdata, RDFa, an API, or server-rendered markup without structured data. Check permitted documented sources and inspect the page structure manually; do not assume that missing JSON-LD means a request failed.
- Fields are empty or unexpectedly nested: Inspect the original permitted response and update the mapper for the actual schema. Handle arrays, nested objects, and multiple offers explicitly; do not fill missing values from a different listing by title similarity alone.
- Prices disagree with checkout or ticket selection: Confirm the selected dates, occupancy, ticket type, currency, and capture time. Determine whether the displayed amount excludes mandatory fees or is only a starting price; represent uncertainty instead of labeling it as a final total.
- Duplicate or stale records appear: Prefer source IDs and canonical URLs for matching, retain retrieval timestamps, and check redirects and listing status. Do not delete a record solely because a page temporarily timed out.
Or skip the browser setup
For visual QA or a saved page image, ScreenshotNeo can capture a page with one GET request. It returns an image or PDF; it is not a listing-data API, so use a permitted data source and parser for structured records. Its cleanup options are for the captured presentation and do not grant permission to collect data.
ScreenshotNeo accepts cookie/consent banners as a visitor and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each cleanup step can be turned off. Bot checks/CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and responses indicate the page verdict and billing status in headers. It also provides an MCP server with take_screenshot, get_page_info, and capture_pdf for AI agents and other MCP clients.
Example cURL capture (replace the target URL and use your API key):
Quick wins for a faster PC:
Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Repair Windows errors before they cause bigger problemsFix Now →curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
See the ScreenshotNeo documentation for API parameters. The same request in Python:
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
And in Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
ScreenshotNeo includes 1,000 screenshots per month free with no card; paid plans start at $5 for 3,000 screenshots. Sign up for free screenshots.
Frequently Asked Questions
Can I combine a licensed feed with permitted HTML collection?
Yes, if each source’s terms allow your intended use. Keep the source identifier and retrieval timestamp on every record so that feed data and page-derived data remain distinguishable.
Should a scraper preserve the source’s original price text?
Yes, when retention is permitted. Keeping the displayed wording alongside parsed amount, currency, fee fields, and timestamp makes it possible to audit normalization without presenting a transformed number as the source’s exact statement.
The Tool Desk
Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

