Recommended Free Tools
You can collect local-business listings with Python when the source permits the access and reuse you intend. Start with a documented API, an authorized export, or permission from the site owner; use HTML scraping only for pages you are allowed to fetch. Google Maps and Places are not a default source for building an independent directory: their terms restrict scraping and reuse, and API rules cover storage, display, and attribution.
Start with the source and the rules
“Scraping” describes a way to collect data, not permission to collect it. Before writing code, identify the exact site or API, the account and region involved, the fields you need, and how you plan to store, display, or reuse the results. Prefer an official data export, documented API, or written permission where available. Collect only fields needed for the intended purpose, and avoid unnecessary personal information.
Review the source’s terms and machine-readable instructions, including its robots.txt file. A robots.txt check can tell you what the published rules say about a particular user agent and URL; it does not establish contractual or legal permission by itself. Python’s urllib.robotparser documentation describes its robots.txt parser, including optional crawl-delay and request-rate directives.
Record the source URL, collection date, intended use, and fields in your project notes. That provenance helps you re-check the rules and identify stale records later.
Windows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallOutdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware match#1 Best Overall
Can you scrape Google Maps with Python?
Do not treat Google Maps scraping as a shortcut for building a separate local-business database. Google’s Terms of Service address automated access that violates machine-readable instructions and scraping content that does not belong to the user. Google Maps Platform terms state: “Customer will not extract, export, or otherwise scrape Google Maps Content for use outside the Services.” The terms give copying business names, addresses, or user reviews as examples. Check the current terms applicable to your product, account, and use rather than assuming a general web-scraping rule overrides them.
If you use the Places API, its policies restrict pre-fetching, caching, and storing Places content beyond stated exceptions; place IDs are exempt from caching restrictions. Attribution requirements apply when displaying API content. The policy points customers with an EEA billing address to region-specific terms, so verify the rules for your billing geography and service.
The Google Business Profile API policies are for managing listings that you own or are authorized by the business owner to manage—not harvesting arbitrary businesses for a directory. The policy describes limited temporary storage of content: it must be secure, unmanipulated or unaggregated, and not exceed 30 calendar days. That provision is specific to the described Business Profile policy and is not a general retention allowance for Maps or Places data. The policy also requires prior specific and express consent for certain automated listing actions.
Choose the right Python collection method
| Method | Use it when | What to check |
|---|---|---|
| Official export or documented API | The source provides a supported way to obtain the records. | Credentials, usage limits, permitted fields, retention, display and attribution rules. |
| Owner-authorized management API | You manage the business’s listings or have the required authorization. | Scope of authorization, consent for automated actions, storage and security obligations. |
| HTML fetch and parse | The specific pages are permitted to fetch and reuse for your purpose. | Terms, robots.txt, page structure, request load, and whether the content is server-rendered HTML. |
A standard-library HTTP request is suitable for a permitted static page. If a page fills its listing data with JavaScript after loading, the initial HTML may not contain the records. This tutorial does not assume a particular directory’s markup or provide selectors that are guaranteed to work on a live site.
Rank #2
Check robots.txt before fetching
The following Python 3 script checks the robots.txt rules for the exact URL and user-agent string you provide. It reports crawl-delay and request-rate directives when present. A “can fetch” result is only one input to the decision: proceed only if the site’s terms and applicable rights also permit the request and use.
from urllib.parse import urlsplit, urlunsplit
from urllib.robotparser import RobotFileParser
page_url = "https://example.com/directory/shops"
user_agent = "MyDirectoryResearchBot"
parts = urlsplit(page_url)
robots_url = urlunsplit((parts.scheme, parts.netloc, "/robots.txt", "", ""))
parser = RobotFileParser(robots_url)
try:
parser.read()
except Exception as exc:
raise SystemExit(f"Could not read {robots_url}: {exc}")
print("robots.txt:", robots_url)
print("Can fetch page:", parser.can_fetch(user_agent, page_url))
print("Crawl delay:", parser.crawl_delay(user_agent))
print("Request rate:", parser.request_rate(user_agent))
Replace the example domain and path with the authorized target. If the result is false, do not fetch that URL with the named user agent. If robots.txt cannot be retrieved, do not treat the failure as permission; resolve the source rules first.
Fetch an allowed static page with Python
Python’s urllib.request can open a URL or a Request object, set request headers, and enforce a timeout. The response body is bytes. Decode it using the page’s declared encoding when available rather than assuming every page is UTF-8.
from urllib.request import Request, urlopen
url = "https://example.com/directory/shops"
request = Request(
url,
headers={"User-Agent": "MyDirectoryResearchBot/1.0 (contact: [email protected])"},
)
try:
with urlopen(request, timeout=20) as response:
content_type = response.headers.get_content_charset()
encoding = content_type or "utf-8"
html = response.read().decode(encoding, errors="replace")
print("Status:", response.status)
print("Encoding:", encoding)
print(html[:500])
except TimeoutError:
print("The server did not respond before the timeout.")
Use a real contact address in the user-agent if you operate a collection service. A finite timeout prevents a stalled request from hanging the run indefinitely. Keep request volume low, avoid repeated unnecessary downloads, and stop rather than trying to evade access denial, blocking, or other restrictions. These are prudent operating practices, not a published rate limit for every site.
Parse only fields supported by the page
HTML parsing is source-specific. Class names, element structure, and embedded data can change, so inspect a permitted page and define selectors only after confirming that the fields are present in its HTML. If the content is absent from the server response, a static fetch cannot extract it. Do not assume that a selector found on one directory works on another.
For a site whose permitted markup you have inspected, Beautiful Soup offers a concise selector workflow. Install it with python -m pip install beautifulsoup4. Replace the example selectors with the exact selectors you verified in the target page; the example domain and selectors below are intentionally illustrative, not a working scraper for a named directory.
from urllib.request import Request, urlopen
from bs4 import BeautifulSoup
url = "https://example.com/directory/shops"
request = Request(url, headers={"User-Agent": "MyDirectoryResearchBot/1.0"})
with urlopen(request, timeout=20) as response:
encoding = response.headers.get_content_charset() or "utf-8"
html = response.read().decode(encoding, errors="replace")
soup = BeautifulSoup(html, "html.parser")
records = []
# Replace these selectors and field rules after inspecting permitted markup.
for card in soup.select(".business-card"):
name_el = card.select_one(".business-name")
address_el = card.select_one(".business-address")
records.append({
"name": name_el.get_text(" ", strip=True) if name_el else None,
"address": address_el.get_text(" ", strip=True) if address_el else None,
"source_url": url,
})
for record in records:
print(record)
Missing elements become None instead of silently producing misleading values. Before relying on output, inspect several records, verify field meanings, and check whether pagination or a source-provided export is available. Store a source-appropriate key for deduplication, retain collection timestamps and provenance, label missing values, and periodically revisit fields that can become stale. These are data-quality practices, not a guarantee that a listing remains accurate.
Plan storage, display, and operations
- Permission and reuse: Confirm that collection, retention, aggregation, and display are all allowed for the chosen source and use. API access does not mean unrestricted reuse.
- Attribution: Follow the source’s display attribution requirements. This is especially important for Places API results.
- Request behavior: Set timeouts, keep volume modest, and stop on denial or blocking. Do not repeatedly request identical pages without a clear need.
- Data quality: Preserve where and when each record came from; deduplicate with an appropriate key; validate missing, malformed, and stale fields.
- Maintenance: Re-check terms and source structure when the collection changes or breaks. HTML selectors are coupled to markup and can fail after a site update.
- Cost and limits: For an API, check the account’s current credentials, quotas, and charges before running a collection. No universal quota or price applies across providers.
Troubleshooting common problems
robots.txt says the page cannot be fetched
Do not proceed with that user agent and URL. Revisit the source’s permitted access route, such as an official API, export, or explicit authorization. A different user-agent string is not a legitimate workaround for a restriction.
Free tools Windows power users keep installed
One-click scans. No signup required.
The request times out or returns an error
Check that the URL is correct and accessible, and keep the timeout finite. If the source denies or blocks access, stop and contact the owner or use its supported access method rather than retrying aggressively or attempting to bypass controls.
The response is not the expected page
Inspect the status, response headers, and a small portion of the returned HTML. The server may return a redirect, an error page, or a shell whose listings are loaded later by JavaScript. A static HTTP request only retrieves the response served for that request; it does not execute a page’s client-side scripts.
Names or addresses are missing
Check the actual markup and the selectors you configured. The page may use a different structure, omit a field, or load it dynamically. Return an explicit missing value and validate sample records; do not infer a business address from unrelated page text.
Text contains replacement characters
Use the response’s declared character set where available, as the examples do. If the site omits or misstates its encoding, inspect the page’s metadata and handle that source deliberately rather than globally assuming UTF-8 or silently discarding characters.
Do these 3 things before closing this tab:
1Clear out junk files and repair common Windows errors2Fix the driver behind crashes, sound loss and screen glitches3Repair Windows errors before they cause bigger problemsBest Value
A Places or Business Profile workflow has retention questions
Use the policy for the exact Google product and account context. Places API caching restrictions and Business Profile’s limited temporary-storage provision are not interchangeable. Check current regional terms and display requirements before storing or showing data.
Or skip the browser setup
If your authorized workflow needs a rendered visual capture rather than structured business records, ScreenshotNeo is a website screenshot API and MCP server. It does not replace a permitted data API or extract a directory into structured fields; it can return a screenshot or PDF of a page. A single GET request in Python is:
ScreenshotNeo API documentation
import requests
r = requests.get(
"https://api.screenshotneo.com/v1/shot",
params={"access_key": "YOUR_API_KEY", "url": "https://example.com/directory/shops"},
timeout=90,
)
open("shot.webp", "wb").write(r.content)
Cookie banners, popups, and chat widgets are removed before the shot; bot checks, blank pages, and failed loads are never billed. Its MCP server lets AI agents take screenshots. The free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. Sign up for 1,000 free screenshots a month with no card.
Frequently Asked Questions
Does robots.txt give me legal permission to collect a business directory?
No. It is a machine-readable access signal for a user agent, not a complete statement of contractual or legal rights. Review the source’s terms and applicable data rights as well.
Will urllib retrieve listings that appear only after a page loads in a browser?
Not by itself. urllib fetches the HTTP response; it does not run the page’s client-side JavaScript. Use a source-supported API or another expressly permitted route if the initial HTML does not contain the data.
Can I use the Google Business Profile API to build a directory of businesses I do not manage?
The Business Profile APIs are scoped to listings you own or are authorized to manage. They are not a general-purpose collection API for arbitrary businesses.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




