Free tools Windows power users keep installed
One-click scans. No signup required.
Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
Start with the data rights, not the parser. Define the fields, geography, freshness and downstream use you need; then check the directory’s official API, open-data offer or licensing product. Extract web pages only when the source permits it and when an API or licensed feed cannot meet the requirement. Google Maps, Google Places and Google Business Profile each impose source-specific restrictions, while Yelp documents search, matching and business-detail endpoints alongside a separate licensing route.
1. Define the dataset before collecting anything
Write a short specification that answers four questions:
- Fields: for example, business name, address, phone, category, coordinates, hours, source identifier and retrieval time. Do not collect fields merely because they appear on a page.
- Geography: specify countries, regions, cities, postal areas or a radius. A bounded area is easier to query, deduplicate and refresh than “all businesses.”
- Freshness: decide whether records may be a month old, must be checked daily, or are a one-time snapshot.
- Use: internal analysis, an authorized client tool, advertising outreach, public display or a replacement directory have different contractual and privacy implications.
Your use case determines which source is appropriate. An API that is suitable for displaying a few search results may not permit a permanent, redistributable database.
Quick wins for a faster PC:
Repair Windows errors before they cause bigger problemsFix Now →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Clear out junk files and repair common Windows errorsFree Scan →2. Check an official API or licence first
Before writing a crawler, open the provider’s current developer documentation and terms. Look for authentication, query limits, available fields, matching support, attribution, storage duration, permitted display and reuse rights. Yelp’s developer documentation describes private-key authentication, business search by keyword, category and location, business matching, business details and up to three review excerpts. It also links to data-licensing products. Confirm the current plan and contractual rights for your intended use at Yelp’s developer documentation before building around it.
#1 Best Overall
An API response is not automatically an unrestricted export. Google’s general API terms restrict scraping, database building, permanent copies and retaining cached copies longer than allowed by cache headers unless the content owner or applicable law expressly permits it. Google’s Maps Additional Terms prohibit mass downloads and bulk feeds and restrict using Maps to create or augment a business-listings database that substitutes for, or is substantially similar to, Google Maps.
Google Places and Maps
Google’s Places policy requires appropriate Google Maps attribution when displaying content and sets storage rules. The policy states: “You can therefore store place ID values indefinitely.” That exception applies specifically to place IDs, not to every field returned for a place. Customers billed in the European Economic Area may be covered by separate EEA terms, so check the terms for the project’s billing region.
Google Business Profile APIs
Business Profile APIs are for creating, managing and reporting on listings that the user owns or is authorized to manage, including tools serving clients with that authorization. They are not a general prospecting or lead-generation database. The policy also limits certain third-party automated access and restricts some stored content to temporary storage of no more than 30 calendar days.
3. Choose sources by coverage and rights
| Decision area | Questions to answer |
|---|---|
| Geographic and category coverage | Does the source cover the places and business types you need, or only selected markets? |
| Fields and freshness | Are phone numbers, hours, coordinates, identifiers and update times available? How often may you refresh? |
| Matching | Can you match a known business to a source record, and what identifier is returned? |
| Storage and reuse | May you cache, republish, enrich or combine records? For how long? |
| Display and attribution | What logos, links or notices are required when results are shown? |
| Regional terms | Do billing country, EEA rules or local privacy requirements change the contract? |
| Total cost | Include request charges, engineering, proxy or browser infrastructure, monitoring and compliance work. Comparable provider pricing and coverage figures are not established here. |
A multi-source design can improve coverage. A Georgia Tech research example describes iterative, location-based collection across Foursquare, Yelp, Google Maps and OpenStreetMap using Python APIs. Treat that as an academic illustration, not proof that every provider currently offers the same access or terms.
4. If page extraction is permitted, use a bounded workflow
- Confirm permission. Read the site’s terms, API policy, robots guidance and applicable privacy obligations. Obtain written authorization for private or client-controlled directories.
- Map the site. Identify category and location pages, pagination or cursors, detail-page links, and any explicit rate limits. Start with a small sample.
- Throttle requests. Use a conservative delay, a descriptive user agent and exponential backoff for transient failures. Never attempt to defeat a CAPTCHA, bot check, login control or access restriction.
- Extract only specified fields. Save the source URL, source identifier and retrieval timestamp with each row.
- Validate. Check required fields, normalize phone and address formats, reject impossible coordinates and record parsing errors rather than silently dropping rows.
- Deduplicate. Prefer a provider identifier. Otherwise combine normalized name, address, phone and coordinates; retain the competing source values and provenance for review.
- Refresh lawfully. Re-check records only at an interval allowed by the source. Keep deletion and correction procedures for businesses that request changes where applicable.
- Audit the output. Store the policy version, query parameters, timestamps, code version and a sample of source pages so another person can reproduce the decision.
Illustrative Python template for an authorized directory
This example is a starting point for a site you are allowed to access. Replace the URL and CSS selectors with the directory’s documented structure; it does not bypass authentication, CAPTCHAs or technical controls.
import csv, time, requests
from bs4 import BeautifulSoup
from urllib.parse import urljoin
START = "https://directory.example/authorized-search?q=plumber"
HEADERS = {"User-Agent": "ResearchCollector/1.0 (contact: [email protected])"}
rows = []
url = START
while url:
response = requests.get(url, headers=HEADERS, timeout=30)
response.raise_for_status()
soup = BeautifulSoup(response.text, "html.parser")
for card in soup.select("article.business-card"):
link = card.select_one("a.details")
rows.append({
"name": card.select_one(".name").get_text(" ", strip=True),
"address": card.select_one(".address").get_text(" ", strip=True),
"phone": card.select_one(".phone").get_text(" ", strip=True),
"source_url": urljoin(url, link["href"]) if link else "",
"retrieved_at": time.strftime("%Y-%m-%dT%H:%M:%SZ", time.gmtime()),
})
next_link = soup.select_one("a[rel=next]")
url = urljoin(url, next_link["href"]) if next_link else None
time.sleep(2)
with open("businesses.csv", "w", newline="", encoding="utf-8") as f:
writer = csv.DictWriter(f, fieldnames=rows[0].keys() if rows else ["name"])
writer.writeheader(); writer.writerows(rows)
For production, add bounded page counts, retry rules, structured logging, schema validation and a queue that can pause collection when the provider changes its terms or markup.
5. Normalize, match and retain provenance
Keep raw values alongside normalized values. Lowercase and Unicode-normalize names, standardize phone digits and address components, and retain the original spelling for display. Exact source IDs should take precedence over fuzzy matching. For uncertain matches, send candidates to a review queue instead of merging automatically. Every record should carry source name, source URL or identifier, retrieval timestamp and the rule that produced the match.
Retention is a policy decision, not merely a database setting. Follow the provider’s permitted cache period, attribution requirements and deletion process. Google’s place-ID exception does not grant permission to retain all associated listing content indefinitely.
6. Common failure modes and fixes
“The API works, but I cannot publish the dataset”
API availability and redistribution rights are separate. Re-read storage, display and licensing clauses; contact the provider for a data-licensing product or redesign the output to show permitted, short-lived results.
“The crawler returns empty pages”
The content may be rendered by JavaScript, blocked for automated access or dependent on a location setting. First look for an official endpoint or export. If page access is authorized, inspect the documented network request and use a supported method rather than trying to evade a block.
“Rows are duplicated”
Pagination, aliases and multiple branches commonly produce duplicates. Match on a stable provider ID where available; otherwise use normalized name, address, phone and coordinates, and route ambiguous pairs for human review.
“Requests are timing out or receiving 429 responses”
Reduce concurrency, honor Retry-After, add exponential backoff and narrow the geography or category. Cache only for the period the source allows. Do not rotate identities to circumvent a rate limit.
“Fields disappear after a redesign”
Version your parser, validate required fields and alert on sudden null-rate changes. Pause the job, review the current documentation and update selectors only after confirming that collection remains permitted.
“A Google listing seems ideal for a lead database”
Google’s Maps and API terms specifically restrict mass downloads, bulk feeds and substitute directories. Business Profile APIs are scoped to listings you own or are authorized to manage. Choose a licensed source instead of assuming that visible data may be copied.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Or skip the browser setup
When your task is to capture a permitted directory page for documentation, QA or an internal record, ScreenshotNeo provides a single website-screenshot request. It accepts the consent banner like a visitor, then removes more than 60 known consent platforms, newsletter popups and chat widgets before capture; each cleanup step can be disabled. Bot checks and CAPTCHAs, blank pages, timeouts, failed loads and cache hits are not billed, and the response identifies the result with X-Page-Verdict and X-Billed headers. Its MCP server offers take_screenshot, get_page_info and capture_pdf tools for Claude, Cursor and other MCP clients.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Use the API documentation at screenshotneo.com/docs/ for all options, including full-page capture, lazy-image loading, CSS-selector element capture, device and viewport settings, dark mode, retina scale, PDF paper and page controls, custom CSS or JavaScript, clicks, waits, request blocking, headers, cookies, user-agent, authorization, timezone, geolocation, transparent backgrounds, resizing, TTL caching, signed links, asynchronous webhooks and bulk capture of up to 100 URLs per call.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://directory.example/listings -o shot.webp
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://directory.example/listings"}, timeout=90)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://directory.example/listings' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
if (!res.ok) throw new Error(`${res.status} ${res.statusText}`);
require('fs').writeFileSync('shot.webp', Buffer.from(await res.arrayBuffer()));
The Free plan includes 1,000 shots per month with no card. Paid plans start at $5 for 3,000 shots; every feature is on every plan. Create a free ScreenshotNeo account.
7. A launch checklist
- Fields, geography, freshness and purpose are written down.
- An official API, open-data feed or licence was checked first.
- Terms, attribution, storage, regional rules and deletion duties are documented.
- Queries are bounded, throttled and monitored.
- Source IDs, URLs and retrieval times are retained with each row.
- Duplicate and change-detection rules are tested on a sample.
- Output is limited to what the source permits you to store and display.
- Documentation and terms are scheduled for review because providers can change them.
Frequently Asked Questions
Can I build a permanent directory from an API response?
Not automatically. Check that provider’s storage and reuse terms; Google’s general API terms and Maps policies impose specific limits, while Yelp directs developers to separate licensing products.
Is Google Business Profile a prospecting API?
No. It is intended for listings that the user owns or is authorized to manage, including authorized client tools.
What should I preserve for auditability?
Keep the source identifier or URL, retrieval timestamp, query parameters, policy version and the transformation or matching rule used.
The Bottom Line
Responsible directory collection is a rights-and-data-design problem before it is a scraping problem: define the dataset, prefer a permitted API or licence, collect narrowly, preserve provenance and refresh only within the source’s rules.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

