Do these 3 things before closing this tab:
1Scan for outdated or missing drivers - takes under a minute2Clear out junk files and repair common Windows errors3Fix the driver behind crashes, sound loss and screen glitchesStart with an official API, export, or licensed feed. Use web crawling only when the site’s current terms and technical instructions permit it. Before collecting anything, define the fields and date range you need, verify access and reuse rights, and design a dataset that preserves provenance, timestamps, pagination and missing records. A few visible pages are not the same as a complete review corpus.
Choose an authorized collection route first
“Scraping” has no platform-neutral permission. The same technique can be allowed on one service and prohibited on another. Yelp says third-party software may not scrape or copy content from its site (Yelp support). Google Maps Platform terms prohibit scraping or exporting Maps content for use outside Google’s services, including copying and saving reviews (Google Maps Platform terms, archived June 4, 2025).
Evaluate routes in this order:
| Route | Use it when | Resolve before coding |
|---|---|---|
| Official API | It covers the content and purpose you need | Fields, quotas, roles, regions, refresh schedule, attribution, retention and reuse |
| Licensed feed or partner | You need broader commercial coverage | Licensed sources, permitted storage, combination, display, redistribution and model-training rights |
| Direct crawling | No suitable authorized feed exists and the site permits automation | Terms, robots.txt, rate, identification, privacy, copyright, database and jurisdiction rules |
Robots.txt is a crawler instruction protocol, not a license. RFC 9309 describes requests to crawlers; it does not replace terms or other legal requirements (RFC 9309).
Define the data boundary
Write the question before the collector
Specify products or businesses, locales, date range, rating scale, review text, question-and-answer fields, identifiers and update frequency. Collect reviewer names, profile URLs or other personal data only when necessary, with a documented basis and retention period.
PC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware match#1 Best Overall
Check eligibility and coverage
Record whether an account, seller role or regional entitlement is required. Yelp’s documented Places API reviews endpoint returns up to three review excerpts per business (Yelp Places API documentation); it is not an unrestricted export of every review. Amazon’s Customer Feedback API is for sellers and vendors, lists the US, UK, France, Italy, Germany, Spain and Japan, is refreshed weekly, returns data only in English, and documents a Brand Analytics role for the operation (Amazon Customer Feedback API).
Build a reproducible collector
API workflow
- Create credentials through the platform’s documented developer or seller console and confirm the required role and region.
- Request the narrowest endpoint and fields that answer your question. Save the API version, query parameters, page tokens and request timestamps.
- Follow documented pagination until the endpoint signals completion. Stop and record the response when you receive 401, 403 or 429 errors; do not rotate credentials or evade a block.
- Use modest concurrency, exponential backoff for transient failures, and an idempotency key or stable source ID so retries do not create duplicates.
- Store raw responses separately from normalized records. Keep source IDs, rating scale, language, timestamps, locale, endpoint and retrieval time.
Permitted crawling workflow
- Read the current terms, privacy notice and robots.txt for every host. A disallow rule or explicit anti-scraping term means you should stop and seek permission or an alternative feed.
- Identify your client with a truthful user agent and contact address where appropriate. Keep request rates low and avoid parallel bursts.
- Fetch only pages needed for the defined scope. Cache your own request queue and honor retry-after responses.
- Capture the source URL, page timestamp, HTTP status, parser version and a hash of the raw document. Never treat a rendered page count as proof of completeness.
Normalize reviews, questions and answers safely
- Stable identity: retain platform, business or ASIN ID, review or question ID when supplied, and the source URL.
- Time: store publication, edit and retrieval timestamps with timezone information.
- Meaning: keep raw text immutable; put cleaned text, translation and tokenized forms in separate fields.
- Ratings: preserve the original scale and labels rather than converting silently to a five-point scale.
- Relationships: link answers to question IDs and reviews to product or business variants.
- Change history: retain edit, deletion and moderation events when the source exposes them.
Deduplicate on a source ID first. If no ID exists, combine conservative keys such as normalized URL, author-independent timestamp and a text hash, then review collisions manually. Do not delete near-duplicates merely because two variants share wording; syndicated or translated content can be legitimate.
Measure coverage and bias
For each run, record pages or API pages fetched, maximum result limits, ranking or selection rules, missing records, failed requests and the source’s reported total. Stratify results by date, language, rating and product or business variant. A Yelp response containing three excerpts, or an Amazon topic insight, cannot be described as all reviews.
Report your sample definition in the dataset and any publication: “reviews returned by endpoint X between dates Y and Z,” not “customer opinion.” Compare observed counts with source totals where the platform supplies them, and keep a failure log so a later rerun can explain changes.
Retention, attribution and republication
API access does not grant unlimited storage or display rights. Google Places policies require author attribution and direct access to source reviews, and restrict caching or storage except for stated exceptions (Google Places policies and attributions). Before publishing, confirm whether you may show full text, excerpts, ratings, author names, translations, derived scores or commercial analyses, and how long each may be retained.
Keep a rights register with the source, policy URL, date checked, allowed purpose, retention limit, attribution format and deletion process. If a platform requires live source access, link readers to that source rather than presenting an old local copy as current.
Rank #3
Protect review and Q&A integrity
Do not edit text to change its message or selectively publish favorable material while calling the result representative. The FTC’s platform guidance says, “Don’t edit reviews to alter the message,” and recommends reasonable authenticity processes and equal treatment of positive and negative reviews (FTC guide for platforms). The Consumer Reviews and Testimonials Rule took effect October 21, 2024; FTC staff says its Q&A is not definitive or comprehensive, so obtain advice for your facts and jurisdictions (FTC Rule Q&A).
Amazon’s policies illustrate source-specific obligations: “Only post your own content or content that you have permission to use on Amazon” (Amazon Community Guidelines). A person with a financial or close personal connection may answer product questions only with clear and conspicuous disclosure (Amazon promotional-content guidance). These rules govern participation and display; they do not by themselves grant scraping permission.
Production checklist
- Purpose, fields, products, locales and date range are written down.
- Terms, API documentation, license and robots.txt were checked on a recorded date.
- Eligibility, quotas, refresh cadence, attribution and retention are documented.
- Raw and normalized data are separated, with source IDs and timestamps.
- Pagination, retries, rate limits and access-denied responses are handled conservatively.
- Coverage limits, missing pages and selection bias are measured and reported.
- Deletion, correction, consent and access requests have an owner and procedure.
- Publication preserves meaning, treats ratings consistently and links required attribution.
Common failures and fixes
“I received only a few reviews”
Cause: endpoint cap, excerpt design, ranking or missing pagination. Fix: read the documented maximum, follow page tokens, and label the result as excerpts or insights. Yelp’s endpoint explicitly documents up to three excerpts per business.
403 Forbidden or a blocked request
Cause: missing role, disallowed automation, invalid origin or terms restriction. Fix: verify credentials and eligibility, stop automated retries, and switch to an authorized API, export or licensed provider.
429 Too Many Requests
Cause: quota or rate limit. Fix: honor Retry-After, reduce concurrency, add exponential backoff and request a documented quota increase rather than evading limits.
Duplicate or missing records
Cause: unstable pagination, edits, deleted content or retries. Fix: upsert by source ID, store page tokens and run-level logs, hash raw records, and report gaps instead of silently filling them.
Free tools Windows power users keep installed
One-click scans. No signup required.
Best Value
Can I publish the collected text?
Not automatically. Recheck the source’s display, attribution, storage and copyright terms for the exact use. If rights are unclear, publish aggregate findings or link to the source instead of reproducing text.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Or skip the browser setup
If your project also needs screenshots of review or Q&A pages, ScreenshotNeo provides a one-request website screenshot API. It accepts consent banners like a visitor and removes more than 60 known consent platforms, newsletter popups and chat widgets before capture; bot checks, blank pages, timeouts, failed loads and cache hits are not billed, and response headers identify the page verdict and billing status. Its MCP server offers take_screenshot, get_page_info and capture_pdf tools for Claude, Cursor and other MCP clients.
cURL:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python:
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
See the ScreenshotNeo documentation for options such as full-page capture, selector capture, custom headers and cookies, waits, blocking, PDF output, caching, bulk jobs and signed webhooks. The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000. Create a free ScreenshotNeo account.
Frequently Asked Questions
Is robots.txt enough permission to scrape reviews?
No. It is a crawler protocol. You must separately check the site’s terms, API or license and applicable privacy, copyright and database rules.
The Tool Desk
Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →What should I call a dataset containing only API excerpts?
Describe the endpoint, date range and limit precisely—for example, “reviews returned by the documented endpoint,” not “all customer reviews.”
How often should a review dataset be refreshed?
Use the source’s documented cadence where available, record each retrieval time, and design for edits and deletions rather than assuming records are permanent.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

