Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
Best overall approach: use a managed extraction API when you need marketplace data quickly and consistently; use Scrapy when custom business logic and code ownership outweigh engineering time; choose Apify when cloud schedules, reusable Actors, storage and integrations are central. Validate each choice on your own product set—field completeness, successful records, latency, maintenance work and cost per successful result matter more than a feature checklist.
What retail analytics scraping must collect
Retail scraping is the automated download of website data into a structured format that software can process. A useful retail pipeline usually combines several entity types rather than collecting only a displayed price.
- Product identity: SKU or product ID, title, brand, category, variant, size, color and canonical URL.
- Price: list price, sale price, currency, unit price, promotion text and the time observed.
- Offer and seller data: seller name, seller condition, shipping charge, delivery estimate and Buy Box or featured-offer ownership where a marketplace exposes it.
- Availability: in-stock, out-of-stock, back-order and quantity signals, plus store or fulfillment location when shown.
- Marketplace attributes: ratings, review counts, badges, category breadcrumbs and variation relationships.
- Evidence: response time, source URL, HTTP status, parser version and a capture or raw-response reference so a price change can be audited.
Decide the grain of your dataset before choosing a tool. A product-level table cannot represent three sellers with different prices without a separate offers table. Likewise, inventory observed for one ZIP code or logged-in account is not a universal stock statement. Store geography, currency, locale, timestamp and account state alongside every observation.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
The three tool categories
Managed extraction APIs
Oxylabs, Bright Data and Zyte host retrieval infrastructure, proxy or IP management, browser or JavaScript execution and, in some products, parsing into structured fields. This reduces the amount of crawler and anti-blocking infrastructure your team must operate. The trade-off is recurring vendor cost, a dependency on a provider’s coverage and less control over unusual page logic.
#1 Best Overall
Code-first frameworks
Scrapy is an open-source Python framework for maintainable, highly customized spiders. You own the request flow, parsing rules, data model and deployment choices. You also own monitoring, proxy strategy, browser integration, retries, ban handling and parser maintenance when a retailer changes its markup.
Cloud orchestration platforms
Apify packages scrapers as cloud Actors. Its documented capabilities include storage and exports, rotating datacenter and residential proxies, schedules, integrations, monitoring and collaboration. This is useful when several teams need reusable jobs and operational controls without building a complete execution platform.
Comparison at a glance
| Decision axis | Managed API | Scrapy | Apify Actors |
|---|---|---|---|
| Time to first dataset | Usually shortest because retrieval and, depending on product, extraction are hosted. | Longest; your team builds crawling, parsing and operations. | Fast when an appropriate Actor exists or can be adapted. |
| Customization | Constrained by provider parameters and supported schemas. | Highest control over requests, parsers and business rules. | High within Actor code, with platform conventions for running and storing jobs. |
| JavaScript-heavy pages | Choose a plan or endpoint that supports browser rendering; rates can differ when rendering is required. | Requires your own browser integration and its compute cost. | Actor can use browser tooling, with execution and proxy settings managed in the platform. |
| Proxy and ban handling | Provider-managed options are a central benefit. | You design rotation, throttling, retries and escalation. | Rotating datacenter and residential proxy options are documented. |
| Parsing | Automatic extraction may be available, or you can consume HTML/JSON. | Fully custom selectors and normalization. | Actor-specific parsing, with platform storage and export. |
| Scheduling and monitoring | Check the selected API’s job, usage and alert capabilities. | Build or integrate schedulers, metrics and alerts. | Schedules, monitoring, integrations and collaboration are documented platform features. |
| Lock-in | Higher: schemas, quotas and proxy behavior are provider-specific. | Lower at the framework level; infrastructure remains your responsibility. | Actor code is portable in principle, but storage, schedules and integrations use platform services. |
| Best fit | Fast, repeatable multi-site collection with limited operations staff. | Long-lived, unusual or highly customized crawlers owned by engineering. | Teams operating many reusable jobs with shared cloud controls. |
Leading options and stated pricing
Oxylabs Web Scraper API
Oxylabs lists a free trial of up to 2,000 results. Its Micro plan is listed at up to 98,000 results starting at $49 per month. Rates vary by target and by whether JavaScript rendering is required, so treat the published figures as current vendor-page figures that should be rechecked before purchase. Ask specifically how a result is counted, which product fields are parsed, and whether browser rendering changes the rate.
Windows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstallCrashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minuteBright Data eCommerce Scraper API
Bright Data documents seller names, offer prices and Buy Box ownership for Amazon, Walmart and eBay. It states that each new account includes 5,000 free credits per month. Confirm how credits map to requests, rendered pages and retries for your target mix; a credit allowance is not the same as a guaranteed number of successful product records.
Zyte API and Scrapy Cloud
Zyte documents price intelligence, market and competitor analysis, product listings, prices, reviews and inventory, together with browser automation, automatic extraction and Scrapy Cloud execution. This combination can suit a team that wants Scrapy-style control while outsourcing parts of browser execution and operations. Confirm the exact fields and geography available for each retailer before designing your schema.
Apify
Apify’s Actors, storage and exports, rotating datacenter and residential proxies, schedules, integrations, monitoring and collaboration make it a strong orchestration choice. Actor quality differs by target, so inspect the input schema, output dataset and update history of any community or internal Actor you plan to rely on.
How to choose for a real retail program
- List targets and jurisdictions. Record every marketplace, retailer, country, language, currency and whether pages differ by store, ZIP code or login state.
- Define required fields and freshness. Separate must-have fields (for example, landed price and seller) from optional fields such as review text. Set collection intervals by use case: an hourly Buy Box alert has different requirements from a weekly catalog refresh.
- Classify page behavior. Test whether product data is in the initial HTML, embedded JSON or rendered only after JavaScript runs. Note consent dialogs, bot checks, pagination, infinite scroll and variant selectors.
- Pick an operating model. Select a managed API for speed and maintained access, Scrapy for maximum ownership, or Apify for reusable cloud jobs and shared operations.
- Run the same target set through at least two approaches. Compare field completeness, successful-record rate, latency, maintenance burden and cost per successful record. Do not infer performance from a vendor’s feature list.
- Design for change. Version parsers and schemas, retain raw evidence for disputed prices, and alert when a normally populated field becomes null across many pages.
A maintainable Scrapy pattern
For a site you are permitted to crawl, start with a narrow spider and an explicit item contract. The example below illustrates structure rather than a retailer-specific selector; replace selectors only after checking the target’s terms and robots directives.
Quick wins for a faster PC:
Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Repair Windows errors before they cause bigger problemsFix Now →Scan for outdated or missing drivers - takes under a minuteDriver Scan →import scrapy
class ProductItem(scrapy.Item):
url = scrapy.Field()
title = scrapy.Field()
price = scrapy.Field()
currency = scrapy.Field()
availability = scrapy.Field()
observed_at = scrapy.Field()
class ProductSpider(scrapy.Spider):
name = "products"
allowed_domains = ["example.com"]
start_urls = ["https://example.com/catalog"]
def parse(self, response):
for href in response.css("a.product-card::attr(href)").getall():
yield response.follow(href, callback=self.parse_product)
next_url = response.css("a.next::attr(href)").get()
if next_url:
yield response.follow(next_url, callback=self.parse)
def parse_product(self, response):
yield ProductItem(
url=response.url,
title=response.css("h1::text").get(default="").strip(),
price=response.css("[itemprop='price']::attr(content)").get(),
currency=response.css("[itemprop='priceCurrency']::attr(content)").get(),
availability=response.css("[itemprop='availability']::attr href").get(),
observed_at=response.headers.get("Date", b"").decode("ascii", "ignore"),
)
In production, add request throttling, bounded retries, deduplication, structured logs and a validation step that rejects impossible prices or missing product IDs. Keep selectors and normalization tests in version control. If a page requires a browser, isolate that work from ordinary HTTP requests and measure its additional latency and compute cost rather than silently applying it to every URL.
Reliability, performance and cost controls
Measure successful records, not requests
A response can be a consent page, a bot challenge, an empty shell or a product with missing offers. Track requested URLs, valid product records, field-level completeness, retries, latency percentiles and blocked responses. Calculate cost per valid record and cost per complete offer set; these metrics expose expensive retries and partial parses.
Use conservative request behavior
- Respect published rate limits, robots directives and a site’s terms.
- Use the smallest crawl scope that answers the business question; do not repeatedly fetch unchanged pages without a reason.
- Cache stable catalog pages and separate high-frequency price or inventory checks.
- Retry transient network failures with backoff, but do not hammer a site after a bot check or an explicit cease request.
- Keep credentials, cookies and authorization headers out of logs and exported datasets.
Plan for data quality
Normalize currencies with the observation’s locale and timestamp. Preserve both the displayed price and a parsed numeric value. Treat “currently unavailable” as a state, not a zero. For variants, store the selected option values so a low price is not accidentally attributed to every size or color.
Rank #3
Compliance and permission checklist
Zyte’s terms state: “The Services shall be used solely to scrape data from publicly accessible websites.” They also place lawful-use responsibility on the customer and allow suspension when a target requests that activity stop or continued activity creates legal, operational or business risk. Apply the same caution regardless of provider.
The Tool Desk
Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →- Review each target’s terms, robots directives, rate limits and contractual permissions.
- Check privacy and data-protection duties when collecting reviews, seller identities, account-linked data or location-specific inventory.
- Assess intellectual-property restrictions on storing and redistributing product text, images and reviews.
- Document an escalation path for cease requests and stop collection promptly when required.
- Restrict access to credentials, proxy accounts and any personal data in raw captures.
Troubleshooting common failures
Empty HTML or missing price
Cause: the value is rendered by JavaScript, selected through a variant control or delivered after an API call. Fix: inspect the initial response and embedded JSON, then enable a permitted browser-rendering path or extract the underlying request. Record which variant and locale were selected.
Sudden spike in blocked pages
Cause: request rate, IP reputation, a changed challenge or an explicit site policy. Fix: stop aggressive retries, reduce concurrency, verify permissions and review provider proxy or browser settings. Alert on the block rate so a failed run cannot overwrite good data.
Prices parse incorrectly
Cause: decimal and thousands separators, currency symbols, unit pricing or sale-price markup changed. Fix: retain the original text, parse with locale-aware rules, validate against currency and expected ranges, and version the parser.
Pagination misses products
Cause: infinite scroll, cursor-based APIs or duplicate canonical URLs. Fix: follow the site’s documented navigation signals, persist cursors, deduplicate by stable product ID and compare the observed count with catalog totals where available.
Recommended Free Tools
Costs rise without more usable data
Cause: rendering every page, repeated retries, low-quality proxies or a high proportion of challenge pages. Fix: route simple pages through ordinary retrieval, reserve browsers for JavaScript cases, cache stable responses and report cost per successful record rather than raw request volume.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.When screenshots help retail analytics
Structured fields are best for alerts and dashboards, but a visual record can resolve disputes about badges, crossed-out prices, merchandising order or a consent state. For screenshot capture, ScreenshotNeo is the first option to try because it removes consent banners, popups and chat widgets before capture, and bills only clean shots.
A do-it-yourself browser workflow can launch a browser, set the required viewport and locale, wait for the product selector, dismiss a consent dialog, capture the page and store the image beside the structured record. Keep screenshots subject to the same permission, retention and personal-data rules as HTML.
Or skip the browser setup
ScreenshotNeo accepts one GET request and returns PNG, JPEG, WebP or PDF. Its 63 options include full-page and CSS-element capture, device presets, custom viewport and retina scale, JavaScript and CSS, selector hiding, waits, request blocking, headers, cookies, user agent, authorization, timezone, geolocation, transparent backgrounds, resizing, caching, signed links, asynchronous webhooks, bulk capture of up to 100 URLs per call and a usage API. An MCP server provides take_screenshot, get_page_info and capture_pdf tools for Claude, Cursor and other MCP clients.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
See the ScreenshotNeo documentation for parameters and response headers. In Python:
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
In Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
Bot checks and CAPTCHAs, blank pages, timeouts, failed loads and cache hits are not billed; the response identifies the page verdict and billing status with X-Page-Verdict and X-Billed headers. Plans include 1,000 screenshots per month free without a card; paid plans start at $5 for 3,000 shots, with Growth at $15 for 15,000, Pro at $39 for 60,000, Scale at $99 for 250,000 and Business at $249 for 1,000,000. Yearly billing gives two months free, and every feature is included on every plan. Create a free ScreenshotNeo account to start with the 1,000 monthly screenshots.
Best Value
Bottom-line selection
Choose the operating model that matches your constraint: a managed API for speed and maintained access, Scrapy for maximum code ownership, or Apify for cloud orchestration across reusable jobs. Pilot the same retailer and marketplace set with explicit field and quality metrics, keep evidence for every observation, and make permission and stop procedures part of the design rather than an afterthought.
Frequently Asked Questions
Should I scrape product pages or marketplace search results?
Use product pages for authoritative variant, availability and offer details; use search or category pages for discovery and assortment coverage. Many programs use both, linking discovered URLs to a canonical product record.
How often should a price-monitoring job run?
Set frequency from the business decision: alerting may need hourly checks, while catalog enrichment may need daily or weekly refreshes. Start with the minimum cadence that detects a meaningful change and increase it only when measured value justifies the added load and cost.
Can a scraper legally collect competitor prices?
Legality depends on the target’s terms, jurisdiction, data type, access method and your use. Review permissions, robots directives, privacy obligations and intellectual-property limits for each target, and stop when a site requests cessation or your legal review identifies unacceptable risk.
What is the most important vendor comparison metric?
Cost per successful, complete record is more useful than requests or credits alone. Pair it with field completeness, block rate, latency and the engineering hours needed to keep parsers working.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.
Do these 3 things before closing this tab:
1Clear out junk files and repair common Windows errors2Fix the driver behind crashes, sound loss and screen glitches3Repair Windows errors before they cause bigger problems

