Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
Use the site’s own API, feed, search endpoint, or bulk export before scraping HTML. It is usually faster for your application and cheaper for the site than crawling pages. When no suitable endpoint exists, choose a hosted scraping API or run a controlled crawler, authenticate securely, respect robots.txt and terms, throttle requests, and validate every record before storing it.
1. Choose the least-invasive access path
Start by looking for a documented API, search endpoint, RSS or Atom feed, sitemap, downloadable export, or partner data feed. Scrapy’s optimization guidance says that “An API, a bulk export or a search endpoint is both faster for you and cheaper for the website than crawling its pages.” An official endpoint also gives you clearer schemas, authentication rules, pagination, and change notices than screen-scraping.
Check these sources first
- Developer or API documentation linked from the site footer.
- Network requests made by the site’s own search or export controls, provided using them is permitted.
robots.txt, sitemaps, feeds, and public download pages.- Terms of service, privacy notices, authentication requirements, and data-licensing conditions.
Do not treat a publicly visible page as permission to collect personal, copyrighted, or access-controlled data. Scope your fields, retention period, and purpose before writing a crawler.
Do these 3 things before closing this tab:
1Fix the driver behind crashes, sound loss and screen glitches2Clear out junk files and repair common Windows errors3Scan for outdated or missing drivers - takes under a minute2. Hosted scraping API or your own crawler?
A hosted service runs workers, browsers, proxies, retries, storage, and scheduling for you. Self-hosting with a framework such as the Scrapy framework gives finer control but leaves infrastructure, monitoring, upgrades, and operational compliance to your team.
#1 Best Overall
| Decision factor | Hosted API | Self-hosted crawler |
|---|---|---|
| Coverage | Depends on supported domains, proxy capacity, and anti-bot handling. | You choose request and browser integrations, subject to your infrastructure and the site’s controls. |
| JavaScript rendering | Available only when the service explicitly provides browser rendering. | You can add a browser worker, but it increases memory, latency, and maintenance. |
| Control | Review support for headers, cookies, selectors, pagination, retries, and schemas. | Full control over requests, callbacks, concurrency, delays, and parsing. |
| Operations | Provider manages capacity and platform upgrades; you monitor runs and usage. | You manage workers, queues, proxies, alerts, and upgrades. |
| Output | Many provide JSON, CSV, JSONL, datasets, or webhooks; verify each service. | You define storage and exports. |
| Scheduling | Often built in. | Requires your scheduler and job controls. |
| Cost | Compare request or result charges with engineering and infrastructure time. | Infrastructure and maintenance costs are yours; no general cost average is authoritative. |
For a managed option, verify that it supports tool discovery, synchronous and asynchronous runs, run-status polling, dataset-item export, and schedules if those are requirements. Never assume an HTML-only API can execute JavaScript.
3. Authenticate without leaking credentials
Create an API key only through the provider’s documented account flow. Send it in the prescribed authorization header or request field, and keep it in a server-side secret store or worker environment variable. Do not put keys in browser JavaScript, screenshots, public repositories, logs, or URLs that may be recorded by proxies.
Typical request sequence for a managed API
- Discover the provider’s tool or actor and read its input schema.
- Submit the target URL, fields, pagination limits, and rendering settings.
- Record the returned run ID.
- Poll status until the run succeeds, fails, or reaches a documented timeout.
- Fetch dataset rows as JSON, CSV, or JSONL where supported.
For POST requests that may be retried, use an idempotency key when the API supports one. Retry idempotent GET requests; do not blindly replay a state-changing request.
The Tool Desk
Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →4. A self-hosted Scrapy implementation
Scrapy downloads a response, passes it to a callback that extracts fields, and lets callbacks yield more requests for pagination or detail pages. A minimal spider might look like this:
import scrapy
class ProductsSpider(scrapy.Spider):
name = "products"
allowed_domains = ["example.com"]
start_urls = ["https://example.com/products"]
custom_settings = {
"CONCURRENT_REQUESTS_PER_DOMAIN": 2,
"DOWNLOAD_DELAY": 1.0,
"AUTOTHROTTLE_ENABLED": True,
"ROBOTSTXT_OBEY": True,
}
def parse(self, response):
for card in response.css("article.product"):
yield {
"name": card.css("h2::text").get(default="").strip(),
"price": card.css(".price::text").get(default="").strip(),
"url": response.urljoin(card.css("a::attr(href)").get()),
}
next_url = response.css("a.next::attr(href)").get()
if next_url:
yield response.follow(next_url, callback=self.parse)
Scrapy advises reading the target site’s robots.txt and translating crawl-delay or request-rate directives into downloader settings; Scrapy does not automatically apply every such directive. Keep concurrency conservative and increase it only while latency and response codes remain healthy.
Extracting JSON instead of HTML
If a documented endpoint returns JSON, request that endpoint directly and validate its schema. If the endpoint is an internal request discovered in a browser, confirm that using it is allowed and that you are not bypassing authentication or access controls. Parse typed fields, retain source URLs and timestamps, and record pagination cursors so an interrupted run can resume.
5. JavaScript-heavy pages
First determine whether the required data is available from a documented API. If it is, use that endpoint rather than rendering a page. When rendering is genuinely required, select a crawler or hosted service that explicitly supports browser execution and document the extra latency, resource use, and terms constraints. Wait for a selector, a known network-idle condition, or a bounded delay; do not use an unbounded sleep.
6. Throttling, blocking, and retries
Begin with low concurrency and a delay between requests. Monitor response times, status codes, and ban-page content. Rising 429, 503, or challenge-page counts indicate that your rate or behavior is not tolerated. Reduce concurrency, increase backoff, and review the site policy rather than attempting to evade controls.
HTTP error policy
| Signal | Meaning | Action |
|---|---|---|
| 401 | Authentication failed or is missing. | Check the key, authorization header, scope, and expiration; do not retry unchanged. |
| 403 | Access is forbidden or policy conditions are unmet. | Stop and review permission and terms. |
| 429 | Rate limit or burst limit reached. | Honor Retry-After when present, apply exponential backoff, and lower concurrency. |
| 5xx | Server or upstream failure. | Retry idempotent requests with bounded backoff; capture the response for diagnosis. |
| 200 with a challenge page | A bot check replaced the expected content. | Classify the result as unusable, stop escalating requests, and seek an approved access path. |
7. Validate and store reliable data
- Require key fields and verify their types, formats, and ranges.
- Detect duplicate records using a stable source identifier where available.
- Check that pagination is complete and that cursors do not repeat.
- Store retrieval timestamps, source URLs, parser version, and relevant response metadata.
- Keep raw responses or content hashes when reproducibility or auditability matters.
- Separate failed, partial, and empty results from valid zero-result datasets.
Normalize only after preserving the raw value. This lets you correct a parser without downloading the entire site again.
8. Performance, reliability, and cost decisions
Measure latency, error rates, records per successful request, and storage growth for your own workload; no universal success rate, speed figure, or cost average applies across sites. Browser rendering consumes more CPU and memory than a direct HTTP request. Caching immutable pages and using conditional requests can reduce load, but follow the provider’s and site’s cache rules. For recurring jobs, use checkpoints, bounded retries, a dead-letter queue for failed URLs, and alerts for sudden drops in item count or schema changes.
Or skip the browser setup
ScreenshotNeo is a website screenshot API and MCP server for developers. It is useful when your workflow needs a visual record rather than parsed fields: one GET request returns PNG, JPEG, WebP, or PDF, with options for full-page and element captures, device and retina settings, custom CSS or JavaScript, waiting conditions, headers and cookies, blocking rules, geolocation, resizing, caching, signed links, asynchronous webhooks, bulk capture, and more.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Cookie and consent banners, newsletter popups, and chat widgets are removed before capture. Bot checks, blank pages, failed loads, timeouts, and cache hits are not billed, and response headers identify the page verdict and billing result. Its MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients.
cURL:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
Python:
import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"}, timeout=90)
open("shot.webp", "wb").write(r.content)
Node.js:
const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://stripe.com' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);
See the ScreenshotNeo API documentation for parameters and response headers. The Free plan includes 1,000 screenshots per month with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.9. Troubleshooting checklist
Empty or missing fields
Confirm the selector or JSON path against the actual response, check whether content is loaded after JavaScript execution, and save a sample response for inspection.
Pagination stops early
Log every cursor or next URL, detect repeated cursors, and verify the final page condition. A successful HTTP status does not prove that all pages were visited.
Frequent timeouts
Reduce concurrency, set a realistic per-request timeout, avoid rendering when an API endpoint exists, and separate slow domains into their own queue.
Unexpected duplicates
Use a stable source ID or canonical URL, normalize trailing slashes consistently, and make writes idempotent.
Credentials appear in logs
Rotate the key immediately, scrub request URLs and exception messages, and move authentication to a secret-managed header or environment variable.
Frequently Asked Questions
Can an API scrape any website?
No. Availability, authentication, robots.txt directives, terms, rate limits, and technical defenses determine what access is allowed and practical.
Should I use browser automation for every page?
No. Use a documented data endpoint whenever it provides the fields you need; render only when the page itself is the required source.
What should I do with a 429 response?
Treat it as a backoff signal: honor Retry-After, reduce concurrency, increase delay, and retry only according to the site or API policy.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

