Recommended Free Tools
Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.
You can collect Algolia results by reproducing an authorized site’s search-only request, then paginating within a fixed limit and storing provenance for every response. A browser-visible search key is not permission to republish the site’s records: obtain the owner’s authorization, keep Admin and indexing credentials off the client, and respect contracts, privacy rules, copyright, robots directives, and applicable law.
What you are actually scraping
Algolia is a hosted index and search API. A site selects records, uploads them to an index, configures relevance, and queries that index from an API client or an InstantSearch interface. The HTML search page is usually only a presentation layer; the useful data is in the JSON response returned by Algolia.
That response can be incomplete by design. Ranking, filters, permissions, omitted attributes, replicas, and update timing all affect what you see. Treat a collection as a snapshot of an authorized search view, not automatically as the site’s complete database.
Check permission and define a bounded job
Before opening developer tools, document who authorized the collection and what you may retain or reuse. Write down the exact index, query families, filters, fields, page range, refresh interval, retention period, and allowed downstream use. Algolia’s Terms of Service (last updated January 12, 2026) govern use of Algolia services, while the target owner’s terms and your contract determine whether copying that target’s records is allowed.
#1 Best Overall
- Get written permission or a contract for the target collection.
- Limit fields to those required for the stated purpose.
- Set a maximum page count and a stop condition before running.
- Choose a refresh interval and retention period in advance.
- Plan how you will honor corrections, deletions, and takedown requests.
Find the authorized search request
- Open the target’s search page. Use browser developer tools, open the Network panel, filter for requests containing query or the Algolia host, and perform one ordinary search.
- Record the request shape. Capture the application ID, index name, search-only key, request endpoint, query body, filters, facets, page settings, and the response attributes. Copy the request as cURL if the browser offers that option.
- Confirm that it is search-only. Never copy an Admin or indexing key into a script distributed to users. If the site uses a backend endpoint instead of a direct Algolia request, use that endpoint only when the owner has authorized it.
- Reproduce one request first. Compare your response with the browser response for the same query before adding pagination or concurrency.
InstantSearch interfaces commonly expose a search box, hits, pagination, refinements, and a configurable hits-per-page value. Reproducing the request—not clicking every rendered result—is usually more stable and sends fewer requests.
Choose a collection architecture
| Approach | Best use | Credential exposure | Main trade-off |
|---|---|---|---|
| Direct search request | A small, authorized, read-only job | Search-only key is visible to the script | Simple, but subject to the target’s key restrictions and limits |
| Backend proxy | Per-user controls, auditing, or stronger filtering | Secrets stay on your server | You must operate rate limits, caching, logging, and an API layer |
| Algolia Crawler or DocSearch | Indexing content you own | Uses your own indexing workflow | Not a general-purpose export of another company’s index |
For owner-operated systems, Algolia says it does not search your source systems directly; you upload the relevant data into an index. Use the indexing API, Crawler, or DocSearch rather than extracting your own rendered results.
Python: a bounded, cached collector
The script below uses the exact endpoint copied from your authorized request. Set ALGOLIA_SEARCH_ENDPOINT to that URL, along with the application ID, search-only key, and index name. It requests only selected attributes, stops when Algolia reports no more pages, caches identical requests, and writes a provenance record beside the normalized hits.
Do these 3 things before closing this tab:
1Scan for outdated or missing drivers - takes under a minute2Repair Windows errors before they cause bigger problems3Fix the driver behind crashes, sound loss and screen glitchesimport hashlib
import json
import os
import time
from datetime import datetime, timezone
from pathlib import Path
import requests
ENDPOINT = os.environ["ALGOLIA_SEARCH_ENDPOINT"]
APP_ID = os.environ["ALGOLIA_APP_ID"]
SEARCH_ONLY_KEY = os.environ["ALGOLIA_SEARCH_ONLY_KEY"]
INDEX_NAME = os.environ["ALGOLIA_INDEX_NAME"]
QUERY = os.getenv("ALGOLIA_QUERY", "")
FILTERS = os.getenv("ALGOLIA_FILTERS", "")
MAX_PAGES = int(os.getenv("ALGOLIA_MAX_PAGES", "10"))
HITS_PER_PAGE = min(int(os.getenv("ALGOLIA_HITS_PER_PAGE", "50")), 1000)
ATTRIBUTES = ["objectID", "title", "url"]
CACHE = Path("algolia-cache")
CACHE.mkdir(exist_ok=True)
session = requests.Session()
session.headers.update({
"X-Algolia-Application-Id": APP_ID,
"X-Algolia-API-Key": SEARCH_ONLY_KEY,
"Content-Type": "application/json",
})
all_hits = []
retrieved_at = datetime.now(timezone.utc).isoformat()
for page in range(MAX_PAGES):
payload = {
"indexName": INDEX_NAME,
"params": {
"query": QUERY,
"page": page,
"hitsPerPage": HITS_PER_PAGE,
"attributesToRetrieve": ATTRIBUTES,
},
}
if FILTERS:
payload["params"]["filters"] = FILTERS
cache_key = hashlib.sha256(json.dumps(payload, sort_keys=True).encode()).hexdigest()
cache_file = CACHE / f"{cache_key}.json"
if cache_file.exists():
response_json = json.loads(cache_file.read_text())
else:
for attempt in range(5):
response = session.post(ENDPOINT, json=payload, timeout=30)
if response.status_code != 429 and response.status_code < 500:
response.raise_for_status()
response_json = response.json()
cache_file.write_text(json.dumps(response_json))
break
time.sleep(2 ** attempt)
else:
raise RuntimeError("Algolia remained unavailable after retries")
all_hits.extend(response_json.get("hits", []))
if page + 1 >= response_json.get("nbPages", 0):
break
time.sleep(0.5)
record = {
"retrievedAt": retrieved_at,
"endpoint": ENDPOINT,
"applicationId": APP_ID,
"index": INDEX_NAME,
"query": QUERY,
"filters": FILTERS,
"pagesRequested": page + 1,
"hitCount": len(all_hits),
"hits": all_hits,
}
Path("algolia-results.json").write_text(json.dumps(record, ensure_ascii=False, indent=2))
print(f"Saved {len(all_hits)} hits")
Install the only dependency with python -m pip install requests, export the five required variables, and run the file. The example caps the job at 10 pages and 50 hits per page; lower those values for a test. Replace ATTRIBUTES with fields that the authorized response actually exposes. A missing attribute is not a reason to request an Admin key.
Equivalent cURL request
Use the request body copied from the browser and substitute your authorized values:
Rank #2
curl -X POST "$ALGOLIA_SEARCH_ENDPOINT"
-H "X-Algolia-Application-Id: $ALGOLIA_APP_ID"
-H "X-Algolia-API-Key: $ALGOLIA_SEARCH_ONLY_KEY"
-H "Content-Type: application/json"
--data '{"indexName":"products","params":{"query":"keyboard","page":0,"hitsPerPage":20,"attributesToRetrieve":["objectID","title","url"]}}'
Do not put an Admin or indexing credential in a shell script that will be shared, committed, or sent to a browser.
Equivalent Node.js request
This uses the built-in fetch available in current Node.js releases:
The Tool Desk
Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →const endpoint = process.env.ALGOLIA_SEARCH_ENDPOINT;
const body = {
indexName: process.env.ALGOLIA_INDEX_NAME,
params: {
query: process.env.ALGOLIA_QUERY || "",
page: 0,
hitsPerPage: 20,
attributesToRetrieve: ["objectID", "title", "url"]
}
};
const res = await fetch(endpoint, {
method: "POST",
headers: {
"X-Algolia-Application-Id": process.env.ALGOLIA_APP_ID,
"X-Algolia-API-Key": process.env.ALGOLIA_SEARCH_ONLY_KEY,
"Content-Type": "application/json"
},
body: JSON.stringify(body)
});
if (!res.ok) throw new Error(`${res.status} ${await res.text()}`);
const data = await res.json();
console.log(JSON.stringify({ nbHits: data.nbHits, nbPages: data.nbPages, hits: data.hits }, null, 2));
Paginate without flooding the service
Use the response’s nbPages value as the upper bound. Start at page 0, request the next page only while one exists, and stop at your own maximum even if more pages are available. Keep hitsPerPage bounded and request only needed fields.
- Cache identical query, filter, page, and field combinations.
- Run sequentially unless the owner has approved a documented concurrency level.
- Sleep between requests and use exponential backoff for transient 429 or 5xx responses.
- Do not probe undocumented parameters, evade key restrictions, or attempt to defeat bot controls.
- Record the response timestamp because index contents can change between pages.
Algolia documents HTTP 429 responses when indexing is overloaded and recommends waiting for servers to catch up. A 429 during search should likewise be treated as a signal to slow down and verify the target’s limits rather than as an invitation to rotate keys or increase concurrency.
Credentials and access controls
Search keys are designed to be public in frontend applications, but public visibility does not grant republication rights. Keep Admin and indexing keys in server-side secret storage; Algolia recommends restricting indexing credentials to the minimum permissions and keeping them secret.
When access must vary by user or expire, have a backend generate a secured key with index, filter, and validUntil restrictions, or put a backend proxy in front of Algolia. Site owners can also use rate-limited keys, bot detection, and a proxy that hides the direct search client. Never attempt to bypass those controls on a site you do not operate.
Preserve provenance and make updates reversible
Store raw responses separately from normalized records. For each request, retain the target URL or endpoint where permitted, application and index identifiers, query, filters, page number, retrieval time, a response hash, and the source record’s own identifier such as objectID. This lets you compare snapshots, explain why a record appeared, and process corrections or takedown requests without destroying the audit trail.
When Crawler or DocSearch is the better tool
If you own the website, use Algolia Crawler or DocSearch to index it instead of scraping the rendered search results. DocSearch’s guidance specifically says operators who run the scraper should create a search-only key and not share an Admin API key.
Algolia’s documented Crawler limits include a maximum document size of 10 MB, 100 manual recrawls per day, one automatic recrawl per day, and a minimum 24-hour interval between updates. The Crawler also documents a limit of 10,000 Google Analytics API requests per day. Current pricing-model applications document 10,000 indexing operations per unit or Record Unit. These are operational limits for the owner’s indexing workflow, not a quota that authorizes extraction from someone else’s index.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Troubleshooting
401 or 403 response
Check that the application ID and search-only key belong together, that the key is still valid, and that its allowed indices, referers, filters, and expiration include this request. Ask the owner for a secured key or proxy instead of trying another credential.
200 response with no hits
Compare the query, index name, filters, and replica with the browser request. An empty result can be correct for that combination. Check that you are sending the same parameter encoding and that the chosen index has not changed.
Hits differ from the page
Copy every relevant parameter from the authorized request, including facets, numeric filters, tags, around-location settings, and attributes. The UI may also merge multiple index requests. Capture each request separately and document which one your job reproduces.
Pagination repeats or skips records
The index may change while you page through it, or your code may be mixing filters between requests. Keep the query and filters immutable for a run, record retrieval times, cache responses, and rerun the snapshot when consistency matters.
429, timeout, or intermittent 5xx
Reduce page size and concurrency, add delay and exponential backoff, and honor the owner’s limits. Do not parallel-flood the endpoint or rotate credentials to evade throttling.
Free tools Windows power users keep installed
One-click scans. No signup required.
The response omits a field you need
Ask the owner to expose that attribute or provide an export. A search-only response cannot be turned into a complete record by changing credentials without authorization.
Best Value
Or skip the browser setup
If your goal is a visual capture of an authorized search page rather than structured Algolia records, ScreenshotNeo provides a single HTTP request. Its cleaning step accepts cookie or consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be disabled. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and the response identifies the page verdict and billing status in X-Page-Verdict and X-Billed headers.
See the ScreenshotNeo API documentation for all options. This cURL call captures an authorized search URL:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://your-authorized-site.example/search?q=algolia -o shot.webp
ScreenshotNeo also offers an MCP server with take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients. The Free plan includes 1,000 shots per month with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Frequently Asked Questions
Does Algolia return every record in an index?
Not necessarily. A search response is shaped by ranking, filters, permissions, replicas, exposed attributes, and pagination, so request an export or feed when you need a complete dataset.
Should I save raw JSON as well as parsed fields?
Yes. Keeping the raw response beside normalized records preserves the exact source needed to investigate changes, corrections, or removal requests.
Can I make a collection job reproducible?
Record the endpoint, index, query, filters, page settings, retrieval time, response hash, and source identifiers, then reuse the same parameters and cache policy for later snapshots.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

