Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

You can collect Algolia results by reproducing an authorized site’s search-only request, then paginating within a fixed limit and storing provenance for every response. A browser-visible search key is not permission to republish the site’s records: obtain the owner’s authorization, keep Admin and indexing credentials off the client, and respect contracts, privacy rules, copyright, robots directives, and applicable law.

What you are actually scraping

Algolia is a hosted index and search API. A site selects records, uploads them to an index, configures relevance, and queries that index from an API client or an InstantSearch interface. The HTML search page is usually only a presentation layer; the useful data is in the JSON response returned by Algolia.

That response can be incomplete by design. Ranking, filters, permissions, omitted attributes, replicas, and update timing all affect what you see. Treat a collection as a snapshot of an authorized search view, not automatically as the site’s complete database.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Check permission and define a bounded job

Before opening developer tools, document who authorized the collection and what you may retain or reuse. Write down the exact index, query families, filters, fields, page range, refresh interval, retention period, and allowed downstream use. Algolia’s Terms of Service (last updated January 12, 2026) govern use of Algolia services, while the target owner’s terms and your contract determine whether copying that target’s records is allowed.

  • Get written permission or a contract for the target collection.
  • Limit fields to those required for the stated purpose.
  • Set a maximum page count and a stop condition before running.
  • Choose a refresh interval and retention period in advance.
  • Plan how you will honor corrections, deletions, and takedown requests.

Find the authorized search request

  1. Open the target’s search page. Use browser developer tools, open the Network panel, filter for requests containing query or the Algolia host, and perform one ordinary search.
  2. Record the request shape. Capture the application ID, index name, search-only key, request endpoint, query body, filters, facets, page settings, and the response attributes. Copy the request as cURL if the browser offers that option.
  3. Confirm that it is search-only. Never copy an Admin or indexing key into a script distributed to users. If the site uses a backend endpoint instead of a direct Algolia request, use that endpoint only when the owner has authorized it.
  4. Reproduce one request first. Compare your response with the browser response for the same query before adding pagination or concurrency.

InstantSearch interfaces commonly expose a search box, hits, pagination, refinements, and a configurable hits-per-page value. Reproducing the request—not clicking every rendered result—is usually more stable and sends fewer requests.

Choose a collection architecture

Approach Best use Credential exposure Main trade-off
Direct search request A small, authorized, read-only job Search-only key is visible to the script Simple, but subject to the target’s key restrictions and limits
Backend proxy Per-user controls, auditing, or stronger filtering Secrets stay on your server You must operate rate limits, caching, logging, and an API layer
Algolia Crawler or DocSearch Indexing content you own Uses your own indexing workflow Not a general-purpose export of another company’s index

For owner-operated systems, Algolia says it does not search your source systems directly; you upload the relevant data into an index. Use the indexing API, Crawler, or DocSearch rather than extracting your own rendered results.

Python: a bounded, cached collector

The script below uses the exact endpoint copied from your authorized request. Set ALGOLIA_SEARCH_ENDPOINT to that URL, along with the application ID, search-only key, and index name. It requests only selected attributes, stops when Algolia reports no more pages, caches identical requests, and writes a provenance record beside the normalized hits.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
import hashlib
import json
import os
import time
from datetime import datetime, timezone
from pathlib import Path

import requests

ENDPOINT = os.environ["ALGOLIA_SEARCH_ENDPOINT"]
APP_ID = os.environ["ALGOLIA_APP_ID"]
SEARCH_ONLY_KEY = os.environ["ALGOLIA_SEARCH_ONLY_KEY"]
INDEX_NAME = os.environ["ALGOLIA_INDEX_NAME"]
QUERY = os.getenv("ALGOLIA_QUERY", "")
FILTERS = os.getenv("ALGOLIA_FILTERS", "")
MAX_PAGES = int(os.getenv("ALGOLIA_MAX_PAGES", "10"))
HITS_PER_PAGE = min(int(os.getenv("ALGOLIA_HITS_PER_PAGE", "50")), 1000)
ATTRIBUTES = ["objectID", "title", "url"]
CACHE = Path("algolia-cache")
CACHE.mkdir(exist_ok=True)

session = requests.Session()
session.headers.update({
    "X-Algolia-Application-Id": APP_ID,
    "X-Algolia-API-Key": SEARCH_ONLY_KEY,
    "Content-Type": "application/json",
})

all_hits = []
retrieved_at = datetime.now(timezone.utc).isoformat()

for page in range(MAX_PAGES):
    payload = {
        "indexName": INDEX_NAME,
        "params": {
            "query": QUERY,
            "page": page,
            "hitsPerPage": HITS_PER_PAGE,
            "attributesToRetrieve": ATTRIBUTES,
        },
    }
    if FILTERS:
        payload["params"]["filters"] = FILTERS

    cache_key = hashlib.sha256(json.dumps(payload, sort_keys=True).encode()).hexdigest()
    cache_file = CACHE / f"{cache_key}.json"
    if cache_file.exists():
        response_json = json.loads(cache_file.read_text())
    else:
        for attempt in range(5):
            response = session.post(ENDPOINT, json=payload, timeout=30)
            if response.status_code != 429 and response.status_code < 500:
                response.raise_for_status()
                response_json = response.json()
                cache_file.write_text(json.dumps(response_json))
                break
            time.sleep(2 ** attempt)
        else:
            raise RuntimeError("Algolia remained unavailable after retries")

    all_hits.extend(response_json.get("hits", []))
    if page + 1 >= response_json.get("nbPages", 0):
        break
    time.sleep(0.5)

record = {
    "retrievedAt": retrieved_at,
    "endpoint": ENDPOINT,
    "applicationId": APP_ID,
    "index": INDEX_NAME,
    "query": QUERY,
    "filters": FILTERS,
    "pagesRequested": page + 1,
    "hitCount": len(all_hits),
    "hits": all_hits,
}
Path("algolia-results.json").write_text(json.dumps(record, ensure_ascii=False, indent=2))
print(f"Saved {len(all_hits)} hits")

Install the only dependency with python -m pip install requests, export the five required variables, and run the file. The example caps the job at 10 pages and 50 hits per page; lower those values for a test. Replace ATTRIBUTES with fields that the authorized response actually exposes. A missing attribute is not a reason to request an Admin key.

Equivalent cURL request

Use the request body copied from the browser and substitute your authorized values:

curl -X POST "$ALGOLIA_SEARCH_ENDPOINT" 
  -H "X-Algolia-Application-Id: $ALGOLIA_APP_ID" 
  -H "X-Algolia-API-Key: $ALGOLIA_SEARCH_ONLY_KEY" 
  -H "Content-Type: application/json" 
  --data '{"indexName":"products","params":{"query":"keyboard","page":0,"hitsPerPage":20,"attributesToRetrieve":["objectID","title","url"]}}'

Do not put an Admin or indexing credential in a shell script that will be shared, committed, or sent to a browser.

Equivalent Node.js request

This uses the built-in fetch available in current Node.js releases:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
const endpoint = process.env.ALGOLIA_SEARCH_ENDPOINT;
const body = {
  indexName: process.env.ALGOLIA_INDEX_NAME,
  params: {
    query: process.env.ALGOLIA_QUERY || "",
    page: 0,
    hitsPerPage: 20,
    attributesToRetrieve: ["objectID", "title", "url"]
  }
};

const res = await fetch(endpoint, {
  method: "POST",
  headers: {
    "X-Algolia-Application-Id": process.env.ALGOLIA_APP_ID,
    "X-Algolia-API-Key": process.env.ALGOLIA_SEARCH_ONLY_KEY,
    "Content-Type": "application/json"
  },
  body: JSON.stringify(body)
});
if (!res.ok) throw new Error(`${res.status} ${await res.text()}`);
const data = await res.json();
console.log(JSON.stringify({ nbHits: data.nbHits, nbPages: data.nbPages, hits: data.hits }, null, 2));

Paginate without flooding the service

Use the response’s nbPages value as the upper bound. Start at page 0, request the next page only while one exists, and stop at your own maximum even if more pages are available. Keep hitsPerPage bounded and request only needed fields.

  • Cache identical query, filter, page, and field combinations.
  • Run sequentially unless the owner has approved a documented concurrency level.
  • Sleep between requests and use exponential backoff for transient 429 or 5xx responses.
  • Do not probe undocumented parameters, evade key restrictions, or attempt to defeat bot controls.
  • Record the response timestamp because index contents can change between pages.

Algolia documents HTTP 429 responses when indexing is overloaded and recommends waiting for servers to catch up. A 429 during search should likewise be treated as a signal to slow down and verify the target’s limits rather than as an invitation to rotate keys or increase concurrency.

Credentials and access controls

Search keys are designed to be public in frontend applications, but public visibility does not grant republication rights. Keep Admin and indexing keys in server-side secret storage; Algolia recommends restricting indexing credentials to the minimum permissions and keeping them secret.

When access must vary by user or expire, have a backend generate a secured key with index, filter, and validUntil restrictions, or put a backend proxy in front of Algolia. Site owners can also use rate-limited keys, bot detection, and a proxy that hides the direct search client. Never attempt to bypass those controls on a site you do not operate.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Preserve provenance and make updates reversible

Store raw responses separately from normalized records. For each request, retain the target URL or endpoint where permitted, application and index identifiers, query, filters, page number, retrieval time, a response hash, and the source record’s own identifier such as objectID. This lets you compare snapshots, explain why a record appeared, and process corrections or takedown requests without destroying the audit trail.

When Crawler or DocSearch is the better tool

If you own the website, use Algolia Crawler or DocSearch to index it instead of scraping the rendered search results. DocSearch’s guidance specifically says operators who run the scraper should create a search-only key and not share an Admin API key.

Algolia’s documented Crawler limits include a maximum document size of 10 MB, 100 manual recrawls per day, one automatic recrawl per day, and a minimum 24-hour interval between updates. The Crawler also documents a limit of 10,000 Google Analytics API requests per day. Current pricing-model applications document 10,000 indexing operations per unit or Record Unit. These are operational limits for the owner’s indexing workflow, not a quota that authorizes extraction from someone else’s index.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Troubleshooting

401 or 403 response

Check that the application ID and search-only key belong together, that the key is still valid, and that its allowed indices, referers, filters, and expiration include this request. Ask the owner for a secured key or proxy instead of trying another credential.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

200 response with no hits

Compare the query, index name, filters, and replica with the browser request. An empty result can be correct for that combination. Check that you are sending the same parameter encoding and that the chosen index has not changed.

Hits differ from the page

Copy every relevant parameter from the authorized request, including facets, numeric filters, tags, around-location settings, and attributes. The UI may also merge multiple index requests. Capture each request separately and document which one your job reproduces.

Pagination repeats or skips records

The index may change while you page through it, or your code may be mixing filters between requests. Keep the query and filters immutable for a run, record retrieval times, cache responses, and rerun the snapshot when consistency matters.

429, timeout, or intermittent 5xx

Reduce page size and concurrency, add delay and exponential backoff, and honor the owner’s limits. Do not parallel-flood the endpoint or rotate credentials to evade throttling.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The response omits a field you need

Ask the owner to expose that attribute or provide an export. A search-only response cannot be turned into a complete record by changing credentials without authorization.

Or skip the browser setup

If your goal is a visual capture of an authorized search page rather than structured Algolia records, ScreenshotNeo provides a single HTTP request. Its cleaning step accepts cookie or consent banners and removes more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be disabled. Bot checks, CAPTCHAs, blank pages, timeouts, failed loads, and cache hits are not billed, and the response identifies the page verdict and billing status in X-Page-Verdict and X-Billed headers.

See the ScreenshotNeo API documentation for all options. This cURL call captures an authorized search URL:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://your-authorized-site.example/search?q=algolia -o shot.webp

ScreenshotNeo also offers an MCP server with take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients. The Free plan includes 1,000 shots per month with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Frequently Asked Questions

Does Algolia return every record in an index?

Not necessarily. A search response is shaped by ranking, filters, permissions, replicas, exposed attributes, and pagination, so request an export or feed when you need a complete dataset.

Should I save raw JSON as well as parsed fields?

Yes. Keeping the raw response beside normalized records preserves the exact source needed to investigate changes, corrections, or removal requests.

Can I make a collection job reproducible?

Record the endpoint, index, query, filters, page settings, retrieval time, response hash, and source identifiers, then reuse the same parameters and cache policy for later snapshots.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.