To scrape search engines with an API, send your query to a search-results endpoint, authenticate with the provider’s credentials, parse the returned JSON, and paginate with the next-page information supplied by the response. This avoids running a browser for every query, but provider choice matters: Google’s Custom Search JSON API is closed to new customers and scheduled for discontinuation on January 1, 2027, while Bing and managed SERP services have different quotas, localization, terms and pricing.
What API-based search scraping actually does
An API scraper does not read pixels from a browser window. Your application submits parameters such as a query, API key, search-engine identifier, language or location to an HTTP endpoint. The service performs the search and returns structured JSON that your code can store or transform.
A normalized record usually contains the result title, destination URL, snippet, rank, language, geography and the time of collection. Keep the provider name, API version, query parameters and timestamp with every record so a later run can be compared with the original.
API output is generally easier to process than HTML, but it is not automatically identical to the results a person sees in a browser. Index coverage, personalization, device, country, language, freshness, safe-search settings and provider-specific ranking rules all affect the response.
#1 Best Overall
Google Custom Search JSON API: requirements and lifecycle
Google’s documented request endpoint is https://www.googleapis.com/customsearch/v1. Google states that you need a configured Programmable Search Engine and an API key. The cx value identifies that engine, while q contains the query.
Important availability dates
Existing Custom Search JSON API customers receive 100 free queries per day. Additional usage is documented at $5 per 1,000 queries, with a limit of 10,000 queries per day. Google says the API is closed to new customers and will be discontinued on January 1, 2027. These are time-sensitive commercial and availability conditions; confirm the current status before committing a new production system.
What the response contains
The JSON response includes search metadata and result items. Metadata can tell you whether another page is available; each item contains the fields needed for a result list, such as title, link and snippet. Treat absent fields as normal and code defensively rather than assuming every result has the same shape.
Rank #2
Build a Google API scraper step by step
- Configure the search engine. Create a Programmable Search Engine and set the sites or web coverage you require. Record its engine ID, called
cx. - Create or obtain an API key. Keep it in a server-side secret store or an environment variable. Never put a production key in browser JavaScript, a public repository or a mobile application.
- Send a GET request. Provide
key,cxandq. URL-encode the query so spaces, punctuation and non-ASCII text are transmitted correctly. - Validate the response. Check the HTTP status and any API error object before reading
items. A successful request can still contain zero results. - Normalize and store results. Save rank, title, URL, snippet, query, locale, collection time and provider. This makes deduplication and later audits possible.
- Paginate deliberately. When the response supplies a next-page query role or URL, follow that value rather than constructing page numbers yourself. Google documents a 100-result maximum, so do not assume an unlimited result set.
cURL
curl -G "https://www.googleapis.com/customsearch/v1"
--data-urlencode "key=$GOOGLE_API_KEY"
--data-urlencode "cx=$GOOGLE_SEARCH_ENGINE_ID"
--data-urlencode "q=privacy focused analytics"
Set the two environment variables in your shell or secret manager before running the command. The response is JSON on standard output, so redirect it to a file or pipe it to a JSON processor.
Free tools Windows power users keep installed
One-click scans. No signup required.
Python
import os
import requests
endpoint = "https://www.googleapis.com/customsearch/v1"
params = {
"key": os.environ["GOOGLE_API_KEY"],
"cx": os.environ["GOOGLE_SEARCH_ENGINE_ID"],
"q": "privacy focused analytics",
}
response = requests.get(endpoint, params=params, timeout=30)
response.raise_for_status()
data = response.json()
for rank, item in enumerate(data.get("items", []), start=1):
print({
"rank": rank,
"title": item.get("title"),
"url": item.get("link"),
"snippet": item.get("snippet"),
})
# Inspect data.get("queries", {}) for the provider's next-page information.
Use a session with connection pooling for repeated requests, and catch timeout, connection and rate-limit exceptions separately from permanent authentication errors.
Node.js
const endpoint = new URL('https://www.googleapis.com/customsearch/v1');
endpoint.search = new URLSearchParams({
key: process.env.GOOGLE_API_KEY,
cx: process.env.GOOGLE_SEARCH_ENGINE_ID,
q: 'privacy focused analytics'
});
const response = await fetch(endpoint);
const data = await response.json();
if (!response.ok) {
throw new Error(`Search API failed (${response.status}): ${JSON.stringify(data)}`);
}
for (const [index, item] of (data.items || []).entries()) {
console.log({
rank: index + 1,
title: item.title,
url: item.link,
snippet: item.snippet
});
}
// Use the next-page information in data.queries when another page is needed.
Pagination, quotas and reproducibility
Follow the provider’s next-page instruction
Do not guess that a page is available merely because the previous page was full. Read the response metadata and follow the documented next-page query role or URL. Stop when it is absent, when you have reached your application’s result limit, or when the provider’s documented maximum has been reached.
Rank #3
- How search engines work: Because knowing your enemy is half the battle.
- SEO Basics: Like how to start a website and submit it to Google.
- Keyword Research: Find the most promising keywords for your business.
- SEO Content: Create content that even search engines want to binge-read.
- On-Page SEO: The most effective way to explain your pages to search engines.
Control request volume
- Count requests by API key, project, query and day.
- Cache repeat queries when freshness requirements permit.
- Deduplicate simultaneous jobs so one popular query does not consume quota repeatedly.
- Set a hard per-job page limit and an overall daily budget.
- Record HTTP status, provider error code and retry count for every attempt.
Make runs repeatable
Store the exact query string, locale, device or user-agent setting, safe-search choice, provider, API version and collection timestamp. Search rankings change, so a later request is a new observation, not a guaranteed replay of the old one.
Choosing a provider
| Option | Best fit | What to verify |
|---|---|---|
| Google Custom Search JSON API | Google-hosted programmable search for existing customers | Requires an API key and cx; 100 free queries per day for existing customers, documented $5 per 1,000 additional queries up to 10,000 per day; closed to new customers and scheduled for discontinuation January 1, 2027. |
| Bing Web Search API | Microsoft-hosted web results and JSON responses | Microsoft documents query parameters, headers, response objects and terms/display requirements. Check the current subscription, quota and regional conditions directly. |
| Managed SERP API such as SerpApi | Multi-engine extraction, localization and anti-bot operations | A 2026 TechRadar Pro review describes location search, proxies, CAPTCHA handling, a 100-search free tier and a 5,000-search/$75 base plan. Treat those figures as review-era information and verify current pricing and limits before purchase. |
| Search Researcher Result API | Eligible research use cases involving search-result analysis | Google says access requires eligibility and an application; it is not a generally available replacement for every web-search workload. |
Compare index coverage, geographic localization, freshness, structured fields, pagination depth, quotas, latency, error behavior, retention, total cost and display obligations. A service that returns browser-emulated HTML has different extraction and compliance characteristics from one that returns normalized JSON.
Reliability and error handling
Retry only transient failures
Use exponential backoff with jitter for timeouts, connection resets and provider responses that explicitly indicate temporary throttling. Do not blindly retry invalid keys, an unknown cx, malformed parameters or an account that has exceeded a hard quota. Those errors require configuration or billing changes.
Rank #4
Typical symptoms and fixes
- 401 or 403: verify the key, API enablement, restrictions and project billing state; ensure the key is being sent server-side.
- Invalid engine identifier: check that
cxbelongs to the intended Programmable Search Engine and contains no copied whitespace. - 400 malformed request: inspect URL encoding and required parameters. Log the final parameter names, but never log the secret value.
- 429 or quota error: slow concurrent workers, honor retry guidance, cache repeated searches and review daily and per-minute limits.
- Empty
items: distinguish a valid zero-result response from an error object; record the query and settings for diagnosis. - Results differ from a browser: compare country, language, personalization, device and time. API ranking is not promised to match an interactive session.
- Pagination stops early: inspect the response’s next-page metadata and the provider’s maximum-result rule instead of incrementing a page counter indefinitely.
Compliance and responsible use
Read the selected provider’s terms, acceptable-use limits and display requirements. Microsoft’s documentation explicitly points developers to these obligations. Store only the data your application needs, protect API credentials, and define retention and deletion rules for collected snippets and URLs.
Do not treat an API as universal legal permission to scrape. The documented product mechanics and terms references do not answer jurisdiction-specific questions about copyright, database rights, privacy, contract or automated access. Obtain advice for your use case and location when those issues matter.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.When browser automation is still necessary
Use an API when you need structured fields, predictable authentication, server-side scheduling and easy pagination. A browser may still be appropriate when you must test the exact rendered page, interact with a form, observe client-side behavior or capture a visual record. Those are different outputs: JSON for analysis versus pixels or a PDF for presentation and QA.
Recommended Free Tools
Best Value
Or skip the browser setup
If your task is to capture a rendered search-results page as an image or PDF—not to obtain structured SERP JSON—ScreenshotNeo provides a website screenshot API and MCP server. It accepts a URL, handles the browser work and returns PNG, JPEG, WebP or PDF. Before capture it accepts cookie or consent banners and removes more than 60 known consent platforms, newsletter popups and chat widgets; each cleanup step can be turned off.
Only clean shots are billed. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads and cache hits cost nothing, and the response identifies the result with X-Page-Verdict and X-Billed headers. Its MCP server exposes take_screenshot, get_page_info and capture_pdf to Claude, Cursor and other MCP clients.
For the full parameter list, see the ScreenshotNeo API documentation. This one-call example captures a search page; it does not replace a SERP API when you need result objects:
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://www.google.com/search?q=privacy+focused+analytics -o shot.webp
ScreenshotNeo’s Free plan includes 1,000 shots per month with no card. Paid plans start at $5 for 3,000 shots, and every feature is available on every plan. Create a free ScreenshotNeo account to try it.
Quick wins for a faster PC:
Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Repair Windows errors before they cause bigger problemsFix Now →Implementation checklist
- Choose a provider whose index, geography and terms fit the workload.
- Keep keys server-side and rotate them if exposed.
- Validate status codes and provider error objects before parsing items.
- Normalize title, URL, snippet, rank, locale, provider and timestamp.
- Follow returned next-page metadata and enforce a 100-result ceiling where Google’s API applies.
- Cache eligible repeat queries and monitor quotas before production launch.
- Use retries with backoff only for transient failures.
- Document retention, display and acceptable-use decisions.
Frequently Asked Questions
Can I call a search API directly from a browser application?
You can technically make a client request only when the provider supports it and the credential is safe to expose. For production systems, keep the key behind your own server or a protected backend endpoint.
Why should I save the query settings with each result?
Search output changes with locale, device, language, provider and time. Recording those inputs lets you explain differences and audit how a dataset was collected.
Is a screenshot service a substitute for a SERP API?
No. A SERP API returns structured result data for analysis; a screenshot service returns a visual rendering. Choose based on the output your application actually needs.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.
Do these 3 things before closing this tab:
1Repair Windows errors before they cause bigger problems2Scan for outdated or missing drivers - takes under a minute3Clear out junk files and repair common Windows errors

