To scrape Google results responsibly, first choose an authorized data path: the Custom Search JSON API when you are eligible, or a permitted browser/HTTP collection workflow for pages you are allowed to fetch. Store each capture with its query, locale, device, timestamp and page number, then parse optional SERP modules instead of assuming every search has the same layout.
Google’s results are query-dependent and its markup changes. A reliable system therefore separates collection from normalization, keeps the raw HTML or JSON, versions selectors, and treats consent, robots instructions, rate limits and retention as part of the design—not as cleanup after a scraper breaks.
What a Google SERP contains
A search-engine-results page (SERP) is a document produced for one query context. The same words can produce different modules for different countries, languages, devices, users and times. Google says the features shown depend on the query, so your parser must accept missing modules and new markup.
Minimum result record
Normalize each organic or feature result into a record with:
Free tools Windows power users keep installed
One-click scans. No signup required.
#1 Best Overall
- position: the displayed rank or module position, preserving whether it is an organic result or a feature.
- title: the visible result title.
- url: the destination URL after permitted redirect resolution.
- display_domain: the source domain shown to the user.
- snippet: visible descriptive text, if present.
- feature_type: a value such as
organic,featured_snippet,video,newsorunknown.
Optional modules
Depending on the query, a page can also include spelling suggestions, promotions, images, videos, news, local results, knowledge panels, related questions and other blocks. Do not assign a fixed CSS position to any of them. Save the module’s raw fragment and a parser version so a later selector change can be audited.
| Layer | What to retain | Why it matters |
|---|---|---|
| Request context | Query, locale, country, language, device, timestamp, result page | Explains why two captures differ |
| Normalized records | Position, title, URL, domain, snippet, feature type | Supports ranking and reporting queries |
| Raw response | HTML or API JSON, status and headers | Allows re-parsing when markup changes |
| Parser metadata | Parser version, run ID and extraction warnings | Makes failures reproducible |
Choose a collection method
| Method | Best fit | Strengths | Trade-offs |
|---|---|---|---|
| Custom Search JSON API | Authorized, structured result data | Stable JSON fields, documented query controls and pagination | Eligibility, quota and feature-scope limits; current overview says it is closed to new customers |
| Browser automation | Rendered modules that an API does not expose | Sees the page a user sees; can capture visual evidence | Slow, resource-heavy, selector maintenance and higher compliance risk |
| HTTP plus HTML parser | Pages you are expressly permitted to fetch | Low overhead and inexpensive at small scale | May miss JavaScript-rendered content; markup changes frequently |
| Managed SERP service | Teams that need infrastructure and parser maintenance outsourced | May provide rotation, retries, rendering and geographic controls | Provider terms, provenance, retention, freshness and pricing must be checked |
Use the authorized Custom Search JSON API when eligible
Google’s documented route requires a Programmable Search Engine and an API key. The API returns JSON metadata and result items rather than the complete visual page. Google’s current overview says the API is closed to new customers; existing customers have until January 1, 2027 to transition. Confirm eligibility, quotas and pricing in Google for Developers documentation before building a new dependency.
Request setup
- Create or select a Programmable Search Engine and configure its search scope.
- Create an API key with the required API enabled and restrict that key to the services and hosts that need it.
- Put the endpoint, key and search-engine identifier in environment variables rather than source control.
- Send a GET request, check the HTTP status and preserve the complete JSON response.
- Normalize
itemswhile retaining top-level fields such asqueries,searchInformation,spellingandpromotions.
Important controls
| Parameter | Use |
|---|---|
q |
Query text |
start |
Pagination offset |
num |
Results requested per page |
safe |
Safe-search setting |
siteSearch and siteSearchFilter |
Include or exclude a site |
exactTerms and excludeTerms |
Require or omit terms |
dateRestrict |
Limit by a relative date range |
| Language and country controls | Set the intended linguistic and geographic context |
The reference documents a default page size of 10 and a maximum of 100 results for a query. Treat those as API limits, not guarantees that every query has that many results.
cURL request
curl -G "$GOOGLE_CSE_ENDPOINT"
--data-urlencode "key=$GOOGLE_API_KEY"
--data-urlencode "cx=$GOOGLE_CSE_ID"
--data-urlencode "q=site:example.com pricing"
--data-urlencode "num=10"
--data-urlencode "start=1"
-o response.json
Set GOOGLE_CSE_ENDPOINT to the endpoint shown in Google’s current REST guide. Keeping it configurable avoids baking a potentially changed endpoint into deployment code.
Do these 3 things before closing this tab:
1Fix the driver behind crashes, sound loss and screen glitches2Repair Windows errors before they cause bigger problems3Scan for outdated or missing drivers - takes under a minutePython normalization example
import os
import requests
params = {
"key": os.environ["GOOGLE_API_KEY"],
"cx": os.environ["GOOGLE_CSE_ID"],
"q": "site:example.com pricing",
"num": 10,
"start": 1,
}
r = requests.get(os.environ["GOOGLE_CSE_ENDPOINT"], params=params, timeout=30)
r.raise_for_status()
data = r.json()
for item in data.get("items", []):
print({
"title": item.get("title"),
"url": item.get("link"),
"display_domain": item.get("displayLink"),
"snippet": item.get("snippet"),
"feature_type": "organic",
})
Node.js request
const endpoint = process.env.GOOGLE_CSE_ENDPOINT;
const q = new URLSearchParams({
key: process.env.GOOGLE_API_KEY,
cx: process.env.GOOGLE_CSE_ID,
q: 'site:example.com pricing',
num: '10',
start: '1'
});
const res = await fetch(`${endpoint}?${q}`);
if (!res.ok) throw new Error(`Custom Search failed: ${res.status}`);
const data = await res.json();
for (const item of data.items ?? []) {
console.log({
title: item.title,
url: item.link,
display_domain: item.displayLink,
snippet: item.snippet,
feature_type: 'organic'
});
}
Browser automation for rendered SERP modules
Use a browser only when you have a lawful, authorized reason to fetch the page and the API does not provide the detail you need. A browser can expose rendered modules, but it also introduces consent dialogs, login state, JavaScript timing, bot checks and frequent markup changes.
Rank #2
- Define a fixed query, country, language, timezone, viewport and device profile.
- Use an authorized account or explicit permission where required; do not bypass a challenge or access control.
- Navigate with a conservative rate and wait for a documented readiness condition, not an arbitrary rapid loop.
- Save status, headers, final URL, timestamp and a screenshot or raw DOM for each capture.
- Parse semantic text and links with fallback selectors. Version every selector set and alert when expected fields disappear.
- Stop and review when a consent page, CAPTCHA, blank response or unusual redirect appears.
Selectors should be treated as versioned code. A selector that works for one query or device is not evidence that it will work for another.
HTTP plus HTML parsing
For pages you are permitted to fetch, an HTTP client can be cheaper than launching a browser. Preserve the response body before parsing and record the canonical URL, status and relevant headers. A successful HTTP response proves delivery, not permission to automate access.
import requests
from bs4 import BeautifulSoup
url = "https://www.google.com/search"
params = {"q": "site:example.com pricing", "num": "10"}
r = requests.get(url, params=params, timeout=30, headers={"User-Agent": "your-authorized-client"})
r.raise_for_status()
html = r.text
soup = BeautifulSoup(html, "html.parser")
records = []
for link in soup.select("a[href]"):
text = " ".join(link.get_text(" ", strip=True).split())
href = link.get("href")
if text and href:
records.append({"title_or_text": text, "url": href})
print(records[:10])
The selector above is intentionally only a starting point: Google’s markup is not a stable contract. Build field-specific selectors, keep fallbacks, and test against saved fixtures rather than live traffic on every code change.
Quick wins for a faster PC:
Clear out junk files and repair common Windows errorsFree Scan →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Repair Windows errors before they cause bigger problemsFix Now →Or skip the browser setup
ScreenshotNeo is useful when you need a clean visual capture of a permitted page rather than structured ranking JSON. It accepts consent banners before capture and removes more than 60 known consent platforms, newsletter popups and chat widgets; each step can be disabled. Bot checks or CAPTCHAs, blank pages, timeouts, failed loads and cache hits are not billed, and response headers report the page verdict and billing status.
One GET request returns PNG, JPEG, WebP or PDF. The API also supports full-page and element captures, device and viewport settings, dark mode, retina scale, custom CSS and JavaScript, waits, request blocking, cookies and headers, caching, signed links, asynchronous webhooks, bulk capture and an MCP server for AI agents.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp
See the ScreenshotNeo documentation for parameters. A free plan includes 1,000 screenshots each month with no card; paid plans start at $5 for 3,000 shots. It is a screenshot service, not a replacement for extracting titles and links from Google’s JSON or HTML. Sign up free.
Rank #3
Model optional features without breaking your parser
Represent a capture as an envelope and a sequence of typed modules:
{
"context": {"query": "...", "locale": "en-US", "device": "desktop", "timestamp": "...", "page": 1},
"modules": [
{"type": "organic", "position": 1, "title": "...", "url": "...", "snippet": "..."},
{"type": "featured_snippet", "position": 0, "title": "...", "url": "...", "snippet": "..."}
],
"raw_ref": "capture-identifier",
"parser_version": "2026-09-29"
}
Keep unknown modules instead of dropping them. For ranking analysis, distinguish displayed position from an internal array index: a featured block can appear above the first organic result, and some modules have no conventional rank.
Reliability, performance and cost
Control request volume
Batch work by query and locale, use bounded concurrency, and add exponential backoff for transient failures. Do not retry consent pages, CAPTCHAs or policy denials as if they were network errors.
Cache deliberately
Cache identical query-context combinations when freshness permits. Record the cache key and age so reports do not imply that a cached page is a new observation.
Measure the pipeline
Track latency, status codes, empty-result rate, field completeness, parser warnings and the percentage of captures classified as blocked or challenged. Keep raw responses for the shortest period that meets your audit needs.
Understand API economics
For existing Custom Search JSON API customers, Google’s overview reported 100 free queries per day and additional requests at $5 per 1,000 up to 10,000 per day; that page was crawled seven months ago, so verify current pricing and availability before budgeting. Browser and managed-service costs vary with rendering, proxy, storage and retry policies; obtain current terms directly from the provider.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Compliance and data governance
Google’s current Terms of Service prohibit automated access that violates machine-readable instructions such as robots.txt and describe scraping content that does not belong to the user as conduct that can cause harm or liability. Before collecting, document:
- the lawful purpose and permission or contractual basis;
- robots instructions and site-specific terms;
- request-rate and concurrency limits;
- what personal data is minimized or removed;
- retention, access controls and deletion procedures;
- the query, locale, timestamp and parser version stored with every record.
If you need Google results at scale, prefer an authorized API or a provider whose contract explicitly covers your intended use. Re-check Google’s API overview and Terms before deployment because eligibility, quotas and wording can change.
Troubleshooting common failures
| Symptom | Likely cause | Fix |
|---|---|---|
| API returns an authorization error | Invalid key, wrong search-engine ID or unavailable product access | Check credentials and eligibility; do not switch to unauthorized automation automatically |
| Fewer than 10 API items | The query has fewer matches or filtering removed items | Inspect searchInformation and pagination; do not assume a missing item is a parser bug |
| HTML parser returns no titles | Markup changed, a consent page was returned or content requires JavaScript | Save the raw response, classify the page, update versioned selectors or use an authorized browser/API path |
| Repeated CAPTCHA or bot-check page | Traffic pattern, account state or policy enforcement | Stop retries, review authorization and rate limits, and contact the provider if appropriate |
| Results differ between runs | Locale, device, time, personalization or query-dependent features changed | Pin context fields and compare captures with their timestamps |
| Browser run is slow or flaky | Heavy rendering, arbitrary sleeps or unbounded concurrency | Wait on explicit conditions, limit concurrency, block unnecessary resources only when permitted, and record timing |
FAQ
Can I assume every query has a featured snippet?
No. Google states that result features vary with the query. Model the feature as optional and preserve an unknown type when a new module appears.
Recommended Free Tools
How many results can the Custom Search JSON API return?
Google’s reference documents 10 results per page by default and no more than 100 results for a query. Quotas, filters and query coverage can produce fewer.
Should I store only parsed fields?
No. Keep raw HTML or API JSON with parser and request metadata so you can audit a change and reprocess records without fetching again.
Is a managed SERP provider automatically compliant?
No. Review its contract, geographic coverage, rate limits, provenance, retention and whether it expressly permits your intended collection.
Frequently Asked Questions
Can a screenshot service provide structured Google rankings?
No. A screenshot service captures the rendered page; extract titles, links and snippets with the API or an authorized HTML/ browser parser.
PC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchWhat should I pin for reproducible SERP data?
Record the exact query, locale, country, language, device, timestamp, page number and parser version for every capture.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.




