October DealsAmazon USOctober deal check: compare before you payAmazon US: current deals, useful picks and tech finds.Check DealsClean PCRecommendedOne scan can reveal what keeps slowing WindowsLook for cleanup and repair opportunities.Run ScanOctober DealsAmazon USDeal season is back - check today's better picksAmazon US: current deals, useful picks and tech finds.See Picks×
Skip to content
Citations

How to Scrape Google AI Mode: Answers, Citations, and Links as JSON

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

There is no single, documented “Google AI Mode JSON API.” For structured answers with source links, use Gemini API Grounding with Google Search. For browser-like Search HTML, apply for Google’s Researcher Result API (which is limited to eligible, non-commercial research). If you need the literal consumer AI Mode interface, use a third-party extraction service and treat its fields and markup as changeable.

What “scraping AI Mode” can mean

Google’s AI Mode is a consumer Search experience for exploratory questions. Google says it can break a question into related searches (query fan-out), then combine information and links. AI Mode and AI Overviews can use different models and techniques, so the answer and cited links can vary between requests.

That creates three technically different projects:

Approach What you receive Access and commercial terms Citation handling
Gemini API Grounding with Google Search Model-generated text plus citation annotations and search steps Official API feature; model and tool availability can change URL, title, and start/end offsets can associate a source with answer text
Search Researcher Result API HTML that Google would return to a browser for a Search URL Eligibility and application required; Google limits the program to non-commercial research HTML is not a documented, stable AI Mode JSON schema
Commercial third-party extraction Provider-defined parsed blocks, references, and sometimes HTML Subject to the provider’s terms and endpoint limits Fields, references, and markup can be optional or change without notice

Choose the first path when your real requirement is an answer, source URLs, and exact citation spans in JSON. Choose the second only for qualifying research that needs browser-style Search output. Use the third when you specifically need the consumer interface and accept parser maintenance.

Path 1: Generate grounded answers with citation spans

Grounding with Google Search is the supported answer-plus-citations workflow. It is not a promise that Gemini’s response is identical to the consumer AI Mode response for the same prompt. Your program receives model output, and the response can include inline URL citation annotations. Google’s examples also expose the searches and result steps used for grounding.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Python example

Install the current Google Gen AI SDK and set GEMINI_API_KEY. Check Google’s current model and tool documentation before pinning a model name; availability changes.

pip install -U google-genai
export GEMINI_API_KEY="YOUR_API_KEY"

The script below asks a grounded question, preserves answer text, and walks the serialized response to collect citation objects. SDK versions may represent annotations slightly differently, so the recursive walk deliberately tolerates extra wrapper fields.

import json
import os
from google import genai
from google.genai import types

client = genai.Client(api_key=os.environ["GEMINI_API_KEY"])
question = "Compare passkeys and passwords for a small business and cite the sources used."

response = client.models.generate_content(
    model="gemini-2.5-flash",
    contents=question,
    config=types.GenerateContentConfig(
        tools=[types.Tool(google_search=types.GoogleSearch())]
    ),
)

# Convert the SDK object to ordinary JSON data.
if hasattr(response, "model_dump"):
    data = response.model_dump(exclude_none=True)
elif hasattr(response, "to_dict"):
    data = response.to_dict()
else:
    data = json.loads(response.model_json())

texts = []
citations = []
search_steps = []

def walk(value):
    if isinstance(value, dict):
        # Text content blocks are retained in encounter order.
        if isinstance(value.get("text"), str):
            texts.append(value["text"])
        citation = value.get("url_citation")
        if isinstance(citation, dict):
            citations.append({
                "url": citation.get("url"),
                "title": citation.get("title"),
                "start_index": citation.get("start_index"),
                "end_index": citation.get("end_index"),
            })
        # Keep grounding/search metadata when the SDK exposes it.
        if "search_entry_point" in value or "web_search_queries" in value:
            search_steps.append(value)
        for child in value.values():
            walk(child)
    elif isinstance(value, list):
        for child in value:
            walk(child)

walk(data)
result = {
    "answer": "".join(texts),
    "citations": citations,
    "search_steps": search_steps,
}
print(json.dumps(result, ensure_ascii=False, indent=2))

A citation’s start_index and end_index identify the answer span it supports. Keep those offsets with the exact text returned by the API; changing whitespace or concatenating blocks differently can make offsets point to the wrong characters.

Turn offsets into linked JSON

For downstream clients, store a citation object rather than replacing text immediately:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
{
  "answer": "Passkeys resist phishing ...",
  "citations": [
    {
      "url": "https://example.com/source",
      "title": "Source title",
      "start_index": 0,
      "end_index": 31
    }
  ]
}

At render time, validate that each index is within the answer length, sort citations by start_index, and escape the URL and title before producing HTML. If overlapping annotations occur, retain both in JSON and decide in your renderer whether to display multiple source links.

Path 2: Google’s Search Researcher Result API

The Researcher Result API returns the HTML Google would send to a browser for Search URLs. It is useful when a research project needs browser-equivalent Search markup, but it does not document a stable JSON representation of AI Mode answers.

  • Access is eligibility-gated and requires an application.
  • Only Search URLs are accepted; non-search URLs and some parameters are rejected.
  • Projects have request limits measured on a rolling 24-hour basis.
  • Google’s Researcher Program terms restrict use to non-commercial purposes.

Therefore, do not use this API as a commercial AI Mode scraper or assume that an answer container, citation list, or CSS class will always exist. Save the request URL, response time, and raw HTML, then parse defensively. Treat an absent AI answer as a valid “no answer returned” result rather than a parsing exception.

Path 3: Third-party AI Mode extraction

Commercial providers expose their own browser or parsing layers. For example, Scrape.do documents an AI Mode endpoint that can return parsed text blocks and references and can optionally include HTML. Its documentation notes that fields are optional, shopping cards and references vary, and Google’s raw markup class names can change without warning.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

That means your adapter should:

  • Check whether an answer block is present before reading it.
  • Accept empty or partial references.
  • Record the provider response and capture time for audits.
  • Validate every returned URL and associate it with the actual text supplied.
  • Keep selectors and field names in one versioned adapter so markup changes do not break the rest of your pipeline.

A vendor’s parsed response is an implementation choice, not a Google interface contract. Confirm current terms, data handling, limits, and response behavior before putting it into production.

Building a reliable JSON pipeline

Use a versioned schema

Store the prompt, locale, request timestamp, answer text, citations, and any search-step metadata. Add a status such as ok, no_answer, blocked, or provider_error. This distinguishes “Google returned no AI content” from “your parser failed.”

Preserve the response that produced each citation

AI Mode links are response-specific, not a fixed bibliography. Query fan-out and model changes can produce different supporting pages for the same wording. Keep the original answer and citation offsets together instead of maintaining a global citation cache.

Validate and sanitize

  • Reject citation indexes that are negative or beyond the answer length.
  • Normalize neither URLs nor titles until after storing the original values.
  • Escape text when converting JSON to HTML.
  • Set timeouts and retry only transient transport failures; repeated retries do not create an answer when Google returns none.
  • Log provider, model, locale, and capture time so a later difference is explainable.

Eligibility and publishing considerations

Google Search Central says pages must be indexed and eligible for a Search snippet to qualify as supporting links in AI Overviews or AI Mode. Google also says there are no special technical requirements or special schema.org markup for these features. Ordinary crawling, indexing, and Search best practices still apply, and meeting them does not guarantee inclusion.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Search Console reports AI-feature appearances within overall Web search traffic rather than as a separate Search type. Do not infer that every page cited in one answer will be cited again; links vary with the question, model, and underlying searches.

Troubleshooting

The SDK rejects the grounding tool

Model and tool names change. Upgrade the SDK, check the current Gemini documentation, and verify that the selected model supports Google Search grounding. Do not silently fall back to an ungrounded response if citations are required.

The response has text but no citations

Grounding may not have been used, the model may have returned no source annotations, or your parser may be looking in the wrong SDK field. Save the serialized response, inspect its content blocks and grounding metadata, and handle an empty citation array explicitly.

Offsets do not line up

Offsets apply to the returned text block. Do not trim, translate, or concatenate text before applying them. Validate bounds and retain block boundaries if your renderer cannot safely merge them.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Researcher API requests fail

Check that the URL is a supported Google Search URL, that your project is approved, and that you have not exceeded the rolling request allowance. The API is not a general web fetcher and is not approved for commercial scraping under the Researcher Program terms.

A third-party parser suddenly returns empty fields

Google can change its markup, and a particular query may genuinely have no AI Mode answer. Preserve the raw response, classify the result as empty versus parser-error, and update the provider adapter only after inspecting the new structure.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup

If you need a visual capture of a Search page for QA or an audit rather than citation-annotated JSON, ScreenshotNeo provides a direct screenshot API and MCP server. It removes cookie banners, newsletter popups, and chat widgets before capture. Bot checks, blank pages, timeouts, failed loads, and cache hits are not billed, and response headers identify the page verdict and billing status. AI agents can use its MCP tools—take_screenshot, get_page_info, and capture_pdf.

For a URL capture, see the full options in the ScreenshotNeo documentation:

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://www.google.com/search?q=passkeys -o shot.webp

Python:

import requests
r = requests.get("https://api.screenshotneo.com/v1/shot", params={"access_key": "YOUR_API_KEY", "url": "https://www.google.com/search?q=passkeys"}, timeout=90)
open("shot.webp", "wb").write(r.content)

Node.js:

const q = new URLSearchParams({ access_key: 'YOUR_API_KEY', url: 'https://www.google.com/search?q=passkeys' });
const res = await fetch(`https://api.screenshotneo.com/v1/shot?${q}`);

The Free plan includes 1,000 screenshots each month with no card; paid plans start at $5 for 3,000 shots. Create a free ScreenshotNeo account.

FAQ

Is Gemini grounding the consumer AI Mode API?

No. It is an official grounded-generation feature that returns Gemini output with Search citations; the available documentation does not establish identical wording or links to the consumer AI Mode page.

Can I use the Researcher Result API for a paid data product?

Google’s Researcher Program terms restrict it to non-commercial use, so it should not be positioned as a commercial scraping service.

Are AI Mode citations a complete source list?

No. They are the sources shown for that particular response, and the set can change with query fan-out, models, and other Search systems.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Frequently Asked Questions

Is Gemini grounding the consumer AI Mode API?

No. It is an official grounded-generation feature that returns Gemini output with Search citations; the available documentation does not establish identical wording or links to the consumer AI Mode page.

Can I use the Researcher Result API for a paid data product?

Google’s Researcher Program terms restrict it to non-commercial use, so it should not be positioned as a commercial scraping service.

Are AI Mode citations a complete source list?

No. They are the sources shown for that particular response, and the set can change with query fan-out, models, and other Search systems.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Leave a Reply

Your email address will not be published. Required fields are marked *

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Read next

Recommended PC Tool
Recommended PC Tool
Windows Errors? Fix Them Before They SpreadFree repair scan
Crashes, No Sound, or Screen Glitches?Free driver scan

Two free Windows tools

One Free Minute Could Fix That PC

Before you go - each of these free tools takes about a minute and tackles what quietly slows a Windows PC down.

Special offer. View Outbyte info, uninstall instructions, EULA, and Privacy Policy.