Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Some links on this page are affiliate links: if you buy through them we may earn a commission, at no extra cost to you.

Gemini can extract and organize information from web pages, but it is not a general-purpose crawler that automatically traverses an entire website. Use URL Context when you already know which public URLs to inspect; use Google Search grounding when you need Gemini to discover relevant public pages and return cited answers. For either route, specify the fields you need, preserve the source evidence, and validate the results in your own code.

Choose the right Gemini retrieval method

Gemini offers two documented ways to bring web content into a request. They solve different page-selection problems, and neither promises complete coverage of a site.

Use URL Context for pages you already know

Google describes URL Context as a way to provide URLs as additional context to a model. Supply the full, publicly accessible URLs you want considered, then ask Gemini to extract, compare, or summarize information from those pages. URL Context retrieves the URLs you supply; it does not follow links found on those pages. Google AI for Developers: URL Context

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Google documents a two-stage retrieval process: URL Context first attempts to use an internal index cache, then falls back to a live fetch if the URL is unavailable in that cache. That is an implementation detail, not a promise that every response reflects the latest version of every page.

Use Google Search grounding when you need discovery

Search grounding lets Gemini use public web content to answer a request and return URL citations. Google says it connects the model to real-time web content and works with all available languages. The model may decide to run one or more searches; your application should not assume that every request issues exactly one query. Returned annotations can associate answer segments with source URLs. Google AI for Developers: Grounding with Google Search

A practical combination is to ask Search grounding to find relevant public pages, then use URL Context to examine selected URLs in more depth. This remains targeted retrieval, not guaranteed exhaustive domain crawling.

Use an external search API for a private or specialized corpus

For content in a private or specialized index, Google Cloud documents a Vertex AI pattern in which a customer-provided search API returns relevant snippets for Gemini to use. This is a separate integration from public Google Search grounding; the high-level documentation does not determine the right deployment, cost, or suitability for a particular workload. Google Cloud: Grounding with your search API

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What URL Context can and cannot retrieve

Google’s current URL Context documentation lists a maximum of 20 URLs in one request and a maximum of 34 MB of retrieved content per URL. These are documented product limits, not independent benchmark results; check the live documentation before implementation because limits can change. Google AI for Developers: URL Context

Access and supported formats

  • Pages must be publicly accessible. A login, paywall, private network address, localhost URL, or tunneling service is not a supported way to supply content.
  • Supported text-oriented formats listed by Google include HTML, JSON, plain text, XML, CSS, JavaScript, CSV, and RTF.
  • Listed image formats include PNG, JPEG, BMP, and WebP; PDF is also supported.
  • Google lists paywalled content, YouTube URLs, Google Workspace files such as Docs and Sheets, and audio or video files as unsupported.
  • Use complete URLs including the protocol, such as https://example.com/page. Do not assume that linking pages will also be retrieved.

Search grounding has a different job: it discovers public web material and returns citations, rather than promising access to any particular URL or every page about a subject. URL citations improve traceability but do not guarantee that an answer is complete or correct.

Build a reliable extraction workflow

  1. Decide how pages will be selected. If you have the URLs, use URL Context. If you need public-web discovery, use Search grounding. For a private index, assess the external search API route.
  2. Check access and format. Confirm that each page is public, supported, within the per-request URL limit, and not blocked by a login or paywall. Split a larger set into multiple requests.
  3. Define the fields and evidence you need. State what counts as a value, what to return when it is missing, and whether each field should include a short supporting excerpt. Avoid asking for vague “all data.”
  4. Request a predictable structure. Specify field names and types, or use a supported JSON schema. Structured output controls the shape of the response, not the truth of its contents.
  5. Preserve provenance. Keep the URL associated with each extracted record. With Search grounding, retain returned URL annotations and map the relevant citation to the answer segment or field in your application.
  6. Validate before use. Check required fields, data types, dates, duplicates, and implausible values in ordinary application code. Treat missing content as a retrieval or extraction outcome to investigate, not automatically as proof that the real-world value is absent.
  7. Route failures explicitly. Retry or investigate pages that could not be retrieved. Do not silently merge failed pages with successful records or interpret a missing result as a negative finding.

Example extraction contract

For a set of known pages, a prompt can make the task and expected output explicit:

For each supplied URL, extract the page title and the publication date if visible on the page.
Return one record per URL with these fields:
- url: the exact supplied URL
- title: string or null
- publication_date: string or null; preserve the date as shown
- evidence: short verbatim text from the page supporting each non-null value
Do not infer a date from unrelated page content. If a value is not present or the page cannot be accessed, explain that in a status field rather than guessing.

For production use, define allowed types and required fields in a schema supported by the model and API configuration you choose. Google documents structured outputs with built-in tools, including URL Context and Search, as a Gemini 3 preview feature. Check the current model and feature documentation before relying on it. Google AI for Developers: Structured outputs

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Use Search grounding and citations responsibly

Grounded answers can include annotations that associate response text with URLs. Preserve those annotations instead of storing only a flattened answer: a later reviewer should be able to tell which source supported a claim. Put citations near the claims they support when presenting results to a person.

A citation is evidence of a source association, not proof that the model interpreted the page correctly or found every relevant page. For consequential fields, verify the cited page and extracted value. If the application needs a fixed set of known sources, direct URL Context retrieval may be easier to audit than relying on model-driven discovery.

How to get structured JSON without trusting it blindly

There are two separate requirements in an extraction application: valid output structure and valid data. A schema can help constrain a response to specified keys, types, and allowed values. It does not establish that a price, date, name, or other extracted fact is correct, current, or complete.

  • Make nullable fields explicit so an unavailable value is not replaced by a plausible guess.
  • Validate the response against the schema in your application before writing it to a database.
  • Keep source URLs and evidence alongside field values, not just in a separate unlinked citation list.
  • Use deterministic code for normalization, date parsing, duplicate detection, and range checks.
  • Send records that fail validation to a retry or human-review path rather than silently accepting them.

When Gemini is not enough for scraping

The documented Gemini routes support per-request retrieval and interpretation. The reviewed documentation does not promise exhaustive crawling, crawl scheduling, robots handling, or robust extraction from arbitrary dynamic sites. If the requirement is recurring collection across a whole domain, assess a dedicated crawler, the site’s own API, or a maintained search index. The right option depends on access rules, required coverage, and how often the data must refresh.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Before collecting from a target site, review its terms, access controls, and the rules that apply to your use and jurisdiction. The retrieval documentation does not decide whether a particular collection is permitted.

Or skip the browser setup

If the actual goal is to capture a clean visual record of pages you already identify, a screenshot API is a different tool from Gemini’s text retrieval. ScreenshotNeo accepts one GET request and returns a PNG, JPEG, WebP, or PDF; it can remove cookie/consent banners, newsletter popups, and chat widgets before capture. It bills only clean shots, not bot checks/CAPTCHAs, blank pages, timeouts, failed loads, or cache hits, and reports the page verdict and billing status in response headers. It also provides an MCP server for AI agents, with tools including take_screenshot, get_page_info, and capture_pdf.

Install Python’s requests package if needed, set your API key, then run:

import requests

r = requests.get(
    "https://api.screenshotneo.com/v1/shot",
    params={"access_key": "YOUR_API_KEY", "url": "https://stripe.com"},
    timeout=90,
)
r.raise_for_status()
open("shot.webp", "wb").write(r.content)

For the API parameters and response details, see the ScreenshotNeo documentation. Its free plan includes 1,000 screenshots a month with no card; paid plans start at $5 for 3,000. Sign up for ScreenshotNeo free.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Common problems and fixes

A supplied URL returns no useful page content

Check that the URL includes its protocol and is publicly accessible without sign-in, a paywall, or a private-network connection. Confirm the format is supported. URL Context does not follow nested links, so supply the target page directly.

A batch exceeds the documented limit

Keep each request to no more than 20 URLs and ensure the retrieved content from each URL is no more than 34 MB under Google’s documented limits. Split larger URL sets into separate requests and track which request produced each record.

The response omits a field or invents a plausible value

Clarify the extraction rule, permit null or an explicit status for missing data, and request supporting evidence. Validate each value and its evidence in application code; a schema alone does not verify correctness.

Search grounding finds too few or unexpected pages

Search grounding is model-directed discovery, not exhaustive crawling. Refine the search objective, inspect the returned citations, and use direct URL Context requests for pages that must be examined. For a requirement of systematic site-wide coverage, use a crawler, site API, or index designed for that job.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A feature or model is unavailable

Google maintains separate supported-model information for URL Context and Search grounding, and availability can change. Check the relevant live documentation and supported-model tables at implementation time rather than relying on a static model name copied into an evergreen integration.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Performance, reliability, and cost considerations

The documentation establishes retrieval limits and behaviors, but it does not establish a universal latency, success rate, or cost for a workload. Measure your own request volume, page mix, retries, and validation overhead, and consult current Gemini API or Vertex AI pricing for the service and model you actually use.

For reliability, bound batches to documented limits, retain URLs and retrieval outcomes, distinguish failed retrieval from a legitimate empty field, and make retry behavior explicit. If freshness matters, account for URL Context’s documented internal-index-first and live-fetch fallback process; it is not a per-page freshness guarantee.

Frequently Asked Questions

Can Gemini scrape a whole website automatically?

The documented URL Context and Search grounding tools do not promise exhaustive site traversal. For whole-domain or recurring collection, evaluate a crawler, site API, or search index.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Does Gemini follow links found on a URL Context page?

No. Supply the exact URLs you want URL Context to retrieve.

Can I use URL Context for pages behind a login?

No. Google documents URL Context for publicly accessible URLs and lists paywalled content as unsupported.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.